arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.00443v1 [cs.CL] 31 Aug 2026

(V)LMs generalize beyond surface co-occurrence:
Evidence from cross-modal number agreement

Zach Studdiford Affiliation: Department of Psychology Affiliation: University of Wisconsin-Madison Email: studdiford@wisc.edu    Kanishka Misra Affiliation: Department of Linguistics Affiliation: The University of Texas at Austin Email: kmisra@utexas.edu
Abstract

Language models learn about grammatical number primarily from co-occurrence, and show frequency effects as a result—sometimes taken to indicate that they do not learn abstract “rules”, and are instead dependent on specific lexical items. Testing generalization with text stimuli alone cannot settle this debate, since distributional cues (is/are, this/these) easily give number away. We instead use cross-modal generalization as a tool to investigate abstractions in LMs that can also accept visual inputs (VLMs), restricting the evidence that diagnoses number to an extra-linguistic modality. We teach VLMs pairs of new nouns by adding new embeddings and only updating them during learning, comparing conditions where number is diagnosed by visual cues alone against ones where it is disambiguated by text. Across behavior, representational dynamics, and causal mechanisms, we find non-trivial evidence for cross-modal generalization across both exposure conditions, and that linguistic vs. extra-linguistic cue conditions are treated in similar ways in the internal mechanisms of the model. This suggests that statistical learners like VLMs can generalize beyond surface-level co-occurrence and show genuine abstraction-compatible behavior.

1 Introduction

Jeff and Dave are walking in an art museum where they encounter a painting of a single, cute looking animal called “snarpus”. Next to it, Jeff sees a painting of multiple such animals, with the label: “snarpi”. Even though Jeff has never heard of snarpus or snarpi before today, he says to Dave that “these snarpi are so cute!”, using the demonstrative “these” as opposed to “this”, and the plural form are as opposed to is, owing to the fact that snarpi was used in reference to more than one of the animals, and is therefore the plural of snarpi. The made up story that we have just described is rather mundane, but it highlights our ability to abstract away our knowledge of grammatical number in a manner that goes beyond specific modalities.

Figure 1: We use cross-modal generalization as a tool to investigate abstractions in VLMs. Our case study focuses on number-agreement, where we teach VLMs new words (here, [wug] and [wugs]) that share the same textual exposure but are differentiated (in terms of number) visually. We show that VLMs can successfully generalize from this learning setup to number-agreement judgments, that generalization is characterized by movement of the new words’ embeddings towards number-based regions in the model’s representational space, and is also accommodated in causal mechanisms that were discovered before the VLMs learned the new words.

Neural network language models (LMs) learn about number agreement from co-occurrence—dog occurs more often with is, barks, etc. and dogs with are, bark, etc. (Wei et al., 2021; Hobbs and McCoy, 2026). This facet of models can (and has) been interpreted broadly in two different ways. Some suggest that by only learning from co-occurrence, LMs are prone to show disproportionately successful number agreement behavior on words for which they have seen a good deal of evidence (Wei et al., 2021; Lasri et al., 2022a; Wilson et al., 2023). This finding is often taken to mean that LMs do not necessarily learn abstract “rules”, and instead are highly dependent on specific lexical items (Lasri et al., 2022a; Wilson et al., 2023; Oba et al., 2024). Others have instead shown that co-occurrence between specific nouns and verbs (e.g., (⟨\langlecat, meows⟩\rangle, ⟨\langlelions, roar⟩\rangle) is in fact an important cue for learners (LMs and Humans alike) to acquire abstract generalization in making number agreement predictions in the first place (Hobbs and McCoy, 2026). Based on this latter viewpoint, abstractions are seen more as an emergent phenomenon that arise via accumulation of evidence from individual exemplars, rather than something that has to be built explicitly into a learner (Ambridge, 2020b; Misra and Kim, 2023; Jian and Manning, 2026; Dubova and Sloman, 2026).11 1 Our aim is not to dispute the presence of clear item-specific frequency effects we see in models (McCoy et al., 2024)—frequency effects are inevitable in any system that performs statistical learning, and are thoroughly prevalent in humans as well (Lupyan, 2013; Ambridge et al., 2015; Lampinen et al., 2024; Studdiford and Lupyan, 2026)

A historically productive method to investigate abstractions in children has been to conduct novel word learning trials (Berko, 1958; Höhle et al., 2004). In a nutshell, these experiments show children new words using a combination of pictures and/or sentence utterances, and then test how they respond to the usage of the word in a different context. This has also been done for LMs in their learning of nouns and verbs (Kim and Smolensky, 2021; Wilson et al., 2023; Misra and Kim, 2026), but in the case of of number, the exposure can easily give away information via orthography or by distributional cues (verbs like is/are, demonstrative like these/this, etc.). To what extent does a model encode abstract knowledge that cannot be explained merely by co-occurrence alone?

To answer this question, we turn to novel word learning in variants of LMs that can process inputs from extra-linguistic modalities, where we can restrict diagnostic information about the target abstraction (number) for the novel words to come from a non-linguistic modality. We then measure the extent to which models are able to learn about the number agreement for these novel words to conclude about the strength of this abstraction. That is, we use cross-modal generalization (Xu et al., 2026) as a tool to explore abstractions in (V)LMs. We do so by providing evidence from tests of model behavior, internal representational dynamics, and a battery of causal interpretability techniques.

Specifically, we teach the model pairs of new nouns that refer to objects introduced in images, by using them in textual descriptions that are uninformative with respect to number. This way surface-form co-occurrence alone cannot diagnose the novel nouns’ number. We do this by inserting new embeddings for these novel words, which then prevents the model from making conclusions using the surface form (e.g., using -s to conclude that it is plural) 22 2 We note that the VLMs we evaluate also do not have access to orthography for any sampled real nouns—i.e., both singular and plural nouns are tokenized into single tokens.. During learning, we only update the embeddings of the novel words, so that we can analyze and conclude about how the learned abstractions (insofar as they exist) in their pre-existing representations pressure the learning of new information. We compare this to a setting where there is no image and the model instead learns from textual cues, where the linguistic cues clearly diagnose number information. We then measure generalization by performing standard minimal pair analysis on English number-agreement stimuli that are disjoint from those in training. We observe that performance on nouns learned from both visual and linguistics cues is significantly above chance, and only degrades when there are intervening attractor nouns (≥\geq 2). This suggests that behaviorally, models are able to demonstrate non-trivial cross-modal generalization of number.

Next, we characterize what underlies this behavior by investigating the dynamics of the embeddings of novel words during learning. We find that learning in both cue-conditions (vision vs. language) corresponds to similar representational behavior, where novel nouns move to regions in embedding space that are inhabited by real nouns that organize themselves in terms of their number features. That is, the embedding for the novel word that is intended to be treated as singular moves towards a region of known singular nouns (e.g., dog, cat, leopard, etc.) and its plural counterpart moves towards a region of known plural nouns (e.g., dogs, cats, leopards, etc.). This suggests that these low dimensional regions sensitive to grammatical number act as basins into which novel words tend to move towards during learning.

Finally, we test if model mechanisms discovered using interpretability techniques before novel words were acquired can readily accommodate knowledge of novel words. Using four different methods, we find this to be true—the causal efficacy of all methods was substantially above chance. We additionally found no difference in the methods’ effectiveness for nouns acquired from language versus vision, suggesting that models integrated both types of nouns into their existing number-agreement mechanism in the same way.

Overall, our findings suggest that even though VLMs learn linguistic features primarily through co-occurrence, and therefore display item-specific frequency effects, this does not necessarily indicate an absence of abstraction-compatible representations. In fact, the representations that have emerged as a result of LM and VLM training can, in principle, flexibly accommodate linguistic information for tokens acquired primarily through extra-linguistic modalities. This suggests that statistical learners like VLMs can go beyond simple surface-level co-occurrence and integrate linguistic cues even from evidence completely devoid of explicit linguistic signal.

2 Related Work

Number agreement has been perhaps the oldest probe for syntactic competence in neural network models of language (Elman, 1991). In more modern instantiations, it has been evaluated via minimal pair judgments (Linzen et al., 2016; Hu et al., 2020; Mueller et al., 2020) as well as through representational analyses (Lakretz et al., 2019; Finlayson et al., 2021; Lasri et al., 2022b; Arora et al., 2024; Marks et al., 2025). It is also a common case-study for investigating language model generalization, and in particular has been shown to be affected by frequency effects (Wei et al., 2021; Lasri et al., 2022a), where models succeed on noun-verb pairs that are sufficiently frequent. Our work builds on the historical precedence of number-agreement by treating it as the target abstraction that we investigate in VLMs. We specifically test if models can demonstrate this knowledge in a manner that cannot be explained by the aforementioned frequency effects. We do this by conducting novel word learning studies where the knowledge of a noun’s number is only diagnosed by extra-linguistic evidence, and comparing it to a case where it is diagnosed by explicit linguistic context. This allows us to test if the representations that result from (V)LM training can demonstrate abstraction-compatible behavior, even though they are primarily acquired from co-occurrence.

The method that we use to conduct novel word learning (section 3.1) also traces back to early connectionist approaches (Rumelhart et al., 1993; Rogers and McClelland, 2004), where new concept and property nodes were added to a trained network that predicted concept-property associations, trained on a set of inputs by only updating the new node representations, and then tested for inductive generalization. Since then, this method has been used to conduct word learning experiments (Lampinen and McClelland, 2017), test for category learning (Kim and Smolensky, 2021; Misra and Kim, 2023), structural alternations (Wilson et al., 2023; Misra and Kim, 2026), and neologisms (Hewitt et al., 2025). We extend this method by adding an additional modality during learning (vision), and conduct further analyses of how knowledge acquired through this method is integrated into the model by investigating representational dynamics of novel word embeddings, as well as how information about their number is accounted for by interpretability methods that discover mechanisms before the novel words were acquired by the model.

3 General Methods

In this section, we describe our method for learning novel embeddings in VLMs, as well as evaluation stimuli, and the models studied. While we focus on number-agreement in this work, our training paradigm can in principle be extended to assess cross-modal generalization for a variety of linguistic abstractions.

3.1 Training Novel Embeddings

Initialization

We begin by adding new embeddings, e1,e2e_{1},e_{2} as new entries in models’ embedding (WeW_{e}) and unembedding (WuW_{u}) matrices, as well as its tokenizer. We initialize the embeddings with gaussian noise with the mean and standard deviation of 28 singular nouns and their 28 plural counterparts.33 3 Because these are balanced around equal numbers of singular and plural nouns, the embeddings themselves do not carry any bias towards a particular number Furthermore, while we use [wug] and [wugs] for convenience, the model does not see their orthography, and does not break them into wug and wug+s. We then freeze the entire model except for these newly added embeddings.

Training

As we will see below, the inputs to the model are either a simple sentence without any image, or an image along with an associated text caption. Given an input, we perform training via backpropagation with the cross-entropy loss on the textual part of the stimuli, by only updating e1e_{1} and e2e_{2}—i.e., no other parameter in the model is updated. In our experiments, we primarily report results over 50 seeds (i.e., 50 different initializations) unless stated otherwise. We train for a maximum of 20 epochs, using early stopping on an evaluation set (described briefly below) with a patience of 5. We conducted large scale hyperparameter tuning for learning rate, with details in Section B.1 and Appendix G.

Cue Conditions

We primarily compare generalization across two types of cues to the novel nouns’ number: 1) Vision, where we provide the model with images that depict one or more chimeric creatures, with a single creature mapped to [wug] and more than one of the same type mapped to [wugs]. We create these images by using the OpenAI API, with prompts shown in , and manually verified them, to maximize the chance that the VLMs have not already seen these images. We pair each image with an associated piece of text (e.g., Do you see the [wug]/[wugs]?), and importantly use text captions that do not give away any distributional cues. We compare this to 2) Language, where the surface form of the text directly disambiguates the number of the nouns. We do so by pairing the novel nouns with verbs/determiners/quantifiers of the appropriate number (e.g., The [wug] runs vs. The [wugs] run). For both cue conditions, we use 15 pairs of stimuli for each noun, amounting to a total of 30 stimuli used for training.

Halting Condition

To select the final state of the embeddings, we manually curate a development set of 280 minimal pair sentences, and evaluate by computing the percentage of time the model finds acceptable sentence more likely than unacceptable ones. Details about these sentences can be found in Appendix A.4.

3.2 Evaluation Stimuli

Following both targeted syntactic evaluation of LMs (Marvin and Linzen, 2018; Gulordava et al., 2018), as well as psycholinguistic precedence (Bock and Cutting, 1992; Franck et al., 2002; Arehalli and Linzen, 2020), we use minimal pair agreement attraction stimuli to quantify number-agreement performance in our experiments. That is, we use declarative sentences with intervening nouns that carry the opposite number feature than that of the subject of the sentence, and therefore “attract” the prediction of a verb that agrees with them instead of agreeing with the subject. In the context of both humans and LMs, we generally see a greater number of errors as the number of attractors increases (Franck et al., 2002; Arehalli and Linzen, 2020, e.g.,). We use the following template, with nn = 0,1,2,3 attractors:

The [adj] [noun]subj{}_{\texttt{subj}} {[prep] the [noun]attr{}_{\texttt{attr}}}n [verb]

where [adj], [noun]subj{}_{\texttt{subj}}, [prep], [noun]attr{}_{\texttt{attr}}, and [verb] denote adjective, subject-noun, preposition, attractor noun, and target verb respectively. We create sentence pairs by sampling disjoint combinations of items from 25 adjectives, 40 nouns, 88 attractor nouns, and 180 verb pairs (all of which are single token), where each sentence pair has the same noun phrase prefix, and only minimally differs in the verb in terms of its agreement with the subject. We sample 700 sentence pairs per attractor, giving us 2,800 total pairs. We generate these pairs with real noun subjects, and then replace them with [wug] and [wugs] when we evaluate models on number agreement for novel nouns.

3.3 Models

Our case study focused on the 2B and the 4B version of Qwen3-VL (Bai et al., 2025), both of which are post-trained VLMs. The embedding matrices are tied for both models—i.e., their unembedding and embedding matrices share the same weights. Future work can extend this method to models with untied embeddings.

4 Behavioral Evidence

Our main experiment involves testing the extent to which VLMs are able to generalize the knowledge of the novel nouns’ numbers from their exposures to cues in the two different conditions as described in Section 3.1. This experiment forms the basis of all subsequent analyses (Section 5 and Section 6). We specifically test models’ knowledge of number agreement for our novel nouns by performing minimal pair analyses on agreement-attraction stimuli as mentioned in section 3.2. We additionally compare these results to those where the subjects of the sentences are real nouns. Since the models have ostensibly encountered these real nouns far more frequently during their training, we naturally expect there to be a gap in the models’ performance. For each cue condition, we evaluate models on novel noun embeddings obtained from 50 different training seeds.

Figure 2: Behavioral accuracy (and 95% CI) of models on number-agreement stimuli with novel nouns as subjects, across different cue conditions (Language and Vision), and across different intervening attractors (0–3). Black dashed line with triangular points indicates performance on the same stimuli but with real nouns.

Figure 2 shows these results on both models across different attractors (0–3), cue-conditions (language vs. vision), and subject word types (real vs. novel). As expected, we see a noticeable gap between number agreement performance on real, natural nouns relative to that on novel nouns, and that this performance degrades with increasing number of attractors. At the same time, regardless of the cue condition, we also see that models are above chance at number agreement for all attractors, suggesting successful generalization. This is especially striking in the vision condition—our results show that even when number cannot be diagnosed by textual co-occurrence (as in the ‘Language’ case), VLMs are able to demonstrate behavior compatible with successful generalization, and by extension, successful encoding of number.

5 Representational Dynamics of Learning

Having shown that our VLMs are able to generalize number agreement for novel words learned from both vision and language cues, we now move onto analyses that characterize the internal dynamics of these words’ representations during learning. Is there a systematic pattern to how these representations change? To this end, we follow Misra and Kim (2023), and track the movement of the novel word embeddings in low dimensional subspaces that capture abstractions relevant to number agreement. We report these results for the 4B model here, and include results for the 2B model in Appendix C.

5.1 Methods

We start by performing a 2-dimensional Principal Component Analysis (PCA) on a subset of the embedding layer of our models consisting of: 1) pairs of real-world nouns whose number information is well-known (e.g., dog–dogs, car-cars, etc.); and 2) the initial and final states of the novel word pairs from all 50 seeds used in previous experiment. We take the randomly initialized versions of these novel words to be the initial state and their states at the end of each training run to be their respective final states. We repeat this for embeddings learned from language as well as vision cues. To obtain our set of real nouns, we query WordNet (Miller, 1995) and sample 500 singular nouns and their plural counterparts (amounting to a total of 1000 nouns), which exist in the single-token vocabulary of our models’ tokenizers.44 4 It is important to note that the embeddings for these real words have not changed, since we only backpropagate over the novel words’ embeddings. Upon reducing our embeddings to 2 dimensions, we then connect each novel word’s initial state to its final state using an arrow, and then visualize this movement with respect to the representations of the real nouns.

To quantify this movement, we compute the scalar projection of each of our novel word embeddings (initial and final states separately) onto the direction that captures noun-number. To obtain the number direction, we took the difference between the vector formed by averaging the embeddings of all our singular nouns (vsgv_{\textsf{sg}}) and that formed from the average embedding of all our plural nouns (vplv_{\textsf{pl}}). For a given embedding of a novel noun (en​o​v​e​le_{novel}), we then compute the scalar projection, and subsequently the movement as:

proj =enovel.vsg−vpl‖vsg−vpl‖\displaystyle=e_{\textsf{novel}}.\frac{v_{\textsf{sg}}-v_{\textsf{pl}}}{||v_{\textsf{sg}}-v_{\textsf{pl}}||} (1)
movement =projfinal−projinitial\displaystyle=\textrm{proj}_{\textsf{final}}-\textrm{proj}_{\textsf{initial}} (2)

That is, we quantify movement as the change in the projection of a novel noun embedding (final - initial) onto the singular-plural direction. Therefore, novel nouns that are supposed to be treated as singular should show positive movement, and those that are supposed to be treated as plural should show negative movement. This method has been used in the past to measure bias (Bolukbasi et al., 2016) as well as encoding of semantic features in vector space models (Grand et al., 2022).

5.2 Results

Figure 3 visualizes the first two principal components of the model embeddings across both cue conditions—language and vision. We first see that real nouns are organized according to their number features—i.e., the set of 500 singular nouns end up clumping together, and similarly so does the set of their plural counterparts. Then, regardless of the cue condition, we see non-trivial movement of the novel nouns towards their respective directions. That is, novel singular and plural nouns across both modalities move towards the regions occupied by real nouns with the respective number features.

Figure 3: Movement (shown using arrows) of the embedding states of novel words when analyzed using a 2D PCA fit on embeddings of real singular (sg) and plural (pl) nouns (e.g., dogs, chairs, blocks, etc., NN=500 each) in the Qwen3-VL-4B model across both cue conditions (Language and Vision). Colors indicate grammatical number for real and novel nouns.
Cue Condition sg pl
Language 0.092±\pm 0.006 -0.061±\pm 0.006
Vision 0.067±\pm 0.007 -0.040±\pm 0.004
Table 1: Mean movement by Cue Condition and number. Movement is calculated as the difference in the projection of the novel word embeddings’ final and initial states onto the singular-plural direction, computed as the vector difference of real singular and plural nouns. Movements are significantly different than 0 (p<p<.001).

Table 1 shows the average movement (averaged across 50 seeds) of the novel embeddings for both cue-conditions as well as for both types of number features. First, all movements are in their intended directions—we see positive movements for novel embeddings that are supposed to be singular and negative movements for embeddings that are supposed to be plural (p<p<.001 for both). Next, we see that movement is on average greater for novel words whose number is diagnosed from language cues than from visual cues. Overall, even though there is no diagnostic information in the language component of the visual cue, we still see non-trivial movement during learning towards the space of desirable exemplars. This suggests that representational movement towards regions inhabited by exemplars that bear the target feature (here, number) underlies abstraction compatible behavior in VLMs, regardless of modality.

6 Mechanistic Evidence

Results from the previous experiment shed light on what happens inside the embeddings of novel words as they are being integrated within the models’ existing embedding layer. How is this information transmitted from the input embedding to the rest of the model in a manner that enables it to produce the right output? For this, we turn to a slew of modern mechanistic interpretability methods. To us, the goal of these methods is to describe the internal mechanisms that underlie an abstraction. If these mechanisms are discovered before our novel nouns (e.g., [wug] and [wugs]) are added to the model, then to what extent do they make consistent predictions about novel words after they are acquired? This also lets us shed light on the generalization capabilities of these methods in a manner that goes beyond stimuli that the models have seen at the time of mechanism discovery.

6.1 Methods

We rely on methods that involve counterfactual interventions on the models’ internal components in a manner that manipulates the model output (Arora, Mueller). Our methods follow the framework of Arora et al. (2024), where each method involves pairs of stimuli, base and source, with their corresponding next word labels, yby_{b} and ysy_{s}. We use these stimuli to intervene on a model component, ff, by taking its value from the source and applying it to base using some sort of a method-dependent transformation to produce f∗f^{*}. Insofar as this intervention succeeds, we should observe a change in the intervened model’s output probabilities pf←f∗p_{f\leftarrow f^{*}} that is compatible with the output when the input was the source stimulus, relative to the original model probabilities pfp_{f}. We measure the efficacy of the intervention using the log odds-ratio which compares the probabilities of the source and base next-word labels before and after intervention:

Odds (f,f∗,⟨b,s,yb,ys⟩)\displaystyle(f,f^{*},\langle b,s,y_{b},y_{s}\rangle)
=log⁡(pf​(yb∣b)pf​(ys∣b)⋅pf←f∗​(ys∣b)pf←f∗​(yb∣b))\displaystyle=\log\left(\frac{p_{f}(y_{b}\mid b)}{p_{f}(y_{s}\mid b)}\cdot\frac{p_{f\leftarrow f^{*}}(y_{s}\mid b)}{p_{f\leftarrow f^{*}}(y_{b}\mid b)}\right)

We validate our methods by fitting/finding interventions on a train set (except for one of the methods, which is unsupervised) and evaluating on a held-out test set, and in this case, we only use stimuli with real words, since we perform these interventions before novel words are added to the model. Our train and test stimuli are sampled using our attractor stimuli method described in section 3. In particular, we split these stimuli (keeping singular and plural ones balanced) such that the set of subjects, attractor nouns, and target verb pairs are completely disjoint between train (NN=700 per attractor) and test (NN=700 per attractor). We evaluate using the average log odds-ratio computed over our test set.

Given the aforementioned analysis design, we run experiments on four intervention methods:55 5 Detailed description of each method is provided in Appendix D.

Refer to caption
Figure 4: Avg. Odds across model layers (when applicable) for the Qwen3-VL-4B model on agreement stimuli with real nouns as subjects. Higher values mean greater causal effect.
Figure 5: Avg. Odds for the Qwen3-VL-4B model on agreement stimuli with subjects that are real vs. novel nouns acquired from both types of Cue conditions. Results shown for layer with best overall avg. odds chosen on results on stimuli with real nouns as subjects. We see generally high agreement in the results across both cue-conditions.

Distributed Alignment Search (DAS)

This method (Geiger et al., 2021; Geiger et al., 2024) learns a subspace in a model’s activations in a manner that maximizes the likelihood of a given counterfactual completion token on performing interventions. We restrict this rotation to be 1-dimensional, and apply it layer-wise. This method is by definition sensitive to changes in the full model behavior.

DiffMean

This method (Marks and Tegmark, 2023) intervenes by adding an offset vector (or its negative) with a coefficient (±α​𝐯\pm\alpha\mathbf{v}) that is computed using the difference in the average vector of activations (at each layer) per class (here, singular and plural). Since this method is unsupervised, we directly applied it to the test set, and use α=50\alpha=50.

Probe

We fit linear probes (Alain and Bengio, 2017; Ettinger et al., 2016)—in our case, using logistic regression—on model activations (across layers) for the train set and then perform interventions by adding the learned weights along with a coefficient (±α​𝐰\pm\alpha\mathbf{w}, where 𝐰\mathbf{w} is the learned weight vector of logistic regression, and α\alpha is set to 50).

Attribution Patching (AtP)

This method (Nanda, 2023) approximates the individual effects of a set of model components using the gradients computed from a single backwards pass over those components. In our case, we compute these gradients with respect to the logit difference of the correct versus incorrect sentence completion (in short, estimating the components with greatest contribution in producing the model completion "The wugs are" versus "The wugs is"). To produce a causal effect in the model, we then extract the top-kk activations (as ranked by the attribution method) in the forward pass for a given singular prompt and patch those activations into the forward pass for its plural complement (and vice versa).

When applicable, we apply the above methods layer-wise66 6 We apply them at every 5th layer and additionally the last layer for each model and on the last token position prior to the model prediction. To evaluate generalization to novel words, we select the layer (when applicable) with the best average log odds-ratio. Then, after adding our novel nouns into the embedding layer, we further evaluate the method on our test set now with the subject nouns appropriately replaced with our novel nouns. We repeat this process for both language and vision cue conditions.

6.2 Results

Figure 4 shows results from our validation experiments, where we report Avg. Odds across methods, and when applicable, layers, on stimuli with real nouns as subjects. We see that the Avg. Odds across all methods are above 0, indicating qualitatively successful interventions. Similarly to Arora et al. (2024), we find DAS to be most causally efficacious, with comparatively higher Avg. Odds across attractors and layers.

Turning now to results reported in Figure 5, we investigate how well these methods incorporate information from novel nouns that have been acquired after the mechanisms were discovered, across both cue conditions (Language and Vision). These results only include Avg. Odds from layers that were selected to be the best in the previous analysis on real nouns. We again observe Avg. Odds that are generally above 0, across all methods and for both cue conditions. More interestingly, with the exception of DAS for 0 attractors, we see no visual difference between the Language and Vision cue conditions. To test this further, we fit a linear mixed effects model on all our results (across all layers, and both models), where we predict the avg. odds using cue condition, attractors, model, and interpretability method as fixed effects, and layer as random effects. Here, we find no significant effect of the cue condition (βc​u​e\beta_{cue}=-0.016, pp=.92).77 7 For full results, see Appendix E. Overall, this suggests that linguistic evidence is not privileged in mechanisms responsible for number-agreement behavior in models, and that words whose number information is learned exclusively from vision are integrated in the same way as those learned from explicit linguistic evidence.

7 General Discussion and Conclusion

Using analyses of model behavior, internal representational dynamics, and causal mechanisms, we find that VLMs can demonstrate cross-modal generalization of number agreement. That is, they can learn about a novel noun’s number from both linguistic and visual cues, especially even if all they know about the noun is from visual evidence alone. This learning is characterized by emergent abstraction of the number feature, which is learned for words already present in the model embeddings and which guide the learning of this information for these new nouns. Finally, the internal mechanisms responsible for number agreement in the model are able to account for data the model acquires after they have been discovered, and they treat nouns learned from language vs. vision alone in the same way.

Overall, these findings indicate that emergent abstractions in data-driven learners like VLMs are more flexible than their initial mode of acquisition—LMs primarily learn about number from linguistic cues: the pairing of nouns with commonly co-occurring verbs. When they eventually learn to map words to images (and vice-versa) during VLM training, their distributional evidence is further bolstered, but it is bolstered for nouns for which they already have strong evidence (for number). Our results show that even when they learn nouns only from vision, the tested VLMs’ learned linguistic abstractions are prepared to accommodate this new knowledge in a manner that is seemingly indistinguishable from nouns that are acquired only from text. This suggests that generalization in these models might be learned from surface-form co-occurrence, but that they can generalize beyond this knowledge and show evidence of abstraction-compatible behavior. Abstraction and item-based learning might not be so different after all (Ambridge, 2020b; Ambridge, 2020a; Misra and Kim, 2023).

Limitations

While we have, in as much detail as possible, provided in-depth evidence for cross-modal generalization for number-agreement, there are several ways in which our study is limited. In particular, our focus is only on two VLMs from the same family with largely lower parameter sizes, with limited number of images, and on a single language (English). Overall, while we cannot make strong claims about VLMs as a class, our results, along with several others who have conducted studies on a single model (Petty et al., 2022; Misra and Kim, 2023; Jian and Manning, 2026, etc.) suggest that it is, in principle, possible for abstractions to emerge via item-based learning.

Furthermore, our experiments only use a single method to conduct novel word learning—this was a deliberate choice, because many alternatives (Teehan et al., 2024; Wang et al., 2025) often invoke a separate model or an atypical training process (e.g., meta learning based fine-tuning). Using these would have prevented us from drawing conclusions about the target model (cf. the idiosyncrasies of a particular method).

References

  • Alain and Bengio (2017) G. Alain and Y. Bengio Understanding intermediate layers using linear classifier probes. In The Fifth International Conference on Learning Representations, External Links: Link Cited by: Appendix D, §6.1.
  • Ambridge et al. (2015) B. Ambridge, E. Kidd, C. F. Rowland, and A. L. Theakston The ubiquity of frequency effects in first language acquisition. Journal of child language 42 (2), pp. 239–273. Cited by: footnote 1.
  • Ambridge (2020a) B. Ambridge Abstractions made of exemplars or ‘You’re all right, and I’ve changed my mind’: Response to commentators. First Language 40 (5-6), pp. 640–659. Cited by: §7.
  • Ambridge (2020b) B. Ambridge Against stored abstractions: A radical exemplar model of language acquisition. First Language 40 (5-6), pp. 509–559. Cited by: §1, §7.
  • Arehalli and Linzen (2020) S. Arehalli and T. Linzen Neural language models capture some, but not all, agreement attractioneffects. In Proceedings of the Annual Meeting of the Cognitive Science Society, Vol. 42. Cited by: §3.2.
  • Arora et al. (2024) A. Arora, D. Jurafsky, and C. Potts CausalGym: benchmarking causal interpretability methods on linguistic tasks. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 14638–14663. External Links: Link, Document Cited by: Appendix D, §2, §6.1, §6.2.
  • Bai et al. (2025) S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, W. Ge, Z. Guo, Q. Huang, J. Huang, F. Huang, B. Hui, S. Jiang, Z. Li, M. Li, M. Li, K. Li, Z. Lin, J. Lin, X. Liu, J. Liu, C. Liu, Y. Liu, D. Liu, S. Liu, D. Lu, R. Luo, C. Lv, R. Men, L. Meng, X. Ren, X. Ren, S. Song, Y. Sun, J. Tang, J. Tu, J. Wan, P. Wang, P. Wang, Q. Wang, Y. Wang, T. Xie, Y. Xu, H. Xu, J. Xu, Z. Yang, M. Yang, J. Yang, A. Yang, B. Yu, F. Zhang, H. Zhang, X. Zhang, B. Zheng, H. Zhong, J. Zhou, F. Zhou, J. Zhou, Y. Zhu, and K. Zhu Qwen3-VL Technical Report. External Links: 2511.21631, Link Cited by: §3.3.
  • Berko (1958) J. Berko The child’s learning of english morphology. Word 14 (2-3), pp. 150–177. Cited by: §A.2, §1.
  • Bock and Cutting (1992) K. Bock and J. C. Cutting Regulating mental energy: performance units in language production. Journal of memory and language 31 (1), pp. 99–127. Cited by: §3.2.
  • Bolukbasi et al. (2016) T. Bolukbasi, K. Chang, J. Zou, V. Saligrama, and A. T. Kalai Man is to computer programmer as woman is to homemaker? debiasing word embeddings. In Advances in Neural Information Processing Systems, D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett (Eds.), Vol. 29, pp. . External Links: Link Cited by: §5.1.
  • Dubova and Sloman (2026) M. Dubova and S. J. Sloman Excess capacity learning. Behavioral and Brain Sciences, pp. 1–77. External Links: Document Cited by: §1.
  • Elman (1991) J. L. Elman Distributed representations, simple recurrent networks, and grammatical structure. Machine learning 7 (2), pp. 195–225. External Links: Link Cited by: §2.
  • Ettinger et al. (2016) A. Ettinger, A. Elgohary, and P. Resnik Probing for semantic evidence of composition by means of simple classification tasks. In Proceedings of the 1st Workshop on Evaluating Vector-Space Representations for NLP, pp. 134–139. Cited by: Appendix D, §6.1.
  • Finlayson et al. (2021) M. Finlayson, A. Mueller, S. Gehrmann, S. Shieber, T. Linzen, and Y. Belinkov Causal analysis of syntactic agreement mechanisms in neural language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), C. Zong, F. Xia, W. Li, and R. Navigli (Eds.), Online, pp. 1828–1843. External Links: Link, Document Cited by: §2.
  • Franck et al. (2002) J. Franck, G. Vigliocco, and J. Nicol Subject-verb agreement errors in french and english: the role of syntactic hierarchy. Language and cognitive processes 17 (4), pp. 371–404. Cited by: §3.2.
  • Geiger et al. (2021) A. Geiger, H. Lu, T. Icard, and C. Potts Causal abstractions of neural networks. In Advances in Neural Information Processing Systems, M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. W. Vaughan (Eds.), Vol. 34, pp. 9574–9586. External Links: Link Cited by: §6.1.
  • Geiger et al. (2024) A. Geiger, Z. Wu, C. Potts, T. Icard, and N. Goodman Finding alignments between interpretable causal variables and distributed neural representations. In Causal Learning and Reasoning, pp. 160–187. Cited by: §6.1.
  • Grand et al. (2022) G. Grand, I. A. Blank, F. Pereira, and E. Fedorenko Semantic projection recovers rich human knowledge of multiple object features from word embeddings. Nature Human Behaviour, pp. 1–13. External Links: Link Cited by: §5.1.
  • Gulordava et al. (2018) K. Gulordava, P. Bojanowski, E. Grave, T. Linzen, and M. Baroni Colorless green recurrent networks dream hierarchically. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), M. Walker, H. Ji, and A. Stent (Eds.), New Orleans, Louisiana, pp. 1195–1205. External Links: Link, Document Cited by: §3.2.
  • Hewitt et al. (2025) J. Hewitt, R. Geirhos, and B. Kim We can’t understand ai using our existing vocabulary. arXiv preprint arXiv:2502.07586. Cited by: §2.
  • Hobbs and McCoy (2026) C. Hobbs and R. T. McCoy Collocational bootstrapping: a hypothesis about the learning of subject-verb agreement in humans and neural networks. In Proceedings of the 30th Conference on Computational Natural Language Learning, C. Bonial and Y. Berzak (Eds.), San Diego, California, USA, pp. 90–103. External Links: Link, Document, ISBN 979-8-89176-410-1 Cited by: §1.
  • Höhle et al. (2004) B. Höhle, J. Weissenborn, D. Kiefer, A. Schulz, and M. Schmitz Functional elements in infants’ speech processing: the role of determiners in the syntactic categorization of lexical elements. Infancy 5 (3), pp. 341–353. Cited by: §1.
  • Hu et al. (2020) J. Hu, J. Gauthier, P. Qian, E. Wilcox, and R. P. Levy A systematic assessment of syntactic generalization in neural language models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault (Eds.), Online, pp. 1725–1744. External Links: Link, Document Cited by: §2.
  • Jafari et al. (2025) F. R. Jafari, O. Eberle, A. Khakzar, and N. Nanda RelP: faithful and efficient circuit discovery in language models via relevance patching. arXiv preprint arXiv:2508.21258. Cited by: §F.2.
  • Jian and Manning (2026) J. Jian and C. D. Manning Humans and transformer LMs: abstraction drives language learning. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), V. Demberg, K. Inui, and L. Marquez (Eds.), Rabat, Morocco, pp. 752–765. External Links: Link, Document, ISBN 979-8-89176-380-7 Cited by: §1, Limitations.
  • Kim and Smolensky (2021) N. Kim and P. Smolensky Testing for grammatical category abstraction in neural language models. In Proceedings of the Society for Computation in Linguistics 2021, A. Ettinger, E. Pavlick, and B. Prickett (Eds.), Online, pp. 467–470. External Links: Link Cited by: §1, §2.
  • Lakretz et al. (2019) Y. Lakretz, G. Kruszewski, T. Desbordes, D. Hupkes, S. Dehaene, and M. Baroni The emergence of number and syntax units in LSTM language models. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), J. Burstein, C. Doran, and T. Solorio (Eds.), Minneapolis, Minnesota, pp. 11–20. External Links: Link, Document Cited by: §2.
  • Lampinen et al. (2024) A. K. Lampinen, I. Dasgupta, S. C. Chan, H. R. Sheahan, A. Creswell, D. Kumaran, J. L. McClelland, and F. Hill Language models, like humans, show content effects on reasoning tasks. PNAS nexus 3 (7), pp. pgae233. Cited by: footnote 1.
  • Lampinen and McClelland (2017) A. K. Lampinen and J. L. McClelland One-shot and few-shot learning of word embeddings. arXiv preprint arXiv:1710.10280. External Links: Link Cited by: §2.
  • Lasri et al. (2022a) K. Lasri, A. Lenci, and T. Poibeau Does BERT really agree ? fine-grained analysis of lexical dependence on a syntactic task. In Findings of the Association for Computational Linguistics: ACL 2022, S. Muresan, P. Nakov, and A. Villavicencio (Eds.), Dublin, Ireland, pp. 2309–2315. External Links: Link, Document Cited by: §1, §2.
  • Lasri et al. (2022b) K. Lasri, T. Pimentel, A. Lenci, T. Poibeau, and R. Cotterell Probing for the usage of grammatical number. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), S. Muresan, P. Nakov, and A. Villavicencio (Eds.), Dublin, Ireland, pp. 8818–8831. External Links: Link, Document Cited by: §2.
  • Linzen et al. (2016) T. Linzen, E. Dupoux, and Y. Goldberg Assessing the ability of LSTMs to learn syntax-sensitive dependencies. Transactions of the Association for Computational Linguistics 4, pp. 521–535. External Links: Document Cited by: §2.
  • Lupyan (2013) G. Lupyan The difficulties of executing simple algorithms: why brains make mistakes computers don’t. Cognition 129 (3), pp. 615–636. Cited by: footnote 1.
  • Marks et al. (2025) S. Marks, C. Rager, E. Michaud, Y. Belinkov, D. Bau, and A. Mueller Sparse feature circuits: discovering and editing interpretable causal graphs in language models. In International Conference on Learning Representations, Vol. 2025, pp. 23888–23923. External Links: Link Cited by: §2.
  • Marks and Tegmark (2023) S. Marks and M. Tegmark The geometry of truth: emergent linear structure in large language model representations of true/false datasets. arXiv preprint arXiv:2310.06824. Cited by: Appendix D, §6.1.
  • Marvin and Linzen (2018) R. Marvin and T. Linzen Targeted syntactic evaluation of language models. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, E. Riloff, D. Chiang, J. Hockenmaier, and J. Tsujii (Eds.), Brussels, Belgium, pp. 1192–1202. External Links: Link, Document Cited by: §3.2.
  • McCoy et al. (2024) R. T. McCoy, S. Yao, D. Friedman, M. D. Hardy, and T. L. Griffiths Embers of autoregression show how large language models are shaped by the problem they are trained to solve. Proceedings of the National Academy of Sciences 121 (41), pp. e2322420121. Cited by: footnote 1.
  • Miller (1995) G. A. Miller WordNet: a lexical database for English. Communications of the ACM 38 (11), pp. 39–41. Cited by: §5.1.
  • Misra and Kim (2023) K. Misra and N. Kim Abstraction via exemplars? A representational case study on lexical category inference in BERT. In BUCLD 48: Proceedings of the 48th annual Boston University Conference on Language Development, Boston, USA. External Links: Link Cited by: §1, §2, §5, §7, Limitations.
  • Misra and Kim (2026) K. Misra and N. Kim Generating novel experimental hypotheses from language models: a case study on cross-dative generalization. arXiv preprint arXiv:2408.05086. Cited by: §1, §2.
  • Misra (2022) K. Misra Minicons: enabling flexible behavioral and representational analyses of transformer language models. arXiv:2203.13112. External Links: Link Cited by: §B.3, Appendix G.
  • Mueller et al. (2020) A. Mueller, G. Nicolai, P. Petrou-Zeniou, N. Talmina, and T. Linzen Cross-linguistic syntactic evaluation of word prediction models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault (Eds.), Online, pp. 5523–5539. External Links: Link, Document Cited by: §2.
  • Nanda (2023) N. Nanda Attribution patching: activation patching at industrial scale. URL: https://www. neelnanda. io/mechanistic-interpretability/attribution-patching 15, pp. 17. Cited by: Appendix D, §6.1.
  • Oba et al. (2024) M. Oba, Y. Oseki, A. Fukatsu, A. Haga, H. Ouchi, T. Watanabe, and S. Sugawara Can language models induce grammatical knowledge from indirect evidence?. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 20591–20603. External Links: Link, Document Cited by: §1.
  • Petty et al. (2022) J. Petty, M. Wilson, and R. Frank Do language models learn position-role mappings?. In Proceedings of the 46th Annual Boston University Conference on Language Development, Y. Gong and F. Kpogo (Eds.), Somerville, MA, pp. 657–671. Cited by: Limitations.
  • Rogers and McClelland (2004) T. T. Rogers and J. L. McClelland Semantic cognition: a parallel distributed processing approach. MIT press. Cited by: §2.
  • Rumelhart et al. (1993) D. E. Rumelhart P. M. Todd et al. Learning and connectionist representations. Attention and performance XIV: Synergies in experimental psychology, artificial intelligence, and cognitive neuroscience 2, pp. 3–30. Cited by: §2.
  • Studdiford and Lupyan (2026) Z. Studdiford and G. Lupyan Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning. arXiv preprint arXiv:2606.13607. Cited by: footnote 1.
  • Teehan et al. (2024) R. Teehan, B. Lake, and M. Ren CoLLEGe: concept embedding generation for large language models. In First Conference on Language Modeling, External Links: Link Cited by: Limitations.
  • Wang et al. (2025) W. Wang, G. Jiang, T. Linzen, and B. Lake Rapid word learning through meta in-context learning. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 32038–32073. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: Limitations.
  • Wei et al. (2021) J. Wei, D. Garrette, T. Linzen, and E. Pavlick Frequency effects on syntactic rule learning in transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M. Moens, X. Huang, L. Specia, and S. W. Yih (Eds.), Online and Punta Cana, Dominican Republic, pp. 932–948. External Links: Link, Document Cited by: §1, §2.
  • Wilson et al. (2023) M. Wilson, J. Petty, and R. Frank How abstract is linguistic generalization in large language models? experiments with argument structure. Transactions of the Association for Computational Linguistics 11, pp. 1377–1395. External Links: Link, Document Cited by: §1, §1, §2.
  • Wolf et al. (2020) T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. Le Scao, S. Gugger, M. Drame, Q. Lhoest, and A. Rush Transformers: state-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Q. Liu and D. Schlangen (Eds.), Online, pp. 38–45. External Links: Link, Document Cited by: Appendix G.
  • Wu et al. (2024) Z. Wu, A. Geiger, A. Arora, J. Huang, Z. Wang, N. Goodman, C. Manning, and C. Potts Pyvene: a library for understanding and improving PyTorch models via interventions. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: System Demonstrations), K. Chang, A. Lee, and N. Rajani (Eds.), Mexico City, Mexico, pp. 158–165. External Links: Link Cited by: Appendix G.
  • Xu et al. (2026) T. Xu, M. Sandoval-Castañeda, K. Livescu, G. Shakhnarovich, and K. Misra Cross-modal taxonomic generalization in (vision-) language models. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 16319–16337. External Links: Link, Document, ISBN 979-8-89176-390-6 Cited by: §1.

Appendix A Additional Details of Embeddings Training

A.1 Embeddings initialization

We initialize all [wug] and [wugs] embeddings instances at the mean of an equal number of real singular and plural nouns (n=28n=28 pairs). Because the set of singular and plural embeddings used to extract the mean is evenly balanced, there is no information in this embedding biasing learning towards singular or plural nouns. Table 2 shows all (single token) nouns used for initialization.

cat cats dog dogs bird birds
bear bears rat rats tree trees
word words thing things car cars
house houses rock rocks chair chairs
table tables cup cups book books
phone phones man men woman women
child children person people door doors
gate gates fence fences pond ponds
lamp lamps corner corners path paths
wall walls
Table 2: The 28 noun pairs used for embeddings initialization.

A.2 Vision condition

Image stimuli

Our image stimuli consist of five images of one [wug] creature (generated using the OpenAI API), and five images of two [wugs] creatures. The full set of images can be seen in Figure 6 below.

Refer to caption
Figure 6: Image stimuli for [wug] and [wugs] training in the vision condition.

Image stimuli generation prompts

We additionally include the prompts used to generate images used for training [wug] and [wugs] embeddings. These queries were used to generate five singular and plural images from the LLM openai-o3 via the OpenAI online API. Because the actual word wug (Berko, 1958) and wug illustrations are very likely part of large model pretraining corpora, we query openai-o3 to generate images of one or multiple “snarples” (a made up by one of the authors), as seen below:

Image-generation prompts used to create singular “snarple” stimuli

generate an image of an imaginary creature called a snarple
now generate an image of the snarple facing a different direction
awesome! Now generate it facing the other direction
now facing forward again
now generate an image of singular snarple

Image-generation prompts used to create plural “snarple” stimuli

now generate an image of two snarples
generate two more images of two snarples each
now do another image of two snarples
now another and they are standing up
awesome now one more sitting again

Text stimuli

Text stimuli in the vision condition were intentionally designed not to include any discriminating syntactic information for the [wug] or [wugs] embeddings. To this end, we manually created 30 sentences and randomly assigned 15 to co-occur with [wug] and the remaining 15 to co-occur with [wugs]. Table 3 shows stimuli used for [wug] and [wugs] in each condition.

Singular ([wug]) Plural ([wugs])
[wug]? [wugs].
[wug], there. [wugs]!
The [wug] from over there. [wugs] here.
[wug] over there. The [wugs] from over here.
[wug] near the fence. [wugs] by the rock.
[wug] around the chair. [wugs] behind the tree.
Do you see the [wug] near the fence? Look at the [wugs] by the rock.
I walked past the [wug] along the path. I noticed the [wugs] behind the tree.
I found the [wug] under the lamp. I pointed to the [wugs] in the corner.
I stood beside the [wug] by the gate. I watched the [wugs] from a distance.
I looked for the [wug] around the chair. I moved toward the [wugs] near the pond.
I noticed the [wug] in the corner by the lamp. Look at the [wugs] near the rock by the fence.
I pointed to the [wug] behind the chair near the wall. I walked past the [wugs] near the gate by the path.
I found the [wug] near the pond by the tree. I watched the [wugs] under the lamp by the door.
Do you see the [wug] by the fence near the gate? I moved toward the [wugs] around the chair by the rock.
Table 3: Training stimuli for the image condition. Each sentence is paired with a singular or plural creature image.

A.3 Language condition

Text stimuli

Text stimuli (n=30n=30) in the language condition were designed to include syntactic cues informative of whether [wug] and [wugs] embeddings corresponded to singular or plural nouns. All stimuli were verified prior to inclusion.

Singular ([wug]) Plural ([wugs])
One [wug]. Several [wugs].
A single [wug]. A few [wugs].
There sits a [wug]. There sit some [wugs].
I spotted a [wug]. I spotted some [wugs].
That is definitely a [wug]. Those are definitely [wugs].
The [wug] shook its tail before running away. The [wugs] shook their tails before running away.
Beneath the stairs crouched a frightened [wug]. Beneath the stairs crouched several frightened [wugs].
Does the [wug] always come back at night? Do the [wugs] always come back at night?
One [wug] carries a small leaf in its mouth. Many [wugs] carry small leaves in their mouths.
The [wug] near the window makes a soft sound. The [wugs] near the window make a soft sound.
A [wug] appeared from behind the curtain. A dozen [wugs] appeared from behind the curtain.
That [wug] seems friendlier than the others. Those [wugs] seem friendlier than the others.
Along the path I discovered a lone [wug]. Along the path I discovered a group of [wugs].
The [wug] was climbing the hill before sunset. The [wugs] were climbing the hill before sunset.
Curled up in the corner was a sleeping [wug]. Curled up in the corner were several sleeping [wugs].
Table 4: Training stimuli for the language condition.

A.4 Dev-Set Sentence Construction

As a criteria for halting embeddings training after some number of epochs, we create a “dev-set” consisting of n=280n=280 minimal pairs sentences evaluating number agreement in [wug] and [wugs] embeddings across a range of constructions. We first manually created a limited set of constructions, and then augmented these futher using a frontier LLM (openai-o3). All stimuli were then manually verified.

A.5 Instruction Template Format

Because Qwen-3-VL-* models post-training included a chat template, we also interpolate our prompts in chat templates both for training and evaluation. We use the following chat template for training in the vision condition:

Chat template for the image condition.

<|im_start|>user
<|vision_start|><|image_pad|><|vision_end|>Caption this image.<|im_end|>
<|im_start|>assistant
[wug], there.<|im_end|>

And the following chat template in all other cases (language condition, evaluations):

Chat template for the language condition and for all evaluations.

<|im_start|>user
Complete the sentence.<|im_end|>
<|im_start|>assistant
One [wug].<|im_end|>

Appendix B Additional Results of Embeddings Training

B.1 LR hyperparameter sweep

We systematically sweep across a range of learning rates for language and vision conditions in both models (training at five random seed initializations at each learning rate evaluated). We find that 10−310^{-3} is the optimal learning rate for both models and conditions, and use this learning rate for all subsequent embeddings training. Figure 7 shows the results of the full LR sweep.

Figure 7: Results of LR sweep for all models and learning conditions.

B.2 Loss curves for learned embeddings

We include loss curves for the mean cross-entropy loss of [wug] and [wugs] tokens in the vision and language conditions for both models, for sentences seen in training. Loss curves are averaged across the full set of n=50n=50 runs for different seed initializations at the best performing learning rate. These loss trajectories can be seen in Figure 8. Interestingly, we note greater within-distribution loss for the language (text based) condition as opposed to vision.

Figure 8: Cross-entropy loss for wug and wugs training, for 50 seeds in each condition.

B.3 Examples of productive generation of novel embeddings

The bulk of our main results compare relative log probabilities between [wug] and [wugs]. However, are models able to actually generate the novel learned embeddings when appropriate? We use minicons library (Misra, 2022) to generate next token completions for several base sentences contain [wug] or [wugs] tokens following embeddings training in Qwen-3-VL-2B-Instruct. Qualitatively, we note that the model is indeed able to productively generate the next token: that is, the relative probabilities for [wug] and [wugs] are not buried in the greater logit distribution (nor does the model deterministically generate only [wug] or [wugs]), at least for these examples. Table 5 shows base sentences and completions for Qwen-3-VL-2B-Instruct.

# Prompt Model generation
1 One [wug] was playing and another came to join it. Now there are One [wug] was playing and another came to join it. Now there are [wugs]
2 I saw a single [wug] yesterday. Today I saw three I saw a single [wug] yesterday. Today I saw three [wugs]
3 There is one [wug] on the left and two There is one [wug] on the left and two [wugs] on the right.
4 I saw one [wug]. Then another joined. now there are I saw one [wug]. Then another joined. now there are [wugs]
5 If a [wug] is lonely then it should find [wug] If a [wug] is lonely then it should find [wugs]
6 There is one [wug] on the left and one There is one [wug] on the left and one [wug] on the right.
Table 5: Generations from Qwen3-VL-2B with the learned syntax [wug]/[wugs] embeddings. Generated tokens are in bold

Appendix C Results on Qwen3-VL-2B

C.1 Replication of embeddings movement during training

We replicate the results shown in Figure 3 for Qwen-3-VL-2B as shown in 9, observing qualitatively similar results. Both [wug] and [wugs] nouns representations diverge along trajectories defined by real singular and plural nouns.

Figure 9: Movement (shown using arrows) of the embedding states of novel words when analyzed using a 2D PCA fit on embeddings of real singular (sg) and plural (pl) nouns (e.g., dogs, chairs, blocks, etc., NN=500 each) in the Qwen3-VL-2B model across both cue conditions (Language and Vision). Colors indicate grammatical number for real and novel nouns.

C.2 Mechanistic evidence for Qwen3-VL-2B

Results on performing mechanistic analyses for Qwen3-VL-2B are shown in figure 10 and section C.2. We see generally similar results as in the 4B model, albeit with slightly weaker avg. odds in the novel word results. Nevertheless, we see further evidence of no difference between language vs. vision cue conditions.

Refer to caption
Figure 10: Avg. Odds across model layers (when applicable) for the Qwen3-VL-2B model on agreement stimuli with real nouns as subjects. Higher values mean greater causal effect.
Figure 11: Avg. Odds for the Qwen3-VL-2B model on agreement stimuli with subjects that are real vs. novel nouns acquired from both types of Cue conditions. Results shown for layer with best overall avg. odds chosen on results on stimuli with real nouns as subjects. We see generally high agreement in the results across both cue-conditions.

Appendix D Details of Intervention Methods

Distributed Alignment Search (DAS) (Arora et al., 2024)

: DAS learns a 1-dimensional subspace in model activations with the objective of maximizing the likelihood of a given counterfactual completion token for some input sequence. Specifically, given a source input 𝐬\mathbf{s} and a base input 𝐛\mathbf{b} we wish to intervene on, DAS learns the rotation:

h′=h𝐛+(h𝐬​𝐚⊤−h𝐛​𝐚⊤)​𝐚,h^{\prime}=h_{\mathbf{b}}+\big(h_{\mathbf{s}}\mathbf{a}^{\top}-h_{\mathbf{b}}\mathbf{a}^{\top}\big)\,\mathbf{a}, (3)

where h𝐛,h𝐬∈ℝ1×dh_{\mathbf{b}},h_{\mathbf{s}}\in\mathbb{R}^{1\times d} are the residual stream activations at layer ℓ\ell and some token position pp for the base and source inputs respectively, and 𝐚∈ℝ1×d\mathbf{a}\in\mathbb{R}^{1\times d} is the learned direction. In our case, 𝐬\mathbf{s} and 𝐛\mathbf{b} are a minimal-difference singular–plural sentence pair ⟨ss,sp⟩\langle s_{s},s_{p}\rangle, where 𝐚\mathbf{a} is learned to maximize the relative likelihood of the counterfactual completion, log⁡p⁡(sscf)>log⁡p⁡(ssbase)\log p(s_{s}^{\text{cf}})>\log p(s_{s}^{\text{base}}). The explicit supervised learning signal in DAS (which learns the rotation h′h^{\prime} over many labeled pairs) guarantees that the method will converge to the 1-dimensional subspace which captures the difference in model behavior between s and b given one exists.

Linear Probe (Alain and Bengio, 2017; Ettinger et al., 2016):

Given the set of residual stream activations for singular and plural sentences ⟨ss,sp⟩\langle s_{s},s_{p}\rangle at a given layer ℓ\ell, we apply a logistic regression to classify activations from labeled singular (y=0y=0) and plural (y=1y=1) sentences. Given the regression for classifying singular and plural sentences:

p⁡(y=1∣h)=σ⁡(h​𝐰⊤)p(y=1\mid h)=\sigma(h\mathbf{w}^{\top}) (4)

we intervene on model representations using the weight vector 𝐰∈ℝ1×d\mathbf{w}\in\mathbb{R}^{1\times d}, adding either +α​𝐰^+\alpha\hat{\mathbf{w}} to the residual stream to increase the likelihood of the plural completion or −α​𝐰^-\alpha\hat{\mathbf{w}} to increase the likelihood of the singular completion, where α\alpha is a steering coefficient.

Difference-of-Means (DiffMean) (Marks and Tegmark, 2023):

Given the set of activations for singular and plural sentences ⟨ℋs,ℋp⟩\langle\mathcal{H}_{s},\mathcal{H}_{p}\rangle at layer ℓ\ell, difference-in-means uses the vector

𝐯=μp−μs\mathbf{v}=\mu_{p}-\mu_{s} (5)

as the intervention on a given minimal-difference pair ⟨ss,sp⟩\langle s_{s},s_{p}\rangle, adding +α​𝐯^+\alpha\hat{\mathbf{v}} to increase the likelihood of the plural completion and −α​𝐯^-\alpha\hat{\mathbf{v}} for the singular.

Attribution patching (Nanda, 2023):

Rather than finding a subspace in the residual stream, Attribution patching first estimates the set of individual model components with the greatest estimated effect on the target behavior via a gradient based approximation of activation patching, retaining the top-kk parameters based on the attribution value:

c^​(n)=𝔼(𝐛,𝐬)[(n⁡(𝐬)−n⁡(𝐛))⊤​∇nℒ​(𝐛)].\hat{c}(n)=\mathop{\mathbb{E}}_{(\mathbf{b},\mathbf{s})}\big[(n(\mathbf{s})-n(\mathbf{b}))^{\top}\nabla_{n}\mathcal{L}(\mathbf{b})\big]. (6)

where n⁡(𝐛),n⁡(𝐬)n(\mathbf{b}),n(\mathbf{s}) are the activations of component nn under the base and source inputs, and ∇nℒ​(𝐛)\nabla_{n}\mathcal{L}(\mathbf{b}) is the gradient of the difference in logits between the completion agreeing with the source sentence and the completion agreeing with the base sentence ℒ=logit⁡(v𝐬)−logit⁡(v𝐛)\mathcal{L}=\mathrm{logit}(v_{\mathbf{s}})-\mathrm{logit}(v_{\mathbf{b}}) (e.g. logit⁡(are)−logit⁡(is)\mathrm{logit}(\textit{are})-\mathrm{logit}(\textit{is})). We then patch the activations for the top-kk model components between base and source sentences. While attribution patching can in principle be applied to any set of model components, we focus on MLP intermediate activations (which we find obtain the best performance on natural sentences, see Appendix F.2).

Appendix E Statistical Analyses for Cross-modal Mechanistic Analyses

To understand the effect of our various considerations in the mechanistic analyses, we conduct a statistical test using linear mixed effects regression model which we fit on results from both models, across all attractors, and all layerwise methods. We predict avg. odds using cue condition, attractors, model, and interpretability method as fixed effects, and layer as random effects:

AvgOdds ∼cue+attractors+method\displaystyle\sim\texttt{cue}+\texttt{attractors}+\texttt{method}
+model+(1∣layer)\displaystyle+\texttt{model}+(1\mid\texttt{layer})

Results from this analysis is shown in Table 6. We see that except for the model and cue condition, all other predictors show significant effects. For instance, both DiffMeans and Probe have worse odds than DAS (as clearly seen in our results), while odds decrease with increase in the number of attractors.

Term β\beta SE df tt pp
(Intercept) 4.613 0.375 13.0 12.31 <0.001
cueVision{}_{\textit{Vision}} −-0.016 0.157 322.7 −-0.10 0.920
attractors −-0.930 0.070 322.7 −-13.25 <0.001
methodDiffMean{}_{\textit{DiffMean}} −-1.861 0.192 322.7 −-9.68 <0.001
methodProbe{}_{\textit{Probe}} −-1.895 0.192 322.7 −-9.86 <0.001
modelQwen3-VL-4B{}_{\textit{Qwen3-VL-4B}} 0.313 0.168 328.8 1.86 0.064
Table 6: Linear mixed-effects model results

Appendix F Additional Mechanistic Results

F.1 Replication of mechanistic results across multiple random seeds

We verify that the generalization of our mechanistic results is not attributable to the random initialization or training dynamics of a particular [wug] and [wugs] embedding, replicating our analyses for four additional embeddings pairs. Figures 12 and 13 show results for Qwen-3-VL-2B and Qwen-3-VL-4B, respectively. We find that these results are highly consistent with our main mechanistic findings as described in Section 6.

Figure 12: Mechanistic results for LR sweeps (n=5n=5 learning rates) in Qwen-3-VL-2B across all conditions.
Figure 13: Mechanistic results for LR sweeps (n=5n=5 learning rates) in Qwen-3-VL-4B for all conditions.

F.2 Selecting top-kk activations and activation basis for attribution patching

It is an empirical question whether a certain activation basis allows for more effective causal interchanges in the attribution patching method, or how the effectiveness of interchanges changes for different numbers of top-kk activations. We compare the mean odds of patching top-kk activations from attention heads, the residual stream, and MLP activations pre and post-linearity, for k=23​…​27k={2^{3}...2^{7}} activations in Qwen-3-VL-2B. These results can be seen in Figure 14. We observe the greatest effectiveness of causal interventions using the top k=128k=128 MLP intermediate activations as an activation basis in natural stimuli, consistent with findings from Jafari et al. (2025). Interestingly, we find differing results in generalization to [wug] and [wugs] embeddings.

Figure 14: Efficacy of various activation bases and values of top-kk activations for causal interventions in Qwen-3-VL-2B.

Appendix G Implementation and Compute Resources

All experiments in this paper were conducted on either an NVIDIA H-100 GPU or an NVIDIA RTX 6000 Ada GPU. We used Adam as the optimization method for training all embeddings and causal intervention representations (for methods with a supervised learning objective). We sampled from the following range of learning rates in our grid search for optimal LR in both Qwen-3-VL-2B and Qwen-3-VL-4B: 0.0001, 0.0003, 0.0005, 0.0007, 0.0009, 0.001, 0.003, 0.005, 0.0075, 0.01, 0.03, 0.05, 0.075, 0.1. All our code is implemented in python, with DAS being implemented in pyvene (Wu et al., 2024), log-probabilities and generation using minicons (Misra, 2022) and transformers libraries (Wolf et al., 2020).

Appendix H AI usage disclosure statement

LLMs (specifically, OpenAI o3) were used in this project for the purposes of generating novel image and text stimuli.