arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2609.00515v1 [cs.CL] 01 Sep 2026

The Interlingua Hypothesis: LLMs Translate
via a Latent Task-agnostic Feature Space

Jacob Brinton Jannik Brinkmann Mark Crovella Aaron Mueller Email: jbrin@bu.edu, jannik.brinkmann@uni-mannheim.de, Email: crovella@bu.edu, amueller@bu.edu Affiliation: Boston University Technische Universität Clausthal
Abstract

Large language models (LLMs) have recently demonstrated improved machine translation performance over strong supervised baselines. This raises questions as to what mechanisms underlie how LLMs perform machine translation between languages. Motivated by recent interpretability findings—namely, that LLMs use massively multilingual latent feature representations to perform language modeling—we propose the interlingua hypothesis. The hypothesis holds that language models translate by reading a source sentence into a latent feature space, and generate a target sentence by reading from the latent feature space. We show three lines of evidence in support of this hypothesis: (1) variance in BLEU across language pairs is largely predictable from language-specific competences with no language pair–specific interaction terms; (2) many model components are causally influential in both monolingual tasks and translation tasks; and (3) fine-tuning on monolingual data recovers a large proportion of translation improvements relative to fine-tuning on aligned documents. Together, these provide convergent evidence in support of the interlingua hypothesis, and suggest new ways of understanding and improving how LLMs can be leveraged to perform translation tasks.

1 Introduction

Machine translation has long been an influential area of natural language processing. The attention mechanism was first proposed to improve machine translation performance (Bahdanau et al., 2015; Luong et al., 2015; Vaswani et al., 2017); one consequence of attention was the emergence of language models that could be efficiently trained on massive corpora (Devlin et al., 2019; Radford et al., 2018). Recently, large language models (LLMs) have revolutionized many areas of natural language processing, but have been relatively slowly adopted in machine translation (MT).

A reason for the slow adoption of LLMs in MT has been a focus on low-resource languages, where LLMs face significant challenges (Hendy et al., 2023; Robinson et al., 2023). Because LLMs require large training corpora, they are difficult to train effectively on low-resource languages. However, they also hold significant promise: even when not trained on a given language, LLMs can be prompted to translate from or into a language they have seen very little of in their training data (Cahyawijaya et al., 2024)—e.g., by prompting with a grammar book (Tanzer et al., 2024). Moreover, their representations of high-resource languages could enable cross-lingual transfer (Conneau et al., 2020).

Whether LLMs can deliver on this promise depends in part on the mechanisms underlying how they translate. Specifically, do LLMs rely on task- and language-agnostic mechanisms to translate, or do they rely on more specialized translation mechanisms or language pair–specific mechanisms? If the former, this suggests clear paths toward improving MT performance, potentially without large parallel corpora. Much recent work provides mechanistic evidence supporting the existence of massively multilingual representations in language modeling contexts (Wendler et al., 2024; Brinkmann et al., 2025; Wu et al., 2025).

We therefore hypothesize that LLMs perform MT in large part by reusing computational mechanisms for monolingual language modeling. The “stages of inference” hypothesis (Lad et al., 2024) holds that LLMs devote the first half of their layers to reading an input and composing increasingly abstract concept representations. Then, the latter half of the model reads these concept representations, and uses them to decide which token should be predicted given prior context. If language models perform translation this way (what we call the interlingua hypothesis), then translation would not necessarily require language pair–specific mechanisms; instead, it only requires that a model be capable of reading abstract features from the source language, and generating text for the target language conditioned on those features.

If this hypothesis is true, it would entail the following predictions: (1) Given language pair (S,T)(S,T), translation performance should be predictable from monolingual capabilities in SS and TT in isolation without cross-linguistic interaction terms. (2) There should exist task-agnostic representations that are causally relevant for predicting correct outputs in both machine translation and monolingual task settings. Finally, (3) adding translation capabilities for a new language should primarily require improvements to monolingual capabilities for that language---i.e., fine-tuning on monolingual corpora should recover a substantial proportion of the improvement in translation performance as fine-tuning on parallel corpora, assuming the model can already effectively handle the other language in the language pair.11 1 Note that our claims relate to the mechanisms a fully trained LLM uses to perform translation. We do not make claims regarding what is contained in the pretraining data; it is possible that parallel pretraining data is necessary for multilingual representations to be learned.

We investigate each of these three implications, and find positive evidence for each. While no single experiment definitively confirms the interlingua hypothesis, each provides a different type of evidence in support of it. These findings could provide a preliminary explanation for the value of exposure to (largely monolingual) documents in producing models more effective at translation.

2 Related Work

Massively multilingual feature representations.

Our work is partially motivated by the observation that grammatical concept representations are highly multilingual in large language models, even across typologically distinct languages (Brinkmann et al., 2025). Similar evidence has been observed in Wendler et al. (2024); Dumas et al. (2025). Notably, intervening on grammatical concept representations has predictable effects on model behavior in both monolingual and machine translation settings (Brinkmann et al., 2025). These findings do not in themselves confirm the interlingua hypothesis, but they do suggest the existence of abstract feature representations not tied to particular tasks nor languages.

Interlingua in multilingual NMT.

A line of work in neural machine translation (NMT) seeks to engineer interlingua to improve cross-lingual transfer via architectural and training improvements, largely to encoder-decoder architectures (e.g., XLM-R, Conneau et al., 2020; M2M, Fan et al., 2020). Prior approaches introduce explicit shared representations via interlingua layers or networks (Lu et al., 2018; Zhu et al., 2020), or bottlenecks that naturally encourage language-independent representations (Vázquez et al., 2019; Mao et al., 2023). These studies stipulate or engineer an interlingua and validate it behaviorally, whereas our study asks whether an interlingua emerges naturally and is causally relevant to MT performance in a decoder-only LM trained on general language data with no such directly implemented incentives.

Precedents to our mechanistic investigation include Vázquez et al. (2020), who investigate whether sentence representations in NMT models with attention bottlenecks are language-independent. Others have investigated whether shared or language-specific components are needed at all (Escolano et al., 2021; Purason and Tättar, 2022); these studies investigate by controlling the degree of parameter sharing across languages, and find that fully shared representations underperformed representations with at least some language-specific components. Thus, in smaller-scale settings using primarily parallel data, interlingua typically need to be imposed via the training objective, and this costs some performance. We investigate whether this also holds in contemporary decoder-only models trained on a larger quantity of task- and domain-general language data.

LLMs for machine translation.

Machine translation is difficult when parallel data is limited (Koehn and Knowles, 2017), and large language models do not solve this problem (Court and Elsner, 2024). Some hope that LLMs could enable cross-lingual transfer. For example, Tanzer et al. (2024) showed that long-context LLMs can translate a previously unseen low-resource language when prompted with a grammar book containing linguistic descriptions and translation examples. However, Aycock et al. (2025) subsequently found that most of the improvement in this setting came from the book’s parallel examples rather than its grammatical explanations, highlighting the importance of parallel data for translation. One of our experiments asks a complementary question: once an LLM has already acquired multilingual representations during pretraining, to what extent can improving its monolingual competence in a low-resource language improve translation without additional parallel data? We provide preliminary evidence for cross-lingual transfer in §5.

Causal mediation analysis.

Some of our evidence relies on estimates of the causal influence (Lewis, 1973) of specific model components. This relies on causal mediation analysis (Pearl, 2001; Vig et al., 2020), a common technique in the mechanistic interpretability literature (Finlayson et al., 2021; Geiger et al., 2021; Heimersheim and Nanda, 2024; Marks et al., 2025; Mueller et al., 2026, e.g.,). While the purpose of this study is not to understand the exact functional role of particular model components, we wish to characterize whether the same representations influence a model’s behavior in multiple task settings and language pairs.

3 Modeling Translation Performance as a Function of Monolingual Capabilities

How much of a model’s ability to translate can be explained by its monolingual language modeling capabilities? If a model uses an interlingua to perform translation, then much of its translation capabilities should be explainable as a function of its ability to parse latent features from an input, and produce coherent text in that language. In other words, one should not need language pair–specific terms to explain translation performance, unless a model deploys some separate translation mechanism that does not involve going through an interlingua. To investigate, we define a linear model that predicts translation performance from monolingual competencies.

We use several measures of monolingual competence. First, to measure a model’s ability to distinguish grammatical from ungrammatical sentences, we use MultiBLiMP (Jumelet et al., 2026), a multilingual version of the BLiMP (Warstadt et al., 2020) benchmark. For a given language, MultiBLiMP accuracy is the fraction of minimal pairs for which the model assigns higher full-sentence log-probability to the grammatical completion over its ungrammatical counterpart. We average this over 200 randomly-selected samples per language. Because Llama and Aya saturate this benchmark for many of the languages in our analysis, we also compute the log-probability margin, which is the mean log⁡p⁡(grammatical)−log⁡p⁡(nongrammatical)\log p(\text{grammatical})-\log p(\text{nongrammatical}) per minimal pair. To overcome issues related to benchmark saturation, we also include accuracy on GlobalMMLU (Singh et al., 2025), a multilingual multiple-choice question answering dataset based on MMLU (Hendrycks et al., 2021). For each task, we filter the set of languages to those in common between the monolingual dataset and FLORES (Table 1).22 2 The following languages are supported across all datasets we use in this section: ara, ces, deu, ell, eng, fas, fra, heb, hin, ita, nld, pol, por, ron, rus, spa, tur, ukr

For translation, we have language pair (S,T)(S,T), where SS is the source and TT is the target language. Translation competence tS​Tt_{ST} is quantified as a BLEU score for (S,T)(S,T). To perform translation, we use a 2-shot prompt containing two randomly sampled sentences from FLORES.33 3 We ensure that the examples do not overlap with the test example. We use sacrebleu (Post, 2018) to compute BLEUs.

Let ll be a monolingual competence estimate using MultiBLiMP or GlobalMMLU. We have lSl_{S} and lTl_{T} for each language pair. We then learn coefficients βS\beta_{S} and βT\beta_{T} as well as bias term β0\beta_{0} to maximize predictive accuracy on BLEU scores tS​Tt_{ST} in the following function:

βS​lS+βT​lT+β0=tS​T\beta_{S}l_{S}+\beta_{T}l_{T}+\beta_{0}=t_{ST} (1)

We also train a bilinear model that is nearly identical, but also contains a multiplicative interaction term:

βS​lS+βT​lT+βS​T​(lS⋅lT)+β0=tS​T\beta_{S}l_{S}+\beta_{T}l_{T}+\beta_{ST}(l_{S}\cdot l_{T})+\beta_{0}=t_{ST} (2)

If the interlingua hypothesis holds, then the bilinear model should not have significantly greater predictive power than the linear model. This would imply that machine translation performance is better explained by monolingual terms, rather than by the existence of a translation mechanism (which should use terms specific to particular language pairs).

Proxy Measures NN
FLORES perplexity fluency / compression (↓\downarrow) 24
MultiBLiMP accuracy grammatical acceptability (↑\uparrow) 18
MultiBLiMP margin grammatical confidence (↑\uparrow) 18
GlobalMMLU accuracy world knowledge (↑\uparrow) 23
Table 1: Monolingual competence proxies. NN is the number of the 24 FLORES languages on which the proxy is defined; ↑\uparrow/↓\downarrow indicates the direction of higher monolingual competence.

Monolingual task performance predicts translation competence.

Fitting the linear model on the common 18-language subset, monolingual behavioral competence proxies are generally strong predictors of translation performance. MultiBLiMP grammatical accuracy reaches R2=0.294R^{2}=0.294/0.2350.235 for Llama/Aya. While significant, this is relatively low; we find that this is largely because MultiBLiMP scores saturate at relatively low BLEU scores, such that it is a good predictor, but only up to middling BLEU scores. In contrast, GlobalMMLU accuracy is a very strong predictor at R2=0.739R^{2}=0.739/0.5100.510 for Llama/Aya. Figure 1 plots monolingual competencies against translation performance for each language shared across each evaluation dataset. Qualitatively, the grammatical and MMLU proxies trend closely with translation quality.

Llama-3.1-8B Aya-23-8B
Proxy Rlin2R^{2}_{\text{lin}} Rbil2R^{2}_{\text{bil}} MAE Rlin2R^{2}_{\text{lin}} Rbil2R^{2}_{\text{bil}} MAE
Per-token perplexity 0.046 0.046 5.48 0.002 0.003 4.89
MultiBLiMP accuracy 0.294 0.294 4.50 0.235 0.236 4.22
MultiBLiMP margin 0.324 0.324 4.82 0.233 0.234 4.27
GlobalMMLU accuracy 0.739 0.741 2.91 0.510 0.510 3.56
Table 2: Comparison of the power of monolingual tasks in predicting BLEU scores using the linear (Rlin2R^{2}_{\text{lin}}) and bilinear (Rbil2R^{2}_{\text{bil}}) models with several monolingual competence proxies on 18 languages. GlobalMMLU is the strongest predictor. The bilinear interaction term does not add significant predictive power over the linear model.

Adding language pair interaction terms does not increase predictive power.

Across each proxy and both models, the bilinear model has virtually the same predictive power as the linear model (Table 2; see also App. A and App. A.1).

The same trend holds when we predict BLEU scores from the language-specific marginal BLEU scores, computed by averaging BLEU scores across all language pairs for a given source or target language. Decomposing the full BLEU matrix into per-language source and target main effects, BLEUS​T≈μ+αS+βT\text{BLEU}_{ST}\approx\mu+\alpha_{S}+\beta_{T}, explains R2=0.932R^{2}=0.932 (Llama) / 0.8790.879 (Aya) of the centered variance, and a rank-1 multiplicative reconstruction recovers ≈\approx90%90\%/83%83\% of the matrix energy (App. A.2). Hence, pairwise translation quality is largely captured by per-language competence, with pairwise interactions adding little predictive power over the simpler model.

Refer to caption
Figure 1: Per-language relationship between two monolingual proxies and the BLEU score for a given target language. Both monolingual terms correlate significantly with BLEU scores, but GlobalMMLU accuracies are a stronger predictor than MultiBLiMP accuracies.

Target language competence matters more than source language competence.

Source and target competence contribute asymmetrically: changing the target language moves BLEU far more than changing the source. The variance across target languages of mean BLEU exceeds the variance across source languages by 9.3×9.3\times for Llama and 3.9×3.9\times for Aya, and the target coefficient is 1.71.7–5×5\times the source coefficient (MMLU target/source =2.98=2.98 for Llama, 1.671.67 for Aya; MultiBLiMP-margin 4.884.88/2.302.30). This may be because it is easier for models to extract meaning from the source language than produce fluent outputs in the target language; alternatively, it may be an artifact of relying on an n-gram matching metric such as BLEU score.

4 Many Translation-relevant Components Also Perform Monolingual Tasks

Our previous results show that monolingual capabilities are predictors of translation performance, but also that language pair interactions are not strong predictors. This provides correlational evidence in support of our hypothesis, but not causal evidence. To obtain causal evidence, we now perform an analysis based on causal mediation analysis (Pearl, 2001; Vig et al., 2020).

Past work has found that language models use the first half of their layers to form progressively more abstract representations of the latent features in a given input (Lad et al., 2024). We hypothesize that language models could repurpose this language modeling machinery to perform translation by reading the source language into a latent feature space, and then using the latter half of its layers to read from this feature space and produce the translation in the output language. If this is true, then we should be able to find model representations or components that are causally influential in both machine translation and language modeling settings. To obtain causal evidence in support of this view, we now perform mechanistic experiments by patching model components. We show that the most influential attention heads for producing correct translations also strongly mediate a language model’s ability to produce grammatical outputs in monolingual settings.

4.1 Using GCM to Identify Translation-relevant Components

Refer to caption
Figure 2: Llama-3.1-8B mean |IE^||\widehat{\mathrm{IE}}| for top attention heads under translation and control prompts, averaged over 56 translation directions. The leading heads have much larger effects in the translation setting than in the controls, indicating that these heads are selective for matching translations and do not directly increase the probability of the given target sentence.

We first search for model components that are causally relevant in promoting the valid translation over an invalid translation. We do so using generative causal mediation (GCM; Sankaranarayanan et al., 2026). Intuitively, GCM measures how intervening on the activation of a model component (e.g., an attention head) influences the model’s relative preference for one continuation over another.

We measure the model’s preference for a single gold translation by swapping the source sentence that the model reads. Let rr be a gold translation of a source sentence, held fixed throughout. Let porigp_{\text{orig}} be a prompt whose final source sentence is the one that rr actually translates, and let pcfp_{\text{cf}} be the same prompt with a distinct final source sentence. We define the preference metric MM:

M=log⁡π⁡(r∣porig)−log⁡π⁡(r∣pcf),M=\log\pi(r\mid p_{\text{orig}})-\log\pi(r\mid p_{\text{cf}}), (3)

where log⁡π⁡(r∣p)=∑tlog⁡p⁡(rt∣p,r<t)\log\pi(r\mid p)=\sum_{t}\log p(r_{t}\mid p,r_{<t}) is the teacher-forced log-probability of a continuation, summed over its tokens. MM is the degree to which the model prefers to generate the gold translation rr when it has read the matching source, rather than a mismatched one; it isolates how much of the production of rr depends on having read the source content.

We wish to know how strongly each attention head influences this preference. To measure this, we take the activation of an attention head; we call this zz, and let zorigz_{\text{orig}} and zcfz_{\text{cf}} be the values it takes when the model reads porigp_{\text{orig}} and pcfp_{\text{cf}} respectively. Let IE​(z)\text{IE}(z) be the causal contribution of zz to MM, measured as the change in MM when we move the head from its mismatched-source state zcfz_{\text{cf}} to its matched-source state zorigz_{\text{orig}} while holding the rest of the computation fixed. We source the counterfactual activation zcfz_{\text{cf}} by running a forward pass on pcfp_{\text{cf}} and caching zz at the final source-token position, then patch it into the corresponding position of the run we score:

IE⁡(z)=M⁡(zorig)−M⁡(zcf).\mathrm{IE}(z)=M(z_{\text{orig}})-M(z_{\text{cf}}). (4)

Patching a single head and rerunning the model tells us that head’s exact contribution, but computing IE​(z)\text{IE}(z) this way for every zz is intractable: it would require O⁡(Z⋅n)O(Z\cdot n) forward passes, where ZZ is the number of mediators and nn the number of examples. We instead use attribution patching (Syed et al., 2023), a first-order linear approximation of the IE based on gradient attributions (Simonyan et al., 2013):

IE^​(z)=∇zM|z=zorig⋅(zorig−zcf).\widehat{\mathrm{IE}}(z)=\nabla_{z}M\big|_{z=z_{\text{orig}}}\cdot(z_{\text{orig}}-z_{\text{cf}}). (5)

The gradient ∇zM\nabla_{z}M for each component can be computed in a single backward pass, and the deltas (zorig−zcf)(z_{\text{orig}}-z_{\text{cf}}) for each component can be computed in two forward passes by caching each component’s activation on the two source prompts.

The magnitude of the indirect effect IE^\widehat{\mathrm{IE}} is how much of a model’s output behavior (as quantified by MM) flows through a component when all else is kept constant. Because MM rewards producing rr under the correct source, components with a positive indirect effect are those that cause the probability of the correct translation rorigr_{\text{orig}} to increase relative to the counterfactual translation rcfr_{\text{cf}}, and those with a negative indirect effect cause the probability of the correct translation to decrease relative to the counterfactual translation.

We instantiate rr as a gold translation from FLORES (Goyal et al., 2022), a massively parallel dataset that supports 101 languages. Given a source and target language, we construct a 2-shot prompt as follows:

{Source}: {shot 1 source}
{Target}: {shot 1 target}
{Source}: {shot 2 source}
{Target}: {shot 2 target}
{Source}: {query source}
{Target}:

and patch at the last source-token position (the final ": "). For each of the 56 ordered pairs where the source and target can be any language in {English, Spanish, German, French, Turkish, Arabic, Hindi, Hebrew} (excluding pairs where the source and target are the same), we compute the IE^\widehat{\mathrm{IE}} for all attention heads over n=100n=100 uniformly sampled pairs.

Naïvely, we may find heads that respond at least in part to variations in the source samples that are unrelated to the translation task. To verify that the components we find are selective for correct source–target pairings, we compare indirect effects with three controls. First, the same-language control refers to cases where the source language is the same as the target, such that the correct “translation” is the same sentence copied from the source, and the counterfactual translation is a randomly sampled sentence in the same language as the source. Higher indirect effects for the translation setting than the same-language control indicate that the component specifically causes correct translations to be more probable, and not copies of the source sequence. The null cross-language control refers to a setup similar to the machine translation setup, but where neither the original nor counterfactual completion are the correct translation. We expect indirect effects here to be very small relative to the translation task. Finally, we have the null same-language control, where both the original and counterfactual completions are in the same language as the source, but neither are the same as the source sentence.

The heads we find have larger effects in the translation setting by far than in any control setting (Figure 2). The mean |IE^||\widehat{\mathrm{IE}}| in the translation task exceeds that of the null cross-language control by 3.0×3.0\times over all heads, and 5.2×5.2\times over the top-10 heads for Llama (and 2.5×2.5\times for all heads/4.3×4.3\times for the top-10 heads for Aya). We observe a similar magnitude of increase relative to both same-language controls. This suggests that the heads we have found are responsible for performing machine translation, and that their effects cannot be explained by their general utility in generating fluent target sequences regardless of the source.

The translation-specific heads are localized to layers 13–14 in Llama and 15–20 in Aya. With respect to the stages of inference hypothesis (Lad et al., 2024), this would correspond to the layers that refine concept representations into increasingly abstract representations.

4.2 Translation Heads Have Similar Effects Across Language Pairs

An interlingua should have multilingual feature representations; this would allow a model to reuse the same representations when reading a source sequence and generating the target sequence. Prior work has established the existence of massively multilingual representations (Wendler et al., 2024), including grammatical concept representations (Brinkmann et al., 2025); here, we confirm these findings for the model we study by investigating their indirect effects across language pairs.

We follow the procedure of §4.1 to compute IE^\widehat{\mathrm{IE}} for each language pair in the translation test dataset. We show the indirect effects for the top heads by absolute IE^\widehat{\mathrm{IE}} across language pairs; if there is feature reuse across pairs, then the sign of the effect should be the same for many language pairs. Note that the magnitude of IE^\widehat{\mathrm{IE}} is not directly comparable across languages, as the initial probability of the target sequence and its translation capabilities are language pair–dependent.

Refer to caption
Figure 3: Signed mean IE for Llama-3.1-8B on machine translation for selected language pairs. Each column is 1 of the 15 top attention heads by IE^\widehat{\mathrm{IE}} across all languages. The top axis shows for how many language pairs the head was in the top-20 set. Several heads have stable effect signs across language pairs, which suggests that the same heads are being reused across many language pairs.

The top attention heads by IE^\widehat{\mathrm{IE}} are largely shared across language pairs. Figure 3 shows the 15 heads that appear most often in a single direction’s top-20 by |IE^||\widehat{\mathrm{IE}}|: the most universal of these are in the top-20 for all 56 directions. The sign of their effect is also stable—a given head keeps the same sign (favoring either correct or incorrect translations) across virtually every direction, regardless of source and target languages (For the full results from both models, see Figures  10 and  11). As predicted by our hypothesis, many of the translation mechanisms employed by the model do not depend on the choice of language pair; in fact, most of the top heads have the same directionality and general magnitude of effect on model performance for all language pairs.

4.3 Ablating Translation Heads Degrades Performance on Translation and Monolingual Tasks

Another mechanistic prediction of our hypothesis is that the heads most responsible for performing machine translation should reuse computational machinery used for monolingual tasks. To test this, we first verify that ablating the top heads by indirect effect harms machine translation performance. Then, we show that ablating the same heads also harms performance on acceptability judgments and multiple-choice question answering, suggesting that these heads are not selective for translation; rather, they may be reusing computational machinery from more general language modeling mechanisms.

Refer to caption
Figure 4: Ablation experiments overview. Ablating the top heads by causal influence on MT (POS-10, in red) performance drives a significant decrease in BLEU, and more than ablating random heads (Control, in grey). For GlobalMMLU, ablating those top heads reduces the model’s ability to distinguish between the correct and incorrect answer more than ablating random heads does. The Control involves ablating the same number of heads in the same layers.

For each target language TT, we aggregate signed IE over the 7 language pairs in which TT is the target, then select the top 10 positive-IE heads (POS-10, the correct translation–favoring set) by signed magnitude. We mean-ablate a head by replacing its output with its average activation over the translation prompts at all token positions. As a control, we ablate the same number of randomly sampled heads from the same layers. We then regenerate FLORES translations and measure the change in BLEU relative to the original model before ablations.

Ablating these heads causes significant reductions in BLEU. Ablating POS-10 degrades BLEU in all 8 target languages for both models (Figure 4, bottom) by roughly twice the amount as the random control. This suggests that the heads we have found are causally relevant to translation performance.

Llama-3.1-8B TinyAya-3B
Direction Base Mono Parallel Base Mono Parallel
Xho →\rightarrow Eng 16.84 24.81 24.88 23.46 26.64 27.80
Fra →\rightarrow Eng 43.27 42.28 38.87 40.70 40.26 38.11
Deu →\rightarrow Eng 42.69 42.76 39.80 40.20 41.21 39.58
Table 3: Monolingual fine-tuning recovers most of the Xhosa translation gains obtained with parallel fine-tuning. Both conditions also retain most performance on French and German translation (despite their not appearing in the fine-tuning data), although monolingual fine-tuning retrains slightly more performance.

Having shown the importance of these heads for translation quality, we now ask whether these same heads are responsible for monolingual capabilities. For this, we again use the GlobalMMLU dataset. We uniformly sample 400 examples per language. Given prompt pp with correct token completion rcorrectr_{\text{correct}} and a randomly chosen incorrect completion rincorrectr_{\text{incorrect}}, each minimal pair is scored by the acceptability margin

Δ=log⁡p⁡(rcorrect∣p)−log⁡p⁡(rincorrect∣p)\Delta=\log p(r_{\text{correct}}\mid p)-\log p(r_{\text{incorrect}}\mid p) (6)

We report the change in Δ\Delta after ablating the same heads ablated in the translation experiments:

Δchange=Δablation−Δoriginal,\Delta_{\text{change}}=\Delta_{\text{ablation}}-\Delta_{\text{original}}, (7)

where a negative Δchange\Delta_{\text{change}} means the ablation weakened the model’s ability to distinguish the correct from the incorrect answer.

Ablating the POS-10 heads from the MT task results in a negative Δchange\Delta_{\text{change}}, whereas ablating random heads has a smaller effect (Figure 4, top). That the same heads affect performance in both tasks provides preliminary evidence that these heads implement computations that are reused across tasks. Thus, these heads appear to implement computations that are not selective for translation alone. This provides further support for our hypothesis.44 4 That said, applying the same ablations in MultiBLiMP results in a smaller effect; see Fig. 14 in App. B. Thus, these heads are not completely task-agnostic.

5 Monolingual Fine-tuning Recovers Most Gains from Parallel Data

If LLMs translate via task-agnostic internal representations, translation quality should depend in part on the model’s monolingual competence in the source and target languages. For example, if we wish to translate between a low-resource and high-resource language, then we should be able to improve translation performance given access only to data that improves language modeling quality on the low-resource language, assuming that we preserve capabilities in the high-resource language. Under this view, low translation performance may reflect a failure to encode the source sentence into an adequate internal representation, or to decode the target sentence fluently from it. Recent work has shown that bilingual or mixed-language signals during pretraining can be important for acquiring translation capabilities in LLMs (Briakou et al., 2023; Qorib et al., 2025; Shao et al., 2026). We ask the complementary question of whether, given a fully pretrained model, improvements in monolingual language modeling can (at least in part) transfer to translation even without parallel data.

We test this idea using Llama-3.1-8B and TinyAya-3B.55 5 We use TinyAya instead of Aya-23-8B because Aya-23-8B only provides an instruction-tuned model, and no base model. In pilot experiments, we found it difficult to achieve good performance after fine-tuning in any language (including high-resource languages) with instruction-tuned models. For each model, we compare how fine-tuning on an underrepresented language (Xhosa) affects translation performance when the training data consist of either parallel data or monolingual text. Specifically, in the monolingual setting, we train on unaligned text consisting of 80% Xhosa and 20% English, French, and German data to avoid catastrophic forgetting of the model’s existing language abilities. In the parallel setting, we use paired Xhosa and English sentences from OPUS MT560 (Gowda et al., 2021), presented in both translation directions.

Both settings use a next-token prediction loss and rank-16 LoRA adapters (Hu et al., 2022). The reported runs use a learning rate of 3×10−53\times 10^{-5}, a maximum sequence length of 512 tokens, and one epoch over 100 million tokens. We evaluate translation with few-shot prompting on the FLORES devtest split (Goyal et al., 2022), using examples from the dev split as in-context demonstrations. Thus, the models are prompted to translate at evaluation time even though the monolingual condition contains no translation examples during training.

Table 3 shows that monolingual fine-tuning recovers most of the improvement obtained with parallel data. For Llama, monolingual fine-tuning raises BLEU from 16.84 to 24.81, compared with 24.88 under parallel fine-tuning. This recovers 99% of the improvement over the base model. For Aya, monolingual fine-tuning raises BLEU from 23.46 to 26.64, compared with 27.80 under parallel fine-tuning, recovering 73% of the improvement. For both models, both fine-tuning conditions largely preserve translation performance from French and German into English, although monolingual fine-tuning remains closer to the base performance.

These results are consistent with the view that once an LLM has been pretrained and acquired multilingual representations, improving its ability to model an underrepresented language can produce most of the available translation gains. Parallel data may still perform better because it can improve language modeling while also strengthening translation-specific mechanisms. Our results therefore do not imply that parallel data are unimportant, particularly during pretraining. They show that, after pretraining, a large share of the gains from parallel fine-tuning can be achieved using only unaligned data.

We provide results for additional languages (German and Thai) and translation directions in Appendix C. In short, changes in performance are smaller for higher-resource languages, but parallel and monolingual fine-tuning still achieve largely comparable results.

6 Discussion and Conclusions

We have provided evidence that machine translation capabilities in LLMs are mediated in significant part by multilingual representations that are also relevant for some monolingual tasks. Across three complementary analyses, we find evidence consistent with this view: Translation performance is largely explained by source- and target-language competence, translation-relevant components overlap with components used for monolingual grammatical behavior, and monolingual Xhosa fine-tuning recovers most of the gains from parallel fine-tuning across both models. Together, these results provide support for the hypothesis that LLMs translate by mapping source-language input into a latent feature space (that can also in theory be used for monolingual next-token predictions), and then generating target-language text from that shared representation.

This interpretation helps explain why monolingual competence is strongly associated with translation performance. If a model translates through a partially language-agnostic feature space, then improving its ability to read or write a language should improve translation involving that language. The observation that target-language competence explains more variance than source-language competence may suggest that fluent and grammatical generation is often the bottleneck. Under this view, poor translation into a language need not imply the absence of a dedicated translation circuit for that language pair; instead, it may reflect weak target-language decoding capabilities given otherwise usable latent representations of the source sentence.

We emphasize that translation-specific mechanisms not based on interlingua are also likely to exist; indeed, Todd et al. (2024) find that there are components selective for word translation. We do not claim that an interlingua is the only means by which LLMs translate, but rather, that it is a significant mechanism that controls a large proportion of model performance on MT tasks. It is likely that LLMs use a mixture of mechanisms (Gur-Arieh et al., 2026) to achieve their machine translation capabilities.

Limitations

The interlingua hypothesis is one plausible explanation for what we have observed, but our experiments do not rule out the existence of other mechanisms. The finding that the target language yields more predictive power than the source language in predicting BLEU score may be an artifact of the nn-gram-based computation of BLEU scores. It is also possible that the features underlying high-quality translations are not well represented in the BLEU score.

Our experiments were limited to two 8B-parameter models and one 3B-parameter model, and may not generalize to smaller or larger LLMs. Future work should investigate whether similar trends hold for a wider variety of language models, and whether recent developments in thinking models affect these findings.

For our GCM experiments, we focus on components that mediate at the last token, missing computations that happen at earlier positions.

Finally, our translation experiments are limited to a relatively small number of language pairs. Future work could scale up this experiment to investigate whether these trends hold across a large number of language pairs, and to what degree monolingual fine-tuning can recover the performance of parallel fine-tuning across many language pairs (and what other factors explain when this works well and when it does not).

Acknowledgments

We are grateful to the members of the BAAIGL lab at Boston University for helpful comments on an earlier iteration of this work. The computational work reported on in this paper was performed largely on the Shared Computing Cluster, which is administered by Boston University’s Research Computing Services. Jannik Brinkmann is supported by the German Federal Ministry for Economic Affairs and the German Federal Ministry of Research, Technology and Space.

References

  • Aycock et al. (2025) S. Aycock, D. Stap, D. Wu, C. Monz, and K. Sima’an Can LLMs really learn to translate a low-resource language from one grammar book?. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §2.
  • Bahdanau et al. (2015) D. Bahdanau, K. Cho, and Y. Bengio Neural machine translation by jointly learning to align and translate. Cited by: §1.
  • Briakou et al. (2023) E. Briakou, C. Cherry, and G. Foster Searching for needles in a haystack: on the role of incidental bilingualism in PaLM’s translation capability. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Rogers, J. Boyd-Graber, and N. Okazaki (Eds.), Toronto, Canada, pp. 9432–9452. External Links: Link, Document Cited by: §5.
  • Brinkmann et al. (2025) J. Brinkmann, C. Wendler, C. Bartelt, and A. Mueller Large language models share representations of latent grammatical concepts across typologically diverse languages. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, pp. 6131–6150. External Links: Link, Document, ISBN 979-8-89176-189-6 Cited by: §1, §2, §4.2.
  • Cahyawijaya et al. (2024) S. Cahyawijaya, H. Lovenia, and P. Fung LLMs are few-shot in-context low-resource language learners. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), K. Duh, H. Gomez, and S. Bethard (Eds.), Mexico City, Mexico, pp. 405–433. External Links: Link, Document Cited by: §1.
  • Conneau et al. (2020) A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott, L. Zettlemoyer, and V. Stoyanov Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault (Eds.), Online, pp. 8440–8451. External Links: Link, Document Cited by: §1, §2.
  • Court and Elsner (2024) S. Court and M. Elsner Shortcomings of LLMs for low-resource translation: retrieval and understanding are both the problem. In Proceedings of the Ninth Conference on Machine Translation, B. Haddow, T. Kocmi, P. Koehn, and C. Monz (Eds.), Miami, Florida, USA, pp. 1332–1354. External Links: Link, Document Cited by: §2.
  • Devlin et al. (2019) J. Devlin, M. Chang, K. Lee, and K. Toutanova BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), J. Burstein, C. Doran, and T. Solorio (Eds.), Minneapolis, Minnesota, pp. 4171–4186. External Links: Link, Document Cited by: §1.
  • Dumas et al. (2025) C. Dumas, C. Wendler, V. Veselovsky, G. Monea, and R. West Separating tongue from thought: activation patching reveals language-agnostic concept representations in transformers. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 31822–31841. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §2.
  • Escolano et al. (2021) C. Escolano, M. R. Costa-jussà, J. A. R. Fonollosa, and M. Artetxe Multilingual machine translation: closing the gap between shared and language-specific encoder-decoders. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics, pp. 944–948. External Links: Document, Link Cited by: §2.
  • Fan et al. (2020) A. Fan, S. Bhosale, H. Schwenk, Z. Ma, A. El-Kishky, S. Goyal, M. Baines, O. Celebi, G. Wenzek, V. Chaudhary, N. Goyal, T. Birch, V. Liptchinsky, S. Edunov, E. Grave, M. Auli, and A. Joulin Beyond english-centric multilingual machine translation. External Links: 2010.11125, Document, Link Cited by: §2.
  • Finlayson et al. (2021) M. Finlayson, A. Mueller, S. Gehrmann, S. Shieber, T. Linzen, and Y. Belinkov Causal analysis of syntactic agreement mechanisms in neural language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), C. Zong, F. Xia, W. Li, and R. Navigli (Eds.), Online, pp. 1828–1843. External Links: Link, Document Cited by: §2.
  • Fiotto-Kaufman et al. (2025) J. F. Fiotto-Kaufman, A. R. Loftus, E. Todd, J. Brinkmann, K. Pal, D. Troitskii, M. Ripa, A. Belfki, C. Rager, C. Juang, A. Mueller, S. Marks, A. S. Sharma, F. Lucchetti, N. Prakash, C. E. Brodley, A. Guha, J. Bell, B. C. Wallace, and D. Bau NNsight and NDIF: democratizing access to open-weight foundation model internals. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: Appendix E.
  • Geiger et al. (2021) A. Geiger, H. Lu, T. Icard, and C. Potts Causal abstractions of neural networks. In Advances in Neural Information Processing Systems, M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. W. Vaughan (Eds.), Vol. 34, pp. 9574–9586. External Links: Link Cited by: §2.
  • Gowda et al. (2021) T. Gowda, Z. Zhang, C. Mattmann, and J. May Many-to-english machine translation tools, data, and pretrained models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, pp. 306–316. External Links: Link, Document Cited by: Appendix D, §5.
  • Goyal et al. (2022) N. Goyal, C. Gao, V. Chaudhary, P. Chen, G. Wenzek, D. Ju, S. Krishnan, M. Ranzato, F. Guzmán, and A. Fan The Flores-101 evaluation benchmark for low-resource and multilingual machine translation. Transactions of the Association for Computational Linguistics 10, pp. 522–538. External Links: Link, Document Cited by: Appendix D, §4.1, §5.
  • Gur-Arieh et al. (2026) Y. Gur-Arieh, M. Geva, and A. Geiger Mixing mechanisms: how language models retrieve bound entities in-context. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §6.
  • Heimersheim and Nanda (2024) S. Heimersheim and N. Nanda How to use and interpret activation patching. External Links: 2404.15255, Link Cited by: §2.
  • Hendrycks et al. (2021) D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt Measuring massive multitask language understanding. In International Conference on Learning Representations, External Links: Link Cited by: §3.
  • Hendy et al. (2023) A. Hendy, M. Abdelrehim, A. Sharaf, V. Raunak, M. Gabr, H. Matsushita, Y. J. Kim, M. Afify, and H. H. Awadalla How good are gpt models at machine translation? a comprehensive evaluation. External Links: 2302.09210, Link Cited by: §1.
  • Hu et al. (2022) E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen LoRA: low-rank adaptation of large language models. In International Conference on Learning Representations, External Links: Link Cited by: §5.
  • Jumelet et al. (2026) J. Jumelet, L. Weissweiler, J. Nivre, and A. Bisazza MultiBLiMP 1.0: a massively multilingual benchmark of linguistic minimal pairs. Transactions of the Association for Computational Linguistics 14, pp. 193–216. External Links: Link, Document Cited by: Appendix D, §3.
  • Koehn and Knowles (2017) P. Koehn and R. Knowles Six challenges for neural machine translation. In Proceedings of the First Workshop on Neural Machine Translation, T. Luong, A. Birch, G. Neubig, and A. Finch (Eds.), Vancouver, pp. 28–39. External Links: Link, Document Cited by: §2.
  • Lad et al. (2024) V. Lad, W. Gurnee, and M. Tegmark The remarkable robustness of LLMs: stages of inference?. In ICML 2024 Workshop on Mechanistic Interpretability, External Links: Link Cited by: §1, §4.1, §4.
  • Lewis (1973) D. Lewis Causation. The journal of philosophy 70 (17), pp. 556–567. Cited by: §2.
  • Lu et al. (2018) Y. Lu, P. Keung, F. Ladhak, V. Bhardwaj, S. Zhang, and J. Sun A neural interlingua for multilingual machine translation. In Proceedings of the Third Conference on Machine Translation: Research Papers, pp. 84–92. External Links: Document, Link Cited by: §2.
  • Luong et al. (2015) T. Luong, H. Pham, and C. D. Manning Effective approaches to attention-based neural machine translation. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, L. Màrquez, C. Callison-Burch, and J. Su (Eds.), Lisbon, Portugal, pp. 1412–1421. External Links: Link, Document Cited by: §1.
  • Mao et al. (2023) Z. Mao, H. Song, R. Dabre, C. Chu, and S. Kurohashi Variable-length neural interlingua representations for zero-shot neural machine translation. In Proceedings of the 1st International Workshop on Multilingual, Multimodal and Multitask Language Generation, A. Barreiro, M. Silberztein, E. Lloret, and M. Paprzycki (Eds.), Tampere, Finland, pp. 16–25. External Links: Link Cited by: §2.
  • Marks et al. (2025) S. Marks, C. Rager, E. J. Michaud, Y. Belinkov, D. Bau, and A. Mueller Sparse feature circuits: discovering and editing interpretable causal graphs in language models. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §2.
  • Mueller et al. (2026) A. Mueller, J. Brinkmann, M. Li, S. Marks, K. Pal, N. Prakash, C. Rager, A. Sankaranarayanan, A. S. Sharma, J. Sun, E. Todd, D. Bau, and Y. Belinkov The quest for the right mediator: surveying mechanistic interpretability for nlp through the lens of causal mediation analysis. Computational Linguistics 52 (1), pp. 331–378. External Links: ISSN 0891-2017, Document, Link, https://direct.mit.edu/coli/article-pdf/52/1/331/2554934/coli.a.572.pdf Cited by: §2.
  • Pearl (2001) J. Pearl Direct and indirect effects. In Proceedings of the Seventeenth Conference on Uncertainty in Artificial Intelligence, UAI’01, San Francisco, CA, USA, pp. 411–420. External Links: ISBN 1558608001 Cited by: §2, §4.
  • Post (2018) M. Post A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, O. Bojar, R. Chatterjee, C. Federmann, M. Fishel, Y. Graham, B. Haddow, M. Huck, A. J. Yepes, P. Koehn, C. Monz, M. Negri, A. Névéol, M. Neves, M. Post, L. Specia, M. Turchi, and K. Verspoor (Eds.), Brussels, Belgium, pp. 186–191. External Links: Link, Document Cited by: Appendix D, §3.
  • Purason and Tättar (2022) T. Purason and A. Tättar Multilingual neural machine translation with the right amount of sharing. In Proceedings of the 23rd Annual Conference of the European Association for Machine Translation, pp. 91–100. External Links: Link Cited by: §2.
  • Qorib et al. (2025) M. R. Qorib, J. Li, and H. T. Ng Just go parallel: improving the multilingual capabilities of large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 33411–33424. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §5.
  • Radford et al. (2018) A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al. Improving language understanding by generative pre-training. Cited by: §1.
  • Robinson et al. (2023) N. R. Robinson, P. Ogayo, D. R. Mortensen, and G. Neubig ChatGPT mt: competitive for high- (but not low-) resource languages. External Links: 2309.07423, Link Cited by: §1.
  • Sankaranarayanan et al. (2026) A. Sankaranarayanan, A. Zur, A. Geiger, and D. Hadfield-Menell Activation steering via generative causal mediation. External Links: 2602.16080, Link Cited by: §4.1.
  • Shao et al. (2026) J. Shao, R. Tang, C. Zhang, K. Sevegnani, P. Stenetorp, J. Yang, and Y. Lu The role of mixed-language documents for multilingual large language model pretraining. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 36807–36818. External Links: Link, Document, ISBN 979-8-89176-390-6 Cited by: §5.
  • Simonyan et al. (2013) K. Simonyan, A. Vedaldi, and A. Zisserman Deep inside convolutional networks: visualising image classification models and saliency maps. arXiv preprint arXiv:1312.6034. Cited by: §4.1.
  • Singh et al. (2025) S. Singh, A. Romanou, C. Fourrier, D. I. Adelani, J. G. Ngui, D. Vila-Suero, P. Limkonchotiwat, K. Marchisio, W. Q. Leong, Y. Susanto, R. Ng, S. Longpre, S. Ruder, W. Ko, A. Bosselut, A. Oh, A. Martins, L. Choshen, D. Ippolito, E. Ferrante, M. Fadaee, B. Ermis, and S. Hooker Global MMLU: understanding and addressing cultural and linguistic biases in multilingual evaluation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 18761–18799. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: Appendix D, §3.
  • Syed et al. (2023) A. Syed, C. Rager, and A. Conmy Attribution patching outperforms automated circuit discovery. External Links: 2310.10348, Link Cited by: §4.1.
  • Tanzer et al. (2024) G. Tanzer, M. Suzgun, E. Visser, D. Jurafsky, and L. Melas-Kyriazi A benchmark for learning to translate a new language from one grammar book. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: §1, §2.
  • Todd et al. (2024) E. Todd, M. L. Li, A. S. Sharma, A. Mueller, B. C. Wallace, and D. Bau Function vectors in large language models. In The Twelfth International Conference on Learning Representations, Note: arXiv:2310.15213 External Links: Link Cited by: §6.
  • Vaswani et al. (2017) A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin Attention is all you need. Advances in neural information processing systems 30. Cited by: §1.
  • Vázquez et al. (2020) R. Vázquez, A. Raganato, M. Creutz, and J. Tiedemann A systematic study of inner-attention-based sentence representations in multilingual neural machine translation. Computational Linguistics 46 (2), pp. 387–424. External Links: Link, Document Cited by: §2.
  • Vázquez et al. (2019) R. Vázquez, A. Raganato, J. Tiedemann, and M. Creutz Multilingual NMT with a language-independent attention bridge. In Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019), I. Augenstein, S. Gella, S. Ruder, K. Kann, B. Can, J. Welbl, A. Conneau, X. Ren, and M. Rei (Eds.), Florence, Italy, pp. 33–39. External Links: Link, Document Cited by: §2.
  • Vig et al. (2020) J. Vig, S. Gehrmann, Y. Belinkov, S. Qian, D. Nevo, Y. Singer, and S. Shieber Investigating gender bias in language models using causal mediation analysis. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33, pp. 12388–12401. External Links: Link Cited by: §2, §4.
  • Warstadt et al. (2020) A. Warstadt, A. Parrish, H. Liu, A. Mohananey, W. Peng, S. Wang, and S. R. Bowman BLiMP: the benchmark of linguistic minimal pairs for English. Transactions of the Association for Computational Linguistics 8, pp. 377–392. External Links: Link, Document Cited by: §3.
  • Wendler et al. (2024) C. Wendler, V. Veselovsky, G. Monea, and R. West Do llamas work in English? on the latent language of multilingual transformers. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 15366–15394. External Links: Link, Document Cited by: §1, §2, §4.2.
  • Wu et al. (2025) Z. Wu, X. V. Yu, D. Yogatama, J. Lu, and Y. Kim The semantic hub hypothesis: language models share semantic representations across languages and modalities. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §1.
  • Zhu et al. (2020) C. Zhu, H. Yu, S. Cheng, and W. Luo Language-aware interlingua for multilingual neural machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault (Eds.), Online, pp. 1650–1655. External Links: Link, Document Cited by: §2.

Appendix A Additional Results for Modeling Translation Performance

Figure 5 shows predicted vs. actual BLEU scores for the linear and bilinear models. In general, linear and bilinear models produce visually indistinguishable predictions.

Refer to caption
Refer to caption
Figure 5: Predicted (xx) vs. actual (yy) BLEU on the common 18-language subset, for each proxy (columns) under the linear (top row) and bilinear (bottom row) models, Llama (top) and Aya (bottom). The two rows are visually indistinguishable for every proxy—the graphical form of the null interaction in Table 2.

A.1 Language-pair interaction tests

Table 4 reports the nested-model FF-test on the lS⋅lTl_{S}\!\cdot\!l_{T} interaction coefficient, on the common 18-language subset and the 17-language no-English subset. The interaction is non-significant in every cell.

Subset Proxy (model) F⁡(1,⋅)F(1,\cdot) pp Δ​R2\Delta R^{2}
18-lang MMLU, Aya 0.001 0.980 <10−5<10^{-5}
18-lang MMLU, Llama 1.74 0.188 0.0015
18-lang MultiBLiMP acc., Aya 0.262 0.609 0.0007
18-lang MultiBLiMP acc., Llama 0.000 0.990 <10−5<10^{-5}
18-lang MultiBLiMP margin, Aya 0.400 0.528 0.0010
18-lang MultiBLiMP margin, Llama 0.028 0.867 0.0001
18-lang Perplexity, Aya 0.439 0.508 0.0015
18-lang Perplexity, Llama 0.009 0.924 <10−4<10^{-4}
17-lang MMLU, Aya 0.014 0.906 <10−4<10^{-4}
17-lang MMLU, Llama 0.955 0.329 0.0013
17-lang MultiBLiMP acc., Aya 0.653 0.420 0.0019
17-lang MultiBLiMP acc., Llama 0.001 0.978 <10−5<10^{-5}
17-lang MultiBLiMP margin, Aya 1.16 0.283 0.0038
17-lang MultiBLiMP margin, Llama 0.414 0.520 0.0014
17-lang Perplexity, Aya 0.497 0.482 0.0018
17-lang Perplexity, Llama 0.170 0.681 0.0006
Table 4: Nested FF-test for the language-pair interaction term. Non-significant everywhere; largest Δ​R2=0.0038\Delta R^{2}=0.0038.

A.2 Predicting BLEU from Language-specific Translation Capabilities

Here, we show BLEU scores for all language pairs (Figure 6, left). We take each source or target language’s average BLEU across language pairs, and fit linear models based on these terms to predict each language pair’s BLEU score (see §3 for details). We observe that the error of a rank-1 linear predictor (Figure 6, right) is generally low at around 10–15%. This provides further evidence that BLEU scores are predictable as a function of language-specific capabilities.

Refer to caption
Figure 6: Left: the observed 24×2424\times 24 Llama BLEU matrix (self-translation diagonal masked, never imputed). Right: its rank-1 reconstruction uS​vTu_{S}v_{T}, fit by masked alternating least squares over the 552 observed off-diagonal cells. Faithfulness =89.8%=89.8\% (Aya: 83.4%83.4\%); a single source-competence vector outer-producted with a single target-competence vector reconstructs most of the matrix.

A.3 Where the grammatical signal concentrates

Restricting the MultiBLiMP margin to subject–verb agreement phenomena improves BLEU prediction (Llama, 18 languages: all-phenomena margin R2=0.324→0.363R^{2}=0.324\to 0.363 for the SV-agreement subset). Per phenomenon, SV-Person reaches R2=0.688R^{2}=0.688, SV-Gender 0.5190.519, SV-Number 0.4120.412, versus subject–predicate agreement at 0.1510.151 / 0.0370.037. The production-side signal (agreement on the generated verb) is where translation predictiveness lives. We caution that per-phenomenon coverage is confounded with which languages each phenomenon is annotated for, so we report this as a strengthening analysis rather than a headline.

A.4 MMLU subject subsets

Unlike the MultiBLiMP phenomenon breakdown, the BLEU-predictive signal in MMLU is not concentrated in any single subject category: aggregate MMLU (R2=0.598R^{2}=0.598 Llama / 0.4990.499 Aya, 23 languages) is stronger than every individual category (best single subset: Humanities 0.5600.560 for Llama, Social Sciences 0.4130.413 for Aya; STEM 0.3600.360/0.1900.190). As with phenomena, subject-category accuracies are highly collinear across languages and largely track per-language resource level.

A.5 Functional form

BLEU is bounded and right-skewed, and grammatical accuracy saturates near 1. Replacing the raw fit with log⁡BLEU∼log⁡(1−aS)+log⁡(1−aT)\log\text{BLEU}\sim\log(1-a_{S})+\log(1-a_{T}) (a logit-like transform that un-saturates accuracy) improves R2R^{2} from 0.2500.250 to 0.3390.339 (Llama) and 0.3940.394 to 0.4440.444 (Aya); the analogous log–log fit for perplexity improves 0.071→0.1270.071\to 0.127 (Llama). We read these gains as variance-stabilization addressing the same ceiling effect that motivates the log-probability margin, rather than evidence of a specific (e.g. exponential) functional form, and so report the raw-scale linear models in the main text.

A.6 Does grammatical competence add signal beyond Global MMLU?

Adding MultiBLiMP source and target terms on top of the Global MMLU-only regression gives a small but significant gain (Table 5). Grammatical competence carries a little translation-relevant signal that general task accuracy doesn’t capture.

Model MMLU-only R2R^{2} Full R2R^{2} Partial R2R^{2} FF-test pp
Llama 0.739 0.752 0.048 5.8×10−45.8\times 10^{-4}
Aya 0.510 0.556 0.094 3.7×10−73.7\times 10^{-7}
Table 5: Does grammatical competence add predictive value beyond general competence? We compare BLEU∼MMLUS+MMLUT\text{BLEU}\sim\text{MMLU}_{S}+\text{MMLU}_{T} to the same model alongside MultiBLiMP source and target terms. Adding MultiBLiMP accuracy adds some small signal for Llama and Aya.

Appendix B Additional GCM Ablation Results

B.1 Translation-specific head decompositions

Figure 7 shows the attention head decompositions as in Fig. 2 in the main text. The top heads concentrate in slightly later layers than Llama.

Refer to caption
Figure 7: Prompt-swap GCM head IE under the four control tasks, top 20 heads sorted by real_cross−null_cross\textsc{real\_cross}-\textsc{null\_cross}, for Aya-23-8B (cf. Fig. 2). The top heads concentrate in layers 15–20. Support and bar definitions as in Fig. 2.

B.2 Top translation-specific heads

Tables 6 and 7 show the details of the top 10 attention heads by translation-specific effect. A majority of the effect is concentrated in the top heads.

mean |IE^||\widehat{\mathrm{IE}}|
Layer Head real_cr null_cr real_sm null_sm Δ\Delta #dir
13 18 1.076 0.204 0.137 0.045 0.871 56
13 0 0.843 0.162 0.102 0.035 0.681 56
14 29 0.751 0.130 0.084 0.017 0.621 56
14 31 0.685 0.126 0.091 0.023 0.558 56
13 13 0.690 0.140 0.064 0.024 0.550 56
13 4 0.575 0.121 0.095 0.029 0.454 56
13 17 0.549 0.110 0.056 0.020 0.438 56
14 13 0.513 0.090 0.104 0.027 0.423 56
14 30 0.519 0.098 0.044 0.012 0.421 56
13 16 0.499 0.099 0.116 0.039 0.400 55
Table 6: Top 10 attention heads by translation-specific effect Δ=real_cross−null_cross\Delta=\textsc{real\_cross}-\textsc{null\_cross} under prompt-swap GCM (Llama), with mean |IE^||\widehat{\mathrm{IE}}| in each of the four control tasks. #dir is the number of the 56 cross-language directions in which the head ranks in that direction’s top-30 by real_cross effect; all but one head appear in every direction.
mean |IE^||\widehat{\mathrm{IE}}|
Layer Head real_cr null_cr real_sm null_sm Δ\Delta #dir
15 19 0.867 0.224 0.179 0.057 0.644 56
18 25 0.436 0.089 0.072 0.020 0.347 56
16 1 0.417 0.090 0.044 0.013 0.327 56
17 18 0.379 0.074 0.067 0.014 0.305 56
16 10 0.375 0.093 0.111 0.042 0.282 56
17 6 0.318 0.069 0.038 0.011 0.248 55
20 29 0.319 0.076 0.054 0.012 0.243 56
18 24 0.305 0.063 0.058 0.016 0.242 55
16 29 0.317 0.085 0.065 0.018 0.232 56
20 20 0.311 0.082 0.105 0.014 0.229 50
Table 7: Top 10 Aya attention heads by translation-specific effect Δ=real_cross−null_cross\Delta=\textsc{real\_cross}-\textsc{null\_cross} under prompt-swap GCM (cf. Table 6). #dir is the number of the 56 cross-language directions in which the head ranks top-30 by real_cross effect.

B.3 SAE-feature decomposition analyses

Top features show the same real_cross-dominant pattern as attention heads but with weaker separation. A few top-ranked features are same-language features (Fig. 8 and 9).

Refer to caption
Figure 8: Llama-3.1-8B mean |IE^||\widehat{\mathrm{IE}}| for top SAE features under translation and control prompts.
Refer to caption
Figure 9: Aya-23-8B mean |IE^||\widehat{\mathrm{IE}}| for top SAE features under translation and control prompts.

B.4 Universality of translation-relevant heads

Translation-relevant heads are universal across translation language pairs; Hebrew and Hindi show more divergence from the rest of the languages (Fig. 10 and 11).

Refer to caption
Figure 10: Full Llama-3.1-8B signed-IE heatmap for the 15 most universal prompt-swap heads across all 56 translation directions.
Refer to caption
Figure 11: Full Aya-23-8B signed-IE heatmap for the 15 most universal prompt-swap heads across all 56 translation directions.

B.5 Head-IE concentration and sparsity

As further illustration of the head-IE sparsity, Fig. 12 and 13 show that the ratio of the largest mean |IE| to the median ranges around 2 orders of magnitude.

Refer to caption
Figure 12: Per-direction head-IE sparsity under prompt-swap GCM: the ratio of the largest head’s mean |IE^||\widehat{\mathrm{IE}}| to the median head’s, for each source→\totarget direction (Llama-3.1-8B; diagonal omitted). The ratio ranges from 60 to 162 (median 116). The top head is roughly two orders of magnitude above the median head in every direction, and is largest for directions into Spanish, German, and French and smallest into Hindi and Hebrew.
Refer to caption
Figure 13: Per-direction head-IE sparsity for Aya-23-8B under prompt-swap GCM (cf. Fig. 12). The top-head/median-head ratio ranges from 51 to 104 (median 83).

The sign of the translation task selection carries over to Multi-BLiMP: ablating the POS-10 heads lowers the grammaticality margin in all 8 target languages for both models, and ablating the NEG-10 heads raises it in all 8 (Fig. 14; Table 8). The magnitude of the POS-10 effect, however, is comparable to that of ablating random heads from the same layers, so it is the consistent direction of the effects, rather than their size, that suggests these heads are reused across tasks.

Refer to caption
Figure 14: Change in the Multi-BLiMP grammaticality margin from mean-ablating the POS-10 and NEG-10 head sets and a random control. Dotted lines mark cross-language means.
Llama-3.1-8B Aya-23-8B
target base POS-10 ctrlP NEG-10 ctrlN base POS-10 ctrlP NEG-10 ctrlN
ara 7.44 −0.055-0.055 −0.066-0.066 +0.077+0.077 +0.050+0.050 6.92 −0.021-0.021 −0.026-0.026 +0.035+0.035 +0.009+0.009
deu 10.18 −0.111-0.111 −0.019-0.019 +0.263+0.263 −0.076-0.076 7.88 −0.037-0.037 −0.026-0.026 +0.089+0.089 −0.023-0.023
eng 5.23 −0.025-0.025 −0.095-0.095 +0.050+0.050 −0.002-0.002 4.09 −0.009-0.009 −0.006-0.006 +0.023+0.023 0.0000.000
fra 9.16 −0.072-0.072 −0.087-0.087 +0.172+0.172 +0.051+0.051 6.33 −0.008-0.008 +0.026+0.026 +0.036+0.036 0.0000.000
heb 11.49 −0.075-0.075 +0.018+0.018 +0.113+0.113 −0.108-0.108 7.75 −0.015-0.015 −0.050-0.050 +0.049+0.049 −0.013-0.013
hin 6.47 −0.155-0.155 −0.188-0.188 +0.232+0.232 +0.012+0.012 7.61 −0.094-0.094 −0.162-0.162 +0.175+0.175 −0.077-0.077
spa 9.61 −0.105-0.105 −0.093-0.093 +0.224+0.224 −0.069-0.069 6.58 −0.017-0.017 −0.031-0.031 +0.056+0.056 −0.015-0.015
tur 9.56 −0.138-0.138 −0.027-0.027 +0.304+0.304 +0.093+0.093 9.45 −0.005-0.005 −0.085-0.085 +0.109+0.109 −0.071-0.071
mean 8.64 −0.092-0.092 −0.069-0.069 +0.179+0.179 −0.006-0.006 7.08 −0.026-0.026 −0.045-0.045 +0.072+0.072 −0.024-0.024
Table 8: Multi-BLiMP margin Δchange\Delta_{\text{change}} relative to baseline from mean-ablating the POS-10 and NEG-10 GCM head sets and their size-matched controls (ctrlP and ctrlN: 10 random heads drawn from the same layers as the corresponding set). Negative means ablation weakened the grammaticality preference. Ablating POS-10 lowers the margin in all 8 targets and ablating NEG-10 raises it in all 8, for both models; the size of the POS-10 effect, however, is comparable to that of its random same-layer control. n=400n=400 pairs per language except heb (n=200n=200) and hin (n=100n=100).

The choice of the number of heads to ablate is unimportant to the overall effect; ablating different numbers of top heads has monotonic effect on performance (Fig. 15).

Refer to caption
Figure 15: Head-count dose response (deu/eng/fra) on Llama. The effect is monotonic with respect to the number of heads ablated.

Appendix C Additional Fine-Tuning Results

This section reports the remaining fine-tuning results. We use the same models, training conditions, and evaluation procedure described in the main text.

C.1 Additional directions involving Xhosa

Table 9 reports translation into Xhosa from English, French, and German. Both fine-tuning conditions improve performance across all three directions and both models. For Llama, parallel fine-tuning performs best when translating from English and French, while monolingual fine-tuning performs best when translating from German. For Aya, monolingual fine-tuning performs slightly better in all three directions. Thus, neither condition performs best in every setting.

Model Method Eng. Fra. Deu.
Llama Base 1.39 0.57 0.91
Llama Monolingual 2.49 1.24 1.95
Llama Parallel 4.25 1.60 1.85
Aya Base 2.47 1.42 0.96
Aya Monolingual 4.00 2.00 2.13
Aya Parallel 3.79 1.97 1.95
Table 9: BLEU for translation into Xhosa. Column labels identify the source language. Bold marks the best fine-tuned result for each model and direction.

C.2 Fine-tuning on higher-resource languages

We also apply the same general pipeline to German and Thai using Llama-3.1-8B. Unlike Xhosa, both languages already have strong translation performance in the base model. As shown in Table 10, neither monolingual nor parallel fine-tuning produces a consistent improvement. Monolingual fine-tuning leaves performance largely unchanged, while parallel fine-tuning reduces performance in several directions.

One possible explanation is that German and Thai were already well represented during pretraining, leaving less room for improvement from continued training. Some of the fine-tuning documents may also have appeared in the pretraining corpus.

Adapted language Other directions
Method Into Eng. From Eng. Fra. into Eng. Xho. into Eng.
German adaptation
Base 42.69 34.94 43.27 16.84
Monolingual 43.48 34.55 43.16 19.13
Parallel 42.95 33.26 42.53 14.90
Thai adaptation
Base 32.60 15.21 43.27 16.84
Monolingual 32.19 15.19 43.34 16.75
Parallel 30.98 10.31 42.46 14.95
Table 10: BLEU after fine-tuning Llama-3.1-8B on German or Thai. The first two result columns involve the adapted language; the final two measure other translation directions.

C.3 Preservation of monolingual abilities

We also measure MultiBLiMP accuracy after fine-tuning Llama on Xhosa. As shown in Table 11, performance is largely preserved. However, the base model is already near the maximum score on all three languages. These results are therefore not strong enough to determine whether fine-tuning causes cross-lingual transfer of linguistic abilities.

Model English French German
Base 99.35 98.70 97.91
Xhosa monolingual 99.48 98.43 98.17
Xhosa parallel 99.22 96.15 95.65
Table 11: MultiBLiMP accuracy after fine-tuning Llama-3.1-8B. Scores are percentages. The near-ceiling base scores make small differences difficult to interpret.

Only the monolingual condition includes replay data from high-resource languages. Differences in retained performance therefore cannot be attributed only to the use of monolingual rather than parallel data. A parallel condition with similar replay data could also reduce forgetting.

Appendix D Artifact Licenses

We use several existing scientific artifacts in our experiments. FLORES-101 (Goyal et al., 2022) is released under a CC BY-SA 4.0 license. MultiBLiMP (Jumelet et al., 2026) is released under a CC BY 4.0 license. GlobalMMLU (Singh et al., 2025) is released under an Apache 2.0 license. For the Xhosa–English fine-tuning experiments, we use the English–Xhosa MT560 sentence-pair dataset (Gowda et al., 2021); because OPUS MT560 aggregates data from multiple sources, downstream users should consult the original OPUS MT560 provenance information before redistributing or extending the dataset.

We also use pretrained model checkpoints and evaluation software. Llama-3.1-8B is released under the Llama 3.1 Community License, and Aya-23-8B is released under CC BY-NC 4.0 with Cohere’s acceptable-use addendum. SacreBLEU (Post, 2018), which we use to compute BLEU scores, is released under an Apache 2.0 license.

In accordance with their licensing terms, all artifacts are used solely for research purposes.

Appendix E Computational Budget

All experiments use 8B-parameter language models: Llama-3.1-8B and Aya-23-8B. We do not train any models from scratch. The compute cost is dominated by translation generation and scoring, activation caching, gradient-based attribution patching, ablation experiments, and LoRA fine-tuning for the Xhosa adaptation experiments.

The total computational budget was approximately 200 GPU hours on NVIDIA A100/H100 GPUs. The largest individual runs were the causal mediation experiments, which require forward passes with activation caching and backward passes for attribution estimates across attention heads, and the LoRA fine-tuning sweeps over learning rate, LoRA rank, and number of training examples.

For activation caching, we used NNsight (Fiotto-Kaufman et al., 2025).