Recurrence Is Not Enough:
Causally Validating Multilingual SAE Translation Features in Gemma 2 and 3

1VinUniversity 2Nanyang Technological University 3University of Science and Technology of China 4Monash University

A Single Translation-Switch Feature Transfers Across Languages in Gemma 2 & 3

Gemma 2 2B IT
Gemma Scope 16k res.
Gemma 3 4B IT
Gemma Scope 2 16k, big L0
Features recurring in all 4 discovery settings 22 22
… that causally transfer 1(L10, 5717) 1(L20, 2456)
ΔCOMET when we amplify it (α = 2) +1.00 to +7.46 +0.06 to +2.50
except ar2en_ar, n.s.
ΔCOMET when we ablate it (α = 0) −0.49 to −8.93 −1.26 to −4.95
except ar2en_ar, n.s.
Language settings tested 23 23

Both models tell the same story: plenty of features recur across languages, but only one per model actually steers translation.

Summary

We asked a simple question: if an SAE feature is found in one language context, does it still do the same job when the language changes? We took the translation-initiation features reported by Wu et al. (AAAI '26), reproduced their discovery method in Gemma 2, and extended it to settings that vary the prompt language, the source language, and the target language. We then amplified and ablated every recurring feature at inference time to see which ones actually change behavior. We repeated the entire study on Gemma 3.

We got the same answer in both models. We found 22 features that recur across all four discovery settings — but when we intervened on them, nearly all had small or inconsistent effects. Only one feature per model survived: Gemma 2's (L10, 5717) and Gemma 3's (L20, 2456). Amplifying it improves COMET across the 23 language settings (every one for Gemma 2, all but ar2en_ar for Gemma 3); ablating it degrades COMET. Feature recurrence overstates cross-lingual transfer; what remains is a language-agnostic translation-initiation direction.

What we did

We reused Wu et al.'s three-stage protocol without changing a single threshold, and ran it independently under four discovery settings. We write each setting as src2tgt_prompt and each feature as (Layer, Index).

1

We discover candidates. We take SAE activations at the boundary token over 98 examples and three prompt strategies, and keep the features that fire often and strongly — up to 150 candidates per setting.

2

We filter for consistency. We ablate each candidate, collect the resulting hidden-state deltas, and keep only features whose first principal component explains ≥95% of the variance.

3

We intervene. On 862 held-out examples we ablate (α = 0) or amplify (α = 2) the feature at the last prompt token and measure the change in COMET.

We designed the four discovery settings to strip English out step by step: en2zh_en is Wu et al.'s original, then we reverse the direction and switch to a Chinese instruction (zh2en_zh), then remove English entirely (zh2ja_zh), then put English back only in the instruction (zh2ja_en). Anything in the four-way intersection therefore cannot depend on English prompts, English source text, or one translation direction.

We drew our data from WMT24++, pivoting through the shared English sentences to build non-English pairs, which gave us 6 languages and 23 settings. We ran everything on Gemma 2 2B IT with Gemma Scope 16k residual SAEs and on Gemma 3 4B IT with Gemma Scope 2 16k (big L0) SAEs.

Consistent features we foundGemma 2Gemma 3
en2zh_en4245
zh2en_zh53104
zh2ja_zh83206
zh2ja_en5965
Intersection (all 4)2222

The 22–22 match is coincidental: it arises from independent runs of the same pipeline on different models and different SAE suites.

What we found

Recurrence bought us almost nothing. When we intervened on the 22 recurring features, 21 of them behaved like a matched null: same filter, same settings, yet effects that were small, inconsistent in sign, or pointing the wrong way — even in the setting they were discovered in. Passing the consistency filter told us nothing about whether a feature actually does anything.

Exactly one feature per model looked like a switch. Amplifying (L10, 5717) gained us +2.33 to +5.22 COMET in all four discovery settings, and amplifying (L20, 2456) gained up to +2.23 while ablation cost −3.29 to −4.15. We found nothing else that moved COMET this far in both directions.

That one feature kept working well outside where we found it. Across all 23 settings, amplification helped in every single one for Gemma 2 (+1.00 to +7.46) and in all but ar2en_ar for Gemma 3, while ablation hurt everywhere for Gemma 2. Gemma 3's gains are smaller, likely a headroom effect — its baseline sits 8–18 COMET points higher.

The direction looks language-agnostic. Gemma 2 and its SAEs were trained on English-centric corpora, yet the effect survived non-English prompts, sources, and targets alike.

Across 23 language settings

Baseline COMET, and what we measured under amplification (α = 2) and ablation (α = 0). = not significant.

Setting Gemma 2 — (L10, 5717) Gemma 3 — (L20, 2456)
Base.Amp. ΔAbl. Δ Base.Amp. ΔAbl. Δ
en2zh_en63.33+4.61−5.8574.89+1.40−3.29
zh2en_zh68.17+2.33−0.4976.25+1.45−3.49
zh2ja_zh62.66+5.22−3.9676.88+2.23−3.78
zh2ja_en67.34+3.66−4.7977.75+0.59−4.15
en2ja_en61.03+5.59−7.4977.19+1.04−4.03
ja2zh_en62.19+3.07−4.1172.39+1.19−4.07
en2ar_en51.70+3.34−5.2569.55+0.58−1.63
en2ar_ar55.15+1.14−5.2369.18+0.71−1.39
ar2en_en67.00+1.36−1.0374.07+0.14†−1.32
ar2en_ar64.20+1.24−1.5669.77−0.41†+0.03†
en2ru_en57.39+7.46−8.9374.09+0.87−1.64
ru2en_ru66.92+1.40−1.2373.84+0.34†−1.80
ru2zh_en58.49+3.61−4.5070.69+1.34−4.18
en2zh_zh68.61+3.70−5.0976.66+2.27−3.29
en2zh_ja57.67+1.57−2.6065.82+2.50−4.58
en2zh_ru68.60+1.67−2.7875.24+1.66−4.70
ja2zh_ja54.90+1.00−0.6555.71+0.71−2.44
ja2zh_zh57.83+2.24−1.1274.56+0.83−3.68
ja2zh_ru59.32+1.32−1.5574.12+0.06†−2.84
en2vi_en63.13+3.44−6.7975.54+0.70−1.26
en2vi_vi66.24+2.47−4.3172.51+1.21−2.28
en2zh_vi65.00+3.61−4.6673.38+1.85−4.95
zh2ja_vi57.01+2.54−2.5072.53+2.44−4.84

We swept the intervention strength

Dose-response curves for (L10, 5717) in Gemma 2 and (L20, 2456) in Gemma 3.

We varied α over {0, 0.5, 1, 1.5, 2, 3, 4}. Suppression always hurt, moderate amplification always helped, and the effect peaked around α = 2–3 in both models — so what we see is not an artifact of one coefficient.

What we checked

Is it really about starting a translation? We classified every output with fastText lid.176 and counted how often the model failed to produce the target language at all. Amplification cut Gemma 2's non-translation rate by up to 13.7 points; ablation raised it by up to 17.3. In the outputs, ablation pushes the model into English assistant-style replies instead of translations. Gemma 3 showed the same direction at a smaller scale.

Did we miss a stronger feature? We also amplified the 20 features that were consistent in en2zh_en but fell outside the intersection. The best of them reached only +0.65 ΔCOMET, against +4.61 for (L10, 5717).

Is it just a general multilingual boost? We applied the same amplification to XL-Sum summarization and WikiANN NER. ROUGE-L moved by only −0.011 to +0.014 and micro-F1 by −0.021 to +0.012 — the gains are translation-specific.

Does the metric matter? We rescored everything with chrF++ and found the same two features standing alone, with ablation lowering chrF++ in both models.

Limitations

Our findings are specific to Gemma 2 2B IT and Gemma 3 4B IT with their Gemma Scope SAEs. We only used English, Chinese, and Japanese for discovery, so we cannot fully separate prompt language from translation direction, and a different discovery set may give a different intersection. We intervened on one feature at a time, measured quality mainly with COMET, and read the outputs by hand rather than annotating them systematically.

BibTeX

@inproceedings{nguyen-etal-2026-recurrence,
    title = "Recurrence Is Not Enough: Causally Validating Multilingual {SAE} Translation Features in {Gemma} 2 and 3",
    author = "Nguyen, Giang Son  and
      Nguyen, Nhi Ngoc-Yen  and
      Buntine, Wray  and
      Le, Dung D.",
    booktitle = "Proceedings of the Ninth Workshop on Analyzing and Interpreting Neural Networks for NLP",
    month = oct,
    year = "2026",
    address = "Budapest, Hungary",
    publisher = "Association for Computational Linguistics",
    note = "To appear"
}

Acknowledgments

This work is part of our wider multilingual machine translation project at VinUniversity, supported by the VinUniversity Cross-College Research Grant (VUNI.2324.CC06). We thank the AUMOVIO-NTU Corporate Lab for lending the GPUs we ran these experiments on, and Wu et al. for their published codebase.