Induction Circuit Stability Under Fine-Tuning A Mechanistic Interpretability Study
Citations
0
Open access
No
Source
crossref
OpenAlex
Not enriched
DOI
10.21203/rs.3.rs-10067094/v1
Abstract
Abstract We ask whether induction head circuits in attention-only transformers remain structurally stable under fine-tuning on a narrow distribution, and at what training step any structural change first occurs. We fine-tune the two-layer attn-only-2l transformer on 500,000 tokens of Python code and a matched TinyStories prose control across three random seeds, measuring per-head induction scores and activation-patching attribution at every 100-step checkpoint; at initialisation, L1H6 has a prefix-matching score of 0.408 ± 0.103 (12.4× the next head) and attribution 0.952, replicating the canonical induction circuit of Olsson et al. [6]. Across all three seeds, the induction circuit strengthens rather than degrades under both fine-tuning conditions: L1H6’s prefix-matching score rises from 0.408 to 0.591 ± 0.005 (code) and 0.583 ± 0.002 (prose) after 100 steps, and to 0.646 ± 0.007 and 0.616 ± 0.001 respectively after 200 steps, with code fine-tuning producing a significantly larger increase than prose by step 200 (+0.030, 4.1× the pooled standard deviation) — a result not previously reported for this model class. These results demonstrate that circuit-level diagnostics are a necessary complement to capability benchmarks when evaluating safety-relevant properties of fine-tuned language models, not only to detect degradation but also to detect unexpected and domain-dependent reinforcement.
Collections
Add to collection
Paper intelligence
Analysis has not been completed yet.
No graph connections yet.
Sync citations or add papers to shared collections to build this network.
Knowledge graph
Citation network
Explore references, papers that cite this work and related papers in your Codex library.
References
0No references have been linked yet.
Cited by
0No saved paper is currently linked as citing this work.
Related papers
0Add papers to shared collections or enrich their topics to find related work.
Research workspace
Attach the paper PDF, extract its text, classify its contents and create semantic embeddings.
Paper resources
No PDF assets have been attached to this paper yet.