← Research papers
2026unread

Induction Circuit Stability Under Fine-Tuning A Mechanistic Interpretability Study

Min Htet Myet
Publisher page
Open graph

Citations

0

Open access

No

Source

crossref

OpenAlex

Not enriched

DOI

10.21203/rs.3.rs-10067094/v1

Abstract

Abstract We ask whether induction head circuits in attention-only transformers remain structurally stable under fine-tuning on a narrow distribution, and at what training step any structural change first occurs. We fine-tune the two-layer attn-only-2l transformer on 500,000 tokens of Python code and a matched TinyStories prose control across three random seeds, measuring per-head induction scores and activation-patching attribution at every 100-step checkpoint; at initialisation, L1H6 has a prefix-matching score of 0.408 ± 0.103 (12.4× the next head) and attribution 0.952, replicating the canonical induction circuit of Olsson et al. [6]. Across all three seeds, the induction circuit strengthens rather than degrades under both fine-tuning conditions: L1H6’s prefix-matching score rises from 0.408 to 0.591 ± 0.005 (code) and 0.583 ± 0.002 (prose) after 100 steps, and to 0.646 ± 0.007 and 0.616 ± 0.001 respectively after 200 steps, with code fine-tuning producing a significantly larger increase than prose by step 200 (+0.030, 4.1× the pooled standard deviation) — a result not previously reported for this model class. These results demonstrate that circuit-level diagnostics are a necessary complement to capability benchmarks when evaluating safety-relevant properties of fine-tuned language models, not only to detect degradation but also to detect unexpected and domain-dependent reinforcement.

Collections

Add to collection

Paper intelligence

Analysis has not been completed yet.

No graph connections yet.

Sync citations or add papers to shared collections to build this network.

Knowledge graph

Citation network

Explore references, papers that cite this work and related papers in your Codex library.

References

0

No references have been linked yet.

Cited by

0

No saved paper is currently linked as citing this work.

Related papers

0

Add papers to shared collections or enrich their topics to find related work.

Research workspace

Attach the paper PDF, extract its text, classify its contents and create semantic embeddings.

Attach PDF

Upload the research paper so Codex can extract, chunk and search its contents.

Paper resources

No PDF assets have been attached to this paper yet.