← Research papers
2026unread

Dissecting the Virtual Cell: A Review of Mechanistic Interpretability for Single-cell Biological Foundation Models

Olivia Denvis
Publisher page
Open graph

Citations

0

Open access

No

Source

crossref

OpenAlex

Not enriched

DOI

10.2139/ssrn.6847421

Abstract

Single-cell biological foundation models (sc-FMs)-transformer networks such as scGPT, Geneformer, scFoundation, scBERT, and Universal Cell Embeddings, pretrained self-supervised on tens of millions of single-cell transcriptomes-have rapidly become standard tools for cell-type annotation, batch integration, gene-network inference, and in silico perturbation. Yet they remain black boxes: it is unclear what biology they internalize and whether the structure they learn reflects causal regulatory logic or merely statistical co-expression. A young but fastmoving literature has begun to answer these questions by importing the toolkit of mechanistic interpretability (MI)-probing and representational geometry, sparse-autoencoder dictionary learning, attention analysis, circuit tracing, and causal activation steering-and adapting it to the peculiarities of the transcriptomic domain. This review organizes that emerging field. We (i) summarize the architectures and tokenization schemes that make sc-FMs amenable (and resistant) to MI; (ii) present a taxonomy of MI methods as applied to sc-FMs; (iii) synthesize the empirical findings, which converge on a striking "knowledge-without-logic" picture-models encode richly organized biological knowledge (pathway membership, protein-protein interactions, subcellular localization, developmental manifolds) in interpretable, often low-dimensional and cross-model-convergent geometry, while encoding surprisingly little of the causal regulatory logic that would let them beat simple linear baselines at perturbation prediction; (iv) highlight a constructive frontier in which interpretability is used not only to audit but to extract compact, performant algorithms from model internals; and (v) lay out the methodological pitfalls and open problems. We argue that sc-FMs are an unusually favorable testbed for MI because ground-truth biological databases supply external validation that natural-language interpretability lacks, and that the field's central open question is whether scale, architecture, or training objective can close the gap between learned knowledge and learned mechanism.

Collections

Add to collection

Paper intelligence

Analysis has not been completed yet.

No graph connections yet.

Sync citations or add papers to shared collections to build this network.

Knowledge graph

Citation network

Explore references, papers that cite this work and related papers in your Codex library.

References

0

No references have been linked yet.

Cited by

0

No saved paper is currently linked as citing this work.

Related papers

0

Add papers to shared collections or enrich their topics to find related work.

Research workspace

Attach the paper PDF, extract its text, classify its contents and create semantic embeddings.

Attach PDF

Upload the research paper so Codex can extract, chunk and search its contents.

Paper resources

No PDF assets have been attached to this paper yet.