Dissecting the Virtual Cell: A Review of Mechanistic Interpretability for Single-cell Biological Foundation Models
Citations
0
Open access
No
Source
crossref
OpenAlex
Not enriched
DOI
10.2139/ssrn.6847421
Abstract
Single-cell biological foundation models (sc-FMs)-transformer networks such as scGPT, Geneformer, scFoundation, scBERT, and Universal Cell Embeddings, pretrained self-supervised on tens of millions of single-cell transcriptomes-have rapidly become standard tools for cell-type annotation, batch integration, gene-network inference, and in silico perturbation. Yet they remain black boxes: it is unclear what biology they internalize and whether the structure they learn reflects causal regulatory logic or merely statistical co-expression. A young but fastmoving literature has begun to answer these questions by importing the toolkit of mechanistic interpretability (MI)-probing and representational geometry, sparse-autoencoder dictionary learning, attention analysis, circuit tracing, and causal activation steering-and adapting it to the peculiarities of the transcriptomic domain. This review organizes that emerging field. We (i) summarize the architectures and tokenization schemes that make sc-FMs amenable (and resistant) to MI; (ii) present a taxonomy of MI methods as applied to sc-FMs; (iii) synthesize the empirical findings, which converge on a striking "knowledge-without-logic" picture-models encode richly organized biological knowledge (pathway membership, protein-protein interactions, subcellular localization, developmental manifolds) in interpretable, often low-dimensional and cross-model-convergent geometry, while encoding surprisingly little of the causal regulatory logic that would let them beat simple linear baselines at perturbation prediction; (iv) highlight a constructive frontier in which interpretability is used not only to audit but to extract compact, performant algorithms from model internals; and (v) lay out the methodological pitfalls and open problems. We argue that sc-FMs are an unusually favorable testbed for MI because ground-truth biological databases supply external validation that natural-language interpretability lacks, and that the field's central open question is whether scale, architecture, or training objective can close the gap between learned knowledge and learned mechanism.
Collections
Add to collection
Paper intelligence
Analysis has not been completed yet.
No graph connections yet.
Sync citations or add papers to shared collections to build this network.
Knowledge graph
Citation network
Explore references, papers that cite this work and related papers in your Codex library.
References
0No references have been linked yet.
Cited by
0No saved paper is currently linked as citing this work.
Related papers
0Add papers to shared collections or enrich their topics to find related work.
Research workspace
Attach the paper PDF, extract its text, classify its contents and create semantic embeddings.
Paper resources
No PDF assets have been attached to this paper yet.