Replication of Matryoshka Sparse Autoencoders
Reproducing feature hierarchy and reconstruction–sparsity experiments
Image will load when scrolled into view
Mechanistic InterpretabilitySparse AutoencodersPyTorchReplication
About the Project
I reproduced the synthetic feature-hierarchy experiment and a TinyStories comparison from the Matryoshka Sparse Autoencoders paper. The public writeup documents feature recovery in the toy setting and the reconstruction–sparsity tradeoff in the language-model experiment.
These experiments establish results in the tested toy and TinyStories settings. They do not establish transfer across modern model families.
This project informs my broader research on the generalizability of mechanistic interpretability methods: understanding which conclusions survive changes in model and experimental setting.
Project Details
StatusPublished Writeup (LessWrong)
Role
Independent Researcher
Stack
PyTorch
Sparse Autoencoders
Synthetic feature hierarchies / TinyStories