Skip to Main Content

Replication of Matryoshka Sparse Autoencoders

Reproducing feature hierarchy and reconstruction–sparsity experiments

Image will load when scrolled into view
Mechanistic InterpretabilitySparse AutoencodersPyTorchReplication

About the Project

I reproduced the synthetic feature-hierarchy experiment and a TinyStories comparison from the Matryoshka Sparse Autoencoders paper. The public writeup documents feature recovery in the toy setting and the reconstruction–sparsity tradeoff in the language-model experiment.

These experiments establish results in the tested toy and TinyStories settings. They do not establish transfer across modern model families.

This project informs my broader research on the generalizability of mechanistic interpretability methods: understanding which conclusions survive changes in model and experimental setting.

Project Details

StatusPublished Writeup (LessWrong)
Role
Independent Researcher
Stack
PyTorch
Sparse Autoencoders
Synthetic feature hierarchies / TinyStories