AI Safety & Ethics · Emerging · Advanced
Sparse Autoencoder Dictionary Learning
Also known as: SAEs in LLMs, Feature Disentanglement
A mechanistic interpretability method that decomposes dense LLM activations into millions of sparse, human-interpretable concepts.
What Sparse Autoencoder Dictionary Learning is
Sparse Autoencoders resolve polysemanticity, uncovering dedicated neural feature directions for concepts, entities, and behaviors.
How it works
Trains autoencoders with L1 regularization on hidden layer activations to extract high-dimensional sparse representations.
Why it matters
Allows researchers to monitor, steer, and edit internal model concepts for improved safety and alignment.
Common uses
- →LLM safety monitoring
- →Mechanistic feature steering
- →Concept disentanglement analysis
More in this collection
Browse all AI Concepts