Definition
Mechanistic interpretability is the research field that reverse-engineers how neural networks compute — identifying circuits, features, and internal representations that causally drive model behavior, as opposed to only observing inputs and outputs.
Key Points
- Jacobian Lens (J-lens): Anthropic technique (July 2026) that linearly transports internal activations to final-layer vocabulary space via average input–output Jacobians (jacobian-lens)
- J-space: Emergent internal “global workspace” in claude — word-linked neural patterns enabling silent reasoning (j-space)
- Global Workspace Theory analogy: Neuroscience framework comparing J-space to consciously accessible working memory — analogy only, not consciousness claim
- Safety applications: Detecting evaluation awareness, monitoring misbehavior, Counterfactual Reflection Training to reduce deception
- Open-source tooling:
anthropics/jacobian-lens(Apache 2.0)
J-space findings are first-party anthropic research not yet externally replicated. Avoid anthropomorphic "consciousness" framing.