Definition

Mechanistic interpretability is the research field that reverse-engineers how neural networks compute — identifying circuits, features, and internal representations that causally drive model behavior, as opposed to only observing inputs and outputs.

Key Points

  • Jacobian Lens (J-lens): Anthropic technique (July 2026) that linearly transports internal activations to final-layer vocabulary space via average input–output Jacobians (jacobian-lens)
  • J-space: Emergent internal “global workspace” in claude — word-linked neural patterns enabling silent reasoning (j-space)
  • Global Workspace Theory analogy: Neuroscience framework comparing J-space to consciously accessible working memory — analogy only, not consciousness claim
  • Safety applications: Detecting evaluation awareness, monitoring misbehavior, Counterfactual Reflection Training to reduce deception
  • Open-source tooling: anthropics/jacobian-lens (Apache 2.0)

J-space findings are first-party anthropic research not yet externally replicated. Avoid anthropomorphic "consciousness" framing.

Sources