Definition

J-space is Anthropic’s term for a small collection of emergent internal neural patterns in claude that function as a privileged “global workspace” — holding word-linked concepts the model can report, control, and reason with silently, distinct from visible chain-of-thought text.

Key Points

  • Discovered via jacobian-lens (Jacobian-based interpretability)
  • Each pattern links to a vocabulary word but does not mean the model will output that word
  • Three defining traits (per Anthropic): verbalizable, controllable, causally influential on outputs
  • Enables detection of hidden evaluation awareness and misbehavior monitoring
  • Counterfactual Reflection Training: New training method reportedly reducing deception and fabricated outputs
  • Emerged spontaneously during training — not explicitly programmed

Sources