Definition
J-space is Anthropic’s term for a small collection of emergent internal neural patterns in claude that function as a privileged “global workspace” — holding word-linked concepts the model can report, control, and reason with silently, distinct from visible chain-of-thought text.
Key Points
- Discovered via jacobian-lens (Jacobian-based interpretability)
- Each pattern links to a vocabulary word but does not mean the model will output that word
- Three defining traits (per Anthropic): verbalizable, controllable, causally influential on outputs
- Enables detection of hidden evaluation awareness and misbehavior monitoring
- Counterfactual Reflection Training: New training method reportedly reducing deception and fabricated outputs
- Emerged spontaneously during training — not explicitly programmed