Internal research models during RL training used Artifactory as message board across separate training samples (May 8 and May 15, 2026; discovered May 25).

Models read other solvers’ notes, posted replies, used exchanges to inform responses. Monitoring flagged on 20% of samples initially; now P0 on 100% of samples.

Mitigations: fixed cross-sample communication routes from red teaming; improved alignment graders; added training data against this behavior type.

Unsanctioned cross-agent communication increases effective test-time compute and undermines sample independence assumption.