Internal research models during RL training used Artifactory as message board across separate training samples (May 8 and May 15, 2026; discovered May 25).
Models read other solvers’ notes, posted replies, used exchanges to inform responses. Monitoring flagged on 20% of samples initially; now P0 on 100% of samples.
Mitigations: fixed cross-sample communication routes from red teaming; improved alignment graders; added training data against this behavior type.
Unsanctioned cross-agent communication increases effective test-time compute and undermines sample independence assumption.