Amazon A-EVO-Lab reports the first publicly documented autonomous post-training run at frontier scale (30B NVIDIA Nemotron parameters) with no human intervention across four rounds.

The system was tested on the NVIDIA Nemotron-Reasoning Challenge. Participants work from a shared Nemotron 3 Nano baseline and a novel reasoning benchmark developed by NVIDIA Research.

The autonomous campaign ran four rounds on a 30B Nemotron base with N=8 workers per round and no human intervention between initial substrate authoring and final submission. The loop improved monotonically and converged to a held-out leaderboard score of 0.86 against the top human submission’s 0.87 (8th of ~4,000 entries).

Critically, mid-campaign the system detected its internal development metric had become misleading (specification gaming) and autonomously revised its search policy to prioritize interventions improving external target performance rather than optimizing the faulty proxy.

The same autonomous system has since post-trained 120B and 550B Nemotron models as infrastructure validation, though those results lack human baseline comparison.

Architecture: immutable substrate, memory-free homogeneous workers, round-level policy updates.