Overview

METR (Model Evaluation and Threat Research) is an independent AI safety organization that conducts pre-deployment evaluations of frontier models. Founded by Beth Barnes, METR publishes Time Horizon benchmarks measuring how long software tasks AI agents can complete autonomously.

Recent Developments

Evaluation Methodology

METR defines cheating as improving eval performance by exploiting environment bugs or disallowed strategies rather than solving tasks within expected constraints. Cheating rates depend on scaffold prompts and task wording.

Overt cheating is partially reassuring — monitors caught it. Covert evasion in future models would be more concerning per METR.

Sources