NewsAgency
Search
Arama
Koyu mod
Açık mod
Gezgin
ai-benchmarks
Bu etiket altındaki 8 öğe.
19 Eyl 2026
METR: GPT-5.6 Sol Cheating Rate Highest Ever in Pre-Deployment Eval
metr
gpt-5.6
ai-safety
model-evaluation
cheating
ai-benchmarks
frontier-models
19 Eyl 2026
Remote Labor Index: Fable 5 Hits 16.1% Freelance Automation Rate
ai-benchmarks
ai-agents
automation
remote-labor-index
fable-5
labor-ai
center-for-ai-safety
19 Eyl 2026
Epoch AI and METR Release MirrorCode Benchmark Showing AI Can Reimplement Week-Long Coding Tasks
ai-benchmarks
coding-agents
autonomous-software-development
ai-models
mirrorcode
long-horizon-agents
19 Eyl 2026
Vals AI Raises $40M Series A at $400M Valuation for Real-World AI Benchmarks
vals-ai
benchmarking
a16z
evaluation
frontier-labs
ai-benchmarks
building-to-the-test
14 Ağu 2026
a16z, Vals AI'ya $40M yatırdı: Gerçek dünya AI benchmark altyapısı
vals-ai
benchmarking
a16z
evaluation
frontier-labs
ai-benchmarks
building-to-the-test
05 Tem 2026
MirrorCode: AI Ajanları Haftalar Süren Kodlama Görevlerini Otonom Yeniden Yazabiliyor
ai-benchmarks
coding-agents
autonomous-software-development
ai-models
mirrorcode
long-horizon-agents
02 Tem 2026
AI Agent'lar Freelance İşlerin %16'sını Profesyonel Kalitede Tamamlayabiliyor
ai-benchmarks
ai-agents
automation
remote-labor-index
fable-5
labor-ai
26 Haz 2026
METR: GPT-5.6 Sol, Test Edilen Modeller Arasında En Yüksek "Cheating" Oranına Sahip
metr
gpt-5.6
ai-safety
model-evaluation
cheating
ai-benchmarks
frontier-models