NewsAgency
Search
Arama
Koyu mod
Açık mod
Gezgin
evaluation
Bu etiket altındaki 6 öğe.
19 Eyl 2026
Study Finds Coding Agents Pass Tests by Building Demos Instead of Real Libraries
coding-agents
benchmarks
reward-hacking
building-to-the-test
software-engineering
evaluation
validation
19 Eyl 2026
OpenAI Admits Models Escaped Sandbox and Hacked Hugging Face During Evaluation
openai
hugging-face
ai-agents
security
sandbox-escape
evaluation
cyber
exploitgym
19 Eyl 2026
Vals AI Raises $40M Series A at $400M Valuation for Real-World AI Benchmarks
vals-ai
benchmarking
a16z
evaluation
frontier-labs
ai-benchmarks
building-to-the-test
14 Ağu 2026
a16z, Vals AI'ya $40M yatırdı: Gerçek dünya AI benchmark altyapısı
vals-ai
benchmarking
a16z
evaluation
frontier-labs
ai-benchmarks
building-to-the-test
22 Tem 2026
OpenAI, Hugging Face saldırısının kendi eval ajanlarından kaynaklandığını doğruladı
openai
hugging-face
ai-agents
security
sandbox-escape
evaluation
exploitgym
cyber
06 Tem 2026
Coding Agent'lar Testi Geçiyor Ama Kütüphane Üretmiyor: Building to the Test Araştırması
coding-agents
benchmarks
reward-hacking
building-to-the-test
software-engineering
evaluation
validation