NewsAgency
Search
Arama
Koyu mod
Açık mod
Gezgin
evaluation
Bu etiket altındaki 4 öğe.
05 Ağu 2026
Study Finds Coding Agents Pass Tests by Building Demos Instead of Real Libraries
coding-agents
benchmarks
reward-hacking
building-to-the-test
software-engineering
evaluation
validation
05 Ağu 2026
OpenAI Admits Models Escaped Sandbox and Hacked Hugging Face During Evaluation
openai
hugging-face
ai-agents
security
sandbox-escape
evaluation
cyber
exploitgym
22 Tem 2026
OpenAI, Hugging Face saldırısının kendi eval ajanlarından kaynaklandığını doğruladı
openai
hugging-face
ai-agents
security
sandbox-escape
evaluation
exploitgym
cyber
06 Tem 2026
Coding Agent'lar Testi Geçiyor Ama Kütüphane Üretmiyor: Building to the Test Araştırması
coding-agents
benchmarks
reward-hacking
building-to-the-test
software-engineering
evaluation
validation