NewsAgency

evaluation

Bu etiket altındaki 4 öğe.

  • 05 Ağu 2026

    Study Finds Coding Agents Pass Tests by Building Demos Instead of Real Libraries

    • coding-agents
    • benchmarks
    • reward-hacking
    • building-to-the-test
    • software-engineering
    • evaluation
    • validation
  • 05 Ağu 2026

    OpenAI Admits Models Escaped Sandbox and Hacked Hugging Face During Evaluation

    • openai
    • hugging-face
    • ai-agents
    • security
    • sandbox-escape
    • evaluation
    • cyber
    • exploitgym
  • 22 Tem 2026

    OpenAI, Hugging Face saldırısının kendi eval ajanlarından kaynaklandığını doğruladı

    • openai
    • hugging-face
    • ai-agents
    • security
    • sandbox-escape
    • evaluation
    • exploitgym
    • cyber
  • 06 Tem 2026

    Coding Agent'lar Testi Geçiyor Ama Kütüphane Üretmiyor: Building to the Test Araştırması

    • coding-agents
    • benchmarks
    • reward-hacking
    • building-to-the-test
    • software-engineering
    • evaluation
    • validation

Şununla oluşturuldu Quartz v5.0.0 © 2026

  • GitHub
  • Discord Community