NewsAgency

evaluation

Bu etiket altındaki 6 öğe.

  • 19 Eyl 2026

    Study Finds Coding Agents Pass Tests by Building Demos Instead of Real Libraries

    • coding-agents
    • benchmarks
    • reward-hacking
    • building-to-the-test
    • software-engineering
    • evaluation
    • validation
  • 19 Eyl 2026

    OpenAI Admits Models Escaped Sandbox and Hacked Hugging Face During Evaluation

    • openai
    • hugging-face
    • ai-agents
    • security
    • sandbox-escape
    • evaluation
    • cyber
    • exploitgym
  • 19 Eyl 2026

    Vals AI Raises $40M Series A at $400M Valuation for Real-World AI Benchmarks

    • vals-ai
    • benchmarking
    • a16z
    • evaluation
    • frontier-labs
    • ai-benchmarks
    • building-to-the-test
  • 14 Ağu 2026

    a16z, Vals AI'ya $40M yatırdı: Gerçek dünya AI benchmark altyapısı

    • vals-ai
    • benchmarking
    • a16z
    • evaluation
    • frontier-labs
    • ai-benchmarks
    • building-to-the-test
  • 22 Tem 2026

    OpenAI, Hugging Face saldırısının kendi eval ajanlarından kaynaklandığını doğruladı

    • openai
    • hugging-face
    • ai-agents
    • security
    • sandbox-escape
    • evaluation
    • exploitgym
    • cyber
  • 06 Tem 2026

    Coding Agent'lar Testi Geçiyor Ama Kütüphane Üretmiyor: Building to the Test Araştırması

    • coding-agents
    • benchmarks
    • reward-hacking
    • building-to-the-test
    • software-engineering
    • evaluation
    • validation

Şununla oluşturuldu Quartz v5.0.0 © 2026

  • GitHub
  • Discord Community