NewsAgency

ai-benchmarks

Bu etiket altındaki 8 öğe.

  • 19 Eyl 2026

    METR: GPT-5.6 Sol Cheating Rate Highest Ever in Pre-Deployment Eval

    • metr
    • gpt-5.6
    • ai-safety
    • model-evaluation
    • cheating
    • ai-benchmarks
    • frontier-models
  • 19 Eyl 2026

    Remote Labor Index: Fable 5 Hits 16.1% Freelance Automation Rate

    • ai-benchmarks
    • ai-agents
    • automation
    • remote-labor-index
    • fable-5
    • labor-ai
    • center-for-ai-safety
  • 19 Eyl 2026

    Epoch AI and METR Release MirrorCode Benchmark Showing AI Can Reimplement Week-Long Coding Tasks

    • ai-benchmarks
    • coding-agents
    • autonomous-software-development
    • ai-models
    • mirrorcode
    • long-horizon-agents
  • 19 Eyl 2026

    Vals AI Raises $40M Series A at $400M Valuation for Real-World AI Benchmarks

    • vals-ai
    • benchmarking
    • a16z
    • evaluation
    • frontier-labs
    • ai-benchmarks
    • building-to-the-test
  • 14 Ağu 2026

    a16z, Vals AI'ya $40M yatırdı: Gerçek dünya AI benchmark altyapısı

    • vals-ai
    • benchmarking
    • a16z
    • evaluation
    • frontier-labs
    • ai-benchmarks
    • building-to-the-test
  • 05 Tem 2026

    MirrorCode: AI Ajanları Haftalar Süren Kodlama Görevlerini Otonom Yeniden Yazabiliyor

    • ai-benchmarks
    • coding-agents
    • autonomous-software-development
    • ai-models
    • mirrorcode
    • long-horizon-agents
  • 02 Tem 2026

    AI Agent'lar Freelance İşlerin %16'sını Profesyonel Kalitede Tamamlayabiliyor

    • ai-benchmarks
    • ai-agents
    • automation
    • remote-labor-index
    • fable-5
    • labor-ai
  • 26 Haz 2026

    METR: GPT-5.6 Sol, Test Edilen Modeller Arasında En Yüksek "Cheating" Oranına Sahip

    • metr
    • gpt-5.6
    • ai-safety
    • model-evaluation
    • cheating
    • ai-benchmarks
    • frontier-models

Şununla oluşturuldu Quartz v5.0.0 © 2026

  • GitHub
  • Discord Community