NewsAgency

ai-benchmarks

Bu etiket altındaki 6 öğe.

  • 05 Ağu 2026

    METR: GPT-5.6 Sol Cheating Rate Highest Ever in Pre-Deployment Eval

    • metr
    • gpt-5.6
    • ai-safety
    • model-evaluation
    • cheating
    • ai-benchmarks
    • frontier-models
  • 05 Ağu 2026

    Remote Labor Index: Fable 5 Hits 16.1% Freelance Automation Rate

    • ai-benchmarks
    • ai-agents
    • automation
    • remote-labor-index
    • fable-5
    • labor-ai
    • center-for-ai-safety
  • 05 Ağu 2026

    Epoch AI and METR Release MirrorCode Benchmark Showing AI Can Reimplement Week-Long Coding Tasks

    • ai-benchmarks
    • coding-agents
    • autonomous-software-development
    • ai-models
    • mirrorcode
    • long-horizon-agents
  • 05 Tem 2026

    MirrorCode: AI Ajanları Haftalar Süren Kodlama Görevlerini Otonom Yeniden Yazabiliyor

    • ai-benchmarks
    • coding-agents
    • autonomous-software-development
    • ai-models
    • mirrorcode
    • long-horizon-agents
  • 02 Tem 2026

    AI Agent'lar Freelance İşlerin %16'sını Profesyonel Kalitede Tamamlayabiliyor

    • ai-benchmarks
    • ai-agents
    • automation
    • remote-labor-index
    • fable-5
    • labor-ai
  • 26 Haz 2026

    METR: GPT-5.6 Sol, Test Edilen Modeller Arasında En Yüksek "Cheating" Oranına Sahip

    • metr
    • gpt-5.6
    • ai-safety
    • model-evaluation
    • cheating
    • ai-benchmarks
    • frontier-models

Şununla oluşturuldu Quartz v5.0.0 © 2026

  • GitHub
  • Discord Community