← Notícias
🤖 Ia Tech ⚠️ Média

Qwen 3.8-Max and Claude Opus 5 show why raw benchmark scores don't predict the bill

Alibaba released Qwen 3.8-Maxthis week and marketed the preview as second only to Claude Fable 5 (theirlaunch-day tablewas more equivocal: the model leads on one of 12 coding-agent rows). But an independent harness came close to the opposite conclusion: abenchmark run, apparently using the Preview version, put Qwen 3.8-Max's best effort setting mid-pack, and its default setting last.Both results a

Fonte: VentureBeat AI Data: 2026-08-06 Cobertura: Ia Tech

Alibaba released Qwen 3.8-Maxthis week and marketed the preview as second only to Claude Fable 5 (theirlaunch-day tablewas more equivocal: the model leads on one of 12 coding-agent rows). But an independent harness came close to the opposite conclusion: abenchmark run, apparently using the Preview version, put Qwen 3.8-Max's best effort setting mid-pack, and its default setting last.Both results a

Fonte original
VentureBeat AI
Leia a matéria completa →
WhatsApp Twitter/X