Claude Opus 5: the new leader in agentic knowledge work
New article published — Claude Opus 5 is the new leader on our agentic knowledge work benchmark, AA-Briefcase, outperforming Claude Fable 5 by nearly 150 Elo while reducing Cost per Task by 20%
Latest updates from supported sources.
New article published — Claude Opus 5 is the new leader on our agentic knowledge work benchmark, AA-Briefcase, outperforming Claude Fable 5 by nearly 150 Elo while reducing Cost per Task by 20%
New language model evaluation results available — Intelligence Index: 61
New language model evaluation results available — Intelligence Index: 60
New language model evaluation results available — Intelligence Index: 59
New language model evaluation results available — Intelligence Index: 56
New language model evaluation results available — Intelligence Index: 51
New article published — Claude Opus 5 is narrowly the most intelligent model on the Artificial Analysis Intelligence Index, offering comparable intelligence to Fable 5 at 26% lower Cost per Task
New language model evaluation results available — Intelligence Index: 39
New language model evaluation results available — Intelligence Index: 16
New article published — Thinking Machines Lab’s Inkling scores an Elo of 836 on on our agentic knowledge work benchmark AA-Briefcase
New article published — Kimi K3 is second only to Fable 5 on AA-Briefcase, our agentic knowledge work benchmark, but costs more than Opus 4.8 to run while averaging nearly an hour per task
New language model evaluation results available — Intelligence Index: 36
New language model evaluation results available — Intelligence Index: 50
New article published — Google has released Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. Both halve time per task relative to their predecessors and increase token efficiency, Gemini 3.5 Flash-Lite improves by 11 Intelligence Index points while Gemini 3.6 Flash does not improve in intelligence over 3.5 Flash
New language model evaluation results available — Intelligence Index: 44
New article published — Yeah. Grok 4.5, GPT-5.6, Muse Spark 1.1, and Kimi K3 all launched within eight days. Six labs now have a model scoring above 50 on the Artificial Analysis Intelligence Index, up from two in early June - and the price of near-frontier intelligence has collapsed
New article published — Benchmarks and Analysis of Kimi K3
New language model evaluation results available — Intelligence Index: 57
New language model evaluation results available — Intelligence Index: 41
New article published — Thinking Machines has released Inkling, the new leading U.S. open weights model, debuting at 41 on the Artificial Analysis Intelligence Index