The highest-scoring AI model on this week’s leaderboard is one you cannot use yet. The next one down costs about four times what a cheaper open-weight model charges to land within a few points of it. That gap between the top score and the practical choice is the whole story this week. This is LLM charts explained, our weekly read of the main AI model leaderboards in plain language: the real numbers, and what they do and do not prove.
We pick the charts that matter most right now, not a fixed set. The conversation this week was about what to actually run: agents doing long-horizon work, on-device and open-weight models, and the cost of all of it. So we read one rigorous leaderboard, Artificial Analysis, three ways: intelligence, price and speed. How smart each model scores, what it costs per million tokens, and how fast it generates. One board, three lenses, with the same models lined up so you can compare them column by column.
One note before the charts. A benchmark score is a proxy, not a promise. It tells you how a model did on a fixed set of tests, not how it will do on your problem. We call out the limit of each chart as we go, because that is the only way the numbers are worth anything. Every figure below is read directly from the source linked under its chart.
Where the Intelligence Index scores land right now
Artificial Analysis runs a fixed set of benchmarks and combines them into one number, the Intelligence Index. It is a weighted average across reasoning, coding, agentic and general-knowledge tests, version 4.1 of the formula. Higher means a higher average across those evals. Here are the top models, one entry per model, with the variant effort settings collapsed to each model’s headline configuration.
Source: Artificial Analysis leaderboard, Intelligence Index v4.1, viewed 25 June 2026.
Read this as a proxy, not a guarantee. The Index is an average of benchmark results, a useful signal for general capability and nothing more; it does not promise a model will do better on your specific task. Version 4.1 is a recent recalibration, so these scores are not comparable to numbers quoted in older editions. The ranking itself is stable since that recalibration, so this is not new movement, it is the same order read more cleanly.
Two things to flag. Claude Fable 5 tops the list at 60, but it is not generally available, so treat its score as a preview signal rather than a model you can ship today. And the eighth bar, Qwen3.7 Max at 46, is tied with Gemini 3.1 Pro Preview, which also scores 46; that slot is a tie on source row order, not a real gap, and we show the generally available model here so it carries through to the next two charts.
What the top models cost to run
For price and speed we narrow the field to the eight models you can deploy today. That drops Claude Fable 5, which is not generally available and lists no speed anyway, and the Gemini 3.1 Pro Preview, and it brings in MiniMax-M3. The order stays the same so this chart and the next line up bar for bar. The numbers below are the blended price, in US dollars per million tokens.
Source: Artificial Analysis leaderboard, blended price in USD per 1M tokens, viewed 25 June 2026.
Blended price is a list-price proxy, not your actual bill. It uses a fixed weighting of input and output token prices; what you pay depends on your own input-output mix, any caching, and your provider. The pattern is the honest point of the week. GLM-5.2, which scored 51, costs $0.90, and Gemini 3.5 Flash, which scored 50, costs $1.31. Both land near the top of the intelligence chart for roughly a quarter of the $3.85 to $4.35 the leading Claude and GPT models charge. A small gap in score, a large gap in price.
MiniMax-M3 is cheaper still at $0.22, but its score of 44 is a more visible step down in capability. Read it as the budget option with a real trade, not a free lunch. The cheap-and-close claim belongs to GLM-5.2, not to the cheapest bar on the chart.
How fast each model generates tokens
Same eight models, same order. This chart shows output speed, the median number of tokens each model generates per second.
Source: Artificial Analysis leaderboard, median output tokens per second, viewed 25 June 2026.
Speed here is throughput, not latency. It measures how fast text comes out once generation starts, not how long you wait for the first token, which is a separate figure. It is also a single provider-measured median that moves with load and provider, so read it as a snapshot. The point lines up with the other two charts. The highest-scoring models you can run, Claude Opus 4.8 and 4.7, are among the slowest at roughly 48 to 65 tokens per second, while Qwen3.7 Max at 208, Gemini 3.5 Flash at 195 and GLM-5.2 at 141 run several times faster. Capability and speed pull in different directions.
What intelligence, price and speed add up to
Put the three charts side by side and they disagree on purpose. The single highest benchmark score belongs to a model you cannot use yet. Among the models you can run, the smartest, Claude Opus 4.8 and GPT-5.5, are the priciest and among the slowest. The cheap open-weight GLM-5.2 and Gemini 3.5 Flash give you a near-top score for about a quarter of the price and several times the speed. MiniMax-M3 is cheaper again, with a wider gap in capability. There is no single best model here, only the right trade for your problem.
That is why we read one board across dimensions instead of crowning a winner. The question is rarely which model tops a chart. It is which trade fits your workload: where capability earns its cost, where speed matters, and where a cheaper open-weight model is plenty.
That is this week’s edition, our weekly read to keep the signal separate from the noise. If you are deciding which model to build on, or whether to build or buy at all, that is the kind of call we help teams make. Read our take on building your own versus buying enterprise software, or get in touch for a no-hype, no-obligation technical conversation. See you next week.