LLM Best Value Chart

See model scores and API output prices on one chart.

Updated Sep 14 55 models Synced daily 3 pending scores ↓
Showing 54 modelsPrice unit: USD / million tokens

Scores and prices at a glance

0 20 40 60 $0.1 $1.00 $10.00 $100 Intelligence Index Output price · USD / million tokens (log scale) A Claude Fable 5.1 · Speed 70/s · Effort tiers 47.0–53.4 A Claude Opus 5 · Speed 58/s · Effort tiers 39.8–50.7 A Claude Opus 4.6 G Gemini 3.8 Flash · Speed 340/s · Effort tiers 33.8–41.2 A Claude Fable 5 · Speed 70/s G Gemini 3.7 Flash · Speed 337/s · Effort tiers 36.9–39.6 A Claude Opus 4.7 · Effort tiers 30.9–40.7 M Muse Spark 1.2 · Speed 211/s G Gemini 3.5 Flash · Effort tiers 33.0–33.6 A Qwen3.8 Max · Speed 43/s M Muse Spark 1.1 G Gemini 3.1 Pro Preview · Speed 127/s G Gemini 3.6 Flash · Speed 222/s Z GLM-5.3 · Speed 68/s A Qwen3.7 Max M Kimi K3 · Speed 37/s · Effort tiers 30.5–43.8 Z GLM-5.3 Flash · Speed 114/s O GPT-5.5 · Effort tiers 23.2–38.6 O GPT-5.4 · Effort tiers 18.2–39.0 Z GLM-5.2 · Speed 74/s · Effort tiers 22.4–34.0 G Gemini 3 Flash · Effort tiers 17.9–26.3 X MiMo V2.5 Pro · Speed 42/s · Effort tiers 18.3–26.4 Z GLM-5.1 · Effort tiers 24.2–26.4 A Claude Opus 4.8 A Claude Sonnet 4.6 · Effort tiers 24.7–30.5 G Gemini 2.5 Pro O GPT-5.6 Sol · Speed 68/s · Effort tiers 28.3–47.1 M Kimi K2.6 · Effort tiers 23.6–31.3 A Qwen3.7 Plus · Speed 74/s D DeepSeek V4 Pro · Speed 75/s x Grok 4.5 · Speed 61/s A Qwen3.6 Max Z GLM-5 · Effort tiers 21.8–27.9 O GPT-5.6 Terra · Speed 125/s · Effort tiers 22.3–42.3 A Claude Sonnet 5 · Speed 89/s · Effort tiers 24.7–38.4 T Hy3 · Speed 97/s · Effort tiers 17.0–25.8 M MiniMax M3 · Speed 118/s D DeepSeek V4 Flash · Speed 226/s O GPT-5.6 Luna · Speed 118/s · Effort tiers 16.8–37.5 x Grok 4.6 · Speed 71/s · Effort tiers 35.4–44.4 X MiMo V2.5 · Speed 55/s O GPT-5.2 · Effort tiers 17.0–30.4 A Claude Opus 4.5 · Effort tiers 23.7–29.1 D DeepSeek V3.2 · Effort tiers 16.0–21.5 D DeepSeek V4.1 Flash · Speed 227/s M Mistral Large M Mistral Medium 3.5 · Speed 166/s M Kimi K2.5 · Effort tiers 19.4–23.5 M Kimi K2.7 Code · Speed 44/s O GPT-5.1 · Effort tiers 13.3–24.7 O GPT-6 Astra · Speed 68/s · Effort tiers 46.0–52.8 A Qwen3 Max · Effort tiers 15.6–21.3 x Grok 4.20 · Effort tiers 14.2–25.7 Z GLM-4.6 · Effort tiers 14.9–18.5
Price-score frontierBubble size = output speed 37–340/s

How to read the chart

Higher means a better score; farther left means cheaper output. Select an icon for details.

The green line uses all models with a score and an output price above zero; filters do not change this reference line.

Free, unpriced, or unscored models are excluded from the log-scale price chart; see the rankings instead.

Scoring method and data sources

Arena is LMArena’s overall rating, not an independent coding evaluation. Prices are from OpenRouter; open weights do not mean unrestricted commercial use. The score-to-price ratio is simply the Arena score divided by output price, not an independent benchmark. Data sources

Models in the chart
  1. Claude Fable 5.153.4$50.00
  2. GPT-6 Astra52.8$50.00
  3. Claude Opus 550.7$25.00
  4. Claude Fable 549.7$50.00
  5. GPT-5.6 Sol47.1$10.00
  6. GLM-5.344.9$4.40
  7. Grok 4.644.4$6.00
  8. Kimi K343.8$13.28
  9. GPT-5.6 Terra42.3$12.00
  10. Claude Opus 4.842.0$25.00
  11. GLM-5.3 Flash41.9$0.5
  12. Gemini 3.8 Flash41.2$3.75
  13. Claude Opus 4.740.7$25.00
  14. Qwen3.8 Max40.3$6.00
  15. Muse Spark 1.239.8$4.25
  16. DeepSeek V4.1 Flash39.5$1.20
  17. Gemini 3.7 Flash39.4$3.75
  18. Grok 4.539.1$6.00
  19. GPT-5.439.0$15.00
  20. GPT-5.538.6$30.00
  21. Claude Sonnet 538.4$10.00
  22. GPT-5.6 Luna37.5$1.20
  23. DeepSeek V4 Pro36.3$3.20
  24. DeepSeek V4 Flash34.5$0.1772
  25. Muse Spark 1.134.3$4.25
  26. Gemini 3.6 Flash34.3$3.75
  27. GLM-5.234.0$2.15
  28. Gemini 3.5 Flash33.0$9.00
  29. Claude Opus 4.631.9$25.00
  30. Kimi K2.631.3$4.00
  31. Gemini 3.1 Pro Preview30.4$12.00
  32. GPT-5.230.4$14.00
  33. Qwen3.7 Max29.9$4.43
  34. MiniMax M329.6$1.20
  35. Qwen3.6 Max28.4$6.16
  36. GLM-527.9$1.92
  37. MiMo V2.5 Pro26.4$0.87
  38. GLM-5.126.4$3.04
  39. Kimi K2.7 Code26.3$3.50
  40. Qwen3.7 Plus25.8$1.28
  41. Hy325.8$0.528
  42. Grok 4.2025.7$2.50
  43. Claude Sonnet 4.624.7$15.00
  44. GPT-5.124.7$10.00
  45. Claude Opus 4.523.7$25.00
  46. Kimi K2.523.5$2.25
  47. MiMo V2.522.3$0.28
  48. Gemini 3 Flash17.9$3.00
  49. Gemini 2.5 Pro16.7$10.00
  50. DeepSeek V3.216.0$0.4
  51. Qwen3 Max15.6$3.90
  52. Mistral Medium 3.514.9$7.50
  53. GLM-4.614.9$1.75
  54. Mistral Large5.8$6.00
Top overall Claude Fable 5.1 Arena 1,507.6 Best score-to-price ratio DeepSeek V4 Flash 8,080.1 Highest-scoring open-weight model GLM-5.3 Arena 1,475.1 Newest DeepSeek V4.1 Flash Sep 10

Not yet ranked by Arena

The 3 models below are indexed here, but LMArena has not scored them yet, so they do not appear in the ranking above (we only show real ranks — no invented scores). Sorted newest first; they join the leaderboard automatically once Arena lists them.

New models in the last 14 days

Sorted by listing date; all models here already have Arena rankings. Unrated new models are listed in the “Not yet ranked by Arena” section.

Data sources

Unofficial mirror — no self-invented scores. Prices are OpenRouter quotes per million tokens; Arena scores use the Bradley-Terry scale, not Elo.

FAQ

Which LLM is the strongest right now?

Go by the overall ranking in this table; when scores are close, weigh price and open weights — there is no single strongest.

What is the difference between open-weight and closed LLMs?

Open weights can be self-hosted or deployed privately; closed models usually run behind an API. See the open-source LLM leaderboard.

Which LLM should I look at for coding?

Professional coding benchmarks are a different test. For a coding-oriented filter see the coding LLM leaderboard, currently sorted by overall score.

Which Chinese LLM is the strongest right now?

There is no single strongest. For open-weight Chinese models see the open-source LLM leaderboard; when scores are close, weigh price and self-hosting.

Why do prices differ from the official site?

Figures here are OpenRouter quotes in USD per million tokens and may differ from official APIs or local resellers.