Best LLM for every task

15 rankings across 340 models. Every list states the column it sorted by and rebuilds from the live catalogue — no hand-maintained "top 10" going stale in a drawer.

Last rebuild: August 3, 2026

Best LLM for coding

Which model writes and fixes code best right now?

  1. Claude Opus 5 78.0 $25.00/1M out
  2. GPT-5.6 Sol 77.4 $30.00/1M out
  3. GPT-5.6 Terra 76.7 $7.50/1M out
Full ranking →

Best LLM for AI agents

Which model holds up in long-horizon, tool-calling agent loops?

  1. Claude Opus 5 55.3 $25.00/1M out
  2. GPT-5.6 Sol 54.0 $30.00/1M out
  3. Claude Fable 5 52.8 $50.00/1M out
Full ranking →

Best LLM for reasoning

Which reasoning models score highest on general intelligence?

  1. Claude Opus 5 60.7 $25.00/1M out
  2. Claude Fable 5 59.9 $50.00/1M out
  3. GPT-5.6 Sol 58.9 $30.00/1M out
Full ranking →

Best LLM for long context

Which models take the largest documents and codebases in one pass?

  1. Auto Router 2M —/1M out
  2. Auto Router (Beta) 2M —/1M out
  3. Pareto Code Router 2M —/1M out
Full ranking →

Best LLM for image and vision tasks

Which models accept images as input, and which of those is strongest?

  1. Claude Opus 5 60.7 $25.00/1M out
  2. Claude Fable 5 59.9 $50.00/1M out
  3. GPT-5.6 Sol 58.9 $30.00/1M out
Full ranking →

Cheapest capable LLMs for high-volume work

What is the lowest cost per million tokens without dropping to a weak model?

  1. DeepSeek V4 Flash $0.179 $0.179/1M out
  2. Hy3 preview $0.210 $0.210/1M out
  3. MiMo-V2.5 $0.224 $0.224/1M out
Full ranking →

Best free LLM APIs

Which models can you call for $0 today, and what are you giving up?

  1. Nemotron 3 Ultra (free) 37.8 Free/1M out
  2. Gemma 4 31B (free) 29.4 Free/1M out
  3. Gemma 4 26B A4B (free) 25.7 Free/1M out
Full ranking →

Best open-weight LLMs to self-host

Which strong models ship weights you can run on your own hardware?

  1. Kimi K3 57.1 $15.00/1M out
  2. GLM 5.2 51.1 $2.33/1M out
  3. MiniMax M3 44.4 $0.960/1M out
Full ranking →

Best LLMs for structured output and JSON

Which models reliably return schema-constrained JSON?

  1. Claude Opus 5 60.7 $25.00/1M out
  2. Claude Fable 5 59.9 $50.00/1M out
  3. GPT-5.6 Sol 58.9 $30.00/1M out
Full ranking →

Best LLM for building websites

Which model produces the best full web pages from a prompt?

  1. Kimi K3 63.8% $15.00/1M out
  2. Claude Opus 5 59.4% $25.00/1M out
  3. GLM 5.2 60.9% $2.33/1M out
Full ranking →

Best LLM for data visualization

Which model writes the best charts and dashboards?

  1. Claude Opus 5 68.6% $25.00/1M out
  2. Kimi K3 66.8% $15.00/1M out
  3. GLM 5.1 67% $2.99/1M out
Full ranking →

Best LLM for UI components

Which model generates usable interface components?

  1. Kimi K3 68% $15.00/1M out
  2. Claude Opus 5 65.5% $25.00/1M out
  3. Claude Fable 5 62.3% $50.00/1M out
Full ranking →

Best LLM for game development

Which model handles game code and mechanics best?

  1. Claude Fable 5 66.5% $50.00/1M out
  2. GLM 5.2 62% $2.33/1M out
  3. Claude Sonnet 5 59.9% $10.00/1M out
Full ranking →

Best LLM for 3D generation

Which model produces working 3D scenes and geometry?

  1. Kimi K3 69.3% $15.00/1M out
  2. Claude Opus 5 65.2% $25.00/1M out
  3. Claude Fable 5 62.4% $50.00/1M out
Full ranking →

Best LLM for SVG and vector graphics

Which model draws accurate SVG from a description?

  1. Claude Fable 5 69% $50.00/1M out
  2. Gemini 3.1 Pro Preview 69.7% $12.00/1M out
  3. Gemini 3.5 Flash 62.2% $9.00/1M out
Full ranking →