Best LLM for image and vision tasks (2026)
Which models accept images as input, and which of those is strongest?
Rebuilt from live model data · August 3, 2026
How we ranked
Models whose input modalities include images, ranked by Artificial Analysis Intelligence Index.
Top 3 for image and vision tasks
- 1
Claude Opus 5
Intelligence index: 60.7 $5.00 in / $25.00 out per 1M 1M context 5 providersClaude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...
Vision Tool calling Reasoning Released Jul 24, 2026 - 2
Claude Fable 5
Intelligence index: 59.9 $10.00 in / $50.00 out per 1M 1M context 4 providersClaude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...
Vision Tool calling Reasoning Released Jun 9, 2026 - 3
GPT-5.6 Sol
Intelligence index: 58.9 $5.00 in / $30.00 out per 1M 1.1M context 2 providersGPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...
Vision Tool calling Reasoning Released Jul 9, 2026
Full ranking
| # | Model | Intelligence index | Output / 1M | Context | Providers | Weights |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 60.7 | $25.00 | 1M | 5 | Closed |
| 2 | Claude Fable 5 | 59.9 | $50.00 | 1M | 4 | Closed |
| 3 | GPT-5.6 Sol | 58.9 | $30.00 | 1.1M | 2 | Closed |
| 4 | Kimi K3 | 57.1 | $15.00 | 1.0M | 6 | Open |
| 5 | Claude Opus 4.8 | 55.7 | $25.00 | 1M | 4 | Closed |
| 6 | GPT-5.6 Terra | 55.0 | $7.50 | 1.1M | 2 | Closed |
| 7 | GPT-5.5 | 54.8 | $30.00 | 1.1M | 2 | Closed |
| 8 | Grok 4.5 | 53.8 | $6.00 | 500K | 1 | Closed |
| 9 | Claude Opus 4.7 | 53.5 | $25.00 | 1M | 4 | Closed |
| 10 | Claude Sonnet 5 | 53.4 | $10.00 | 1M | 4 | Closed |
| 11 | GPT-5.4 | 51.4 | $15.00 | 1.1M | 2 | Closed |
| 12 | GPT-5.6 Luna | 51.2 | $3.00 | 1.1M | 2 | Closed |
FAQ
Which models accept images as input, and which of those is strongest?
Claude Opus 5 leads for image and vision tasks with intelligence index 60.7, at $25.00 per 1M output tokens. Claude Fable 5 (59.9) and GPT-5.6 Sol (58.9) follow. Models whose input modalities include images, ranked by Artificial Analysis Intelligence Index.
How is this ranking produced?
Models whose input modalities include images, ranked by Artificial Analysis Intelligence Index. The list rebuilds from the live model catalogue on every deploy, so a model that launches or changes price appears here without an editor rewriting the page. Last rebuild: August 3, 2026.
What does the top pick cost to run?
Claude Opus 5 costs about $12.50 for a workload of 1M input + 300K output tokens, based on the cheapest of 5 provider(s).