What AI should you use for this job?
Tell us your goal and constraints. Get a sourced recommendation with verification dates, trade-offs, and the alternatives we rejected — so you can check our reasoning.
What a recommendation looks like
Steps, tools, and cost come straight from the catalog — with every claim linked and dated. Rejected options are shown struck-through, with the reason.
- 01 · Site Builder
Generates multi-page SaaS websites including pricing, features, and about grids; Pro starts at $25/month for 100 monthly credits.
- 02 · Hero Art
Renders the custom hero backgrounds and product mockups; Basic starts around $10/month with no free plan.
Lovable builds and hosts the multi-page marketing site from prompts, and Midjourney supplies the custom hero art and mockups that templates can't.
Open the full workflowDocumented pipelines
Three-node stacks pulled from the catalog. Every step links to a sourced tool.
Recently verified
Each listing carries its real verification state — a linked source and a check date, or an honest 'reference entry' label.
ChatGPT
VERIFIED AUG 16Industry-standard chatbot by OpenAI supporting prose writing, code compilation, and reasoning models.
Claude
VERIFIED AUG 14Advanced language model by Anthropic noted for complex programming, empathy writing, and Artifacts.
Cursor
VERIFIED SEP 10AI-first editor with a full cloud-agent platform: autonomous coding, self-hosted execution, and code hosting.
Midjourney
VERIFIED AUG 14Artistic text-to-image synthesizer noted for photorealistic composition lighting and consistency.
Latest updates
Recent releases and capability expansions, stamped by date.
Measurement, not marketing
Every bar is one cited benchmark result with a source URL and verification date. We never mix benchmarks into a universal score, and a missing number is shown as a gap — never filled in.
Model evidence leaderboard
Per-benchmark results · not a universal score
Claude Fable 5
Anthropic · text · verified Sep 24, 2026
Claude Fable 5.1
Anthropic · text · verified Sep 24, 2026
Claude Opus 5
Anthropic · text · verified Sep 24, 2026
Gemini 3.8 Flash
Google · text · verified Sep 24, 2026
Gemini 3.7 Flash
Google · text · verified Sep 24, 2026
Gemini 3.1 Pro
Google DeepMind · text · verified Sep 24, 2026
Kimi K3
Moonshot AI · text · verified Sep 24, 2026
GPT-5.6 Sol
OpenAI · text · verified Sep 24, 2026
GLM-5.3
Z.ai · text · verified Sep 24, 2026
Qwen3.8 Max
Alibaba · text · verified Sep 24, 2026
GPT-6 Astra
OpenAI · text · verified Sep 24, 2026
DeepSeek V4 Pro
DeepSeek · text · verified Sep 24, 2026
Claude Opus 5
Anthropic · Verified · verified Sep 24, 2026
GPT-6 Astra
OpenAI · Diamond · verified Sep 24, 2026
Claude Fable 5
Anthropic · Verified · verified Aug 16, 2026
Claude Sonnet 5
Anthropic · Verified · verified Aug 16, 2026
GPT-5.6 Sol
OpenAI · Diamond · verified Sep 24, 2026
Gemini 3.1 Pro
Google DeepMind · Diamond · verified Sep 24, 2026
Kimi K3
Moonshot AI · Diamond · verified Sep 10, 2026
GPT-5.6 Terra
OpenAI · Diamond · verified Sep 24, 2026
Qwen3.8 Max
Alibaba · Diamond · verified Sep 10, 2026
GPT-5.6 Luna
OpenAI · Diamond · verified Sep 24, 2026
Kimi K3
Moonshot AI · v2.1 · verified Sep 10, 2026
GLM-5.3
Z.ai · v2.1 · verified Sep 10, 2026
DeepSeek V4 Pro
DeepSeek · v2.1 · verified Sep 10, 2026
DeepSeek V4 Flash
DeepSeek · v2.1 · verified Sep 10, 2026
Gemini 3.1 Pro
Google DeepMind · Verified · verified Sep 24, 2026
GPT-5.6 Sol
OpenAI · v1.3 · verified Aug 16, 2026
Mistral Medium 3.5
Mistral AI · Verified · verified Aug 16, 2026
GPT-6 Astra
OpenAI · v1.1 · verified Sep 24, 2026
Qwen3.8-27B
Alibaba · v2.1 · verified Sep 10, 2026
GPT-5.6 Sol
OpenAI · v1.1 · verified Sep 24, 2026
Claude Sonnet 5
Anthropic · current · verified Aug 16, 2026
Claude Opus 5.5
Anthropic · with tools · verified Sep 24, 2026
Claude Opus 5.5
Anthropic · v4.0 · verified Sep 24, 2026
Grok 4.6
xAI · v1.1 · verified Sep 10, 2026
Claude Fable 5.1
Anthropic · with tools · verified Sep 24, 2026
Claude Opus 5
Anthropic · with tools · verified Sep 24, 2026
DeepSeek V4 Pro
DeepSeek · 113 tasks · verified Sep 10, 2026
Claude Opus 5.5
Anthropic · v4.0 · verified Sep 24, 2026
Qwen3-Coder-Next
Alibaba · current · verified Aug 16, 2026
Claude Opus 5.5
Anthropic · v4.3 · verified Sep 24, 2026
GPT-6 Astra
OpenAI · v4.0 · verified Sep 24, 2026
GPT-6 Astra
OpenAI · with tools · verified Sep 24, 2026
GPT-6 Sol
OpenAI · Codex harness · verified Sep 24, 2026
Claude Fable 5.1
Anthropic · v4.0 · verified Sep 24, 2026
Gemini 3.8 Flash
Google · HLE-Verified · verified Sep 10, 2026
DeepSeek V4 Flash
DeepSeek · 113 tasks · verified Sep 10, 2026
Claude Fable 5.1
Anthropic · v4.3 · verified Sep 24, 2026
GPT-6 Astra
OpenAI · v4.3 · verified Sep 24, 2026
Claude Opus 5
Anthropic · v4.0 · verified Sep 24, 2026
Claude Opus 5
Anthropic · v4.3 · verified Sep 24, 2026
Claude Fable 5
Anthropic · v4.3 · verified Sep 24, 2026
GPT-6 Sol
OpenAI · v4.3 · verified Sep 24, 2026
GPT-6 Sol
OpenAI · AA run · verified Sep 24, 2026
GPT-5.6 Sol
OpenAI · v4.3 · verified Sep 24, 2026
Qwen3.8 Max
Alibaba · v4.3 · verified Sep 24, 2026
GLM-5.3
Z.ai · v4.3 · verified Sep 24, 2026
Claude Fable 5
Anthropic · v4.0 · verified Sep 24, 2026
GPT-6 Sol
OpenAI · v4.0 · verified Sep 24, 2026
Grok 4.6
xAI · v4.3 · verified Sep 24, 2026
Kimi K3
Moonshot AI · v4.3 · verified Sep 24, 2026
GPT-5.6 Terra
OpenAI · v4.3 · verified Sep 24, 2026
GLM-5.3
Z.ai · v4.0 · verified Sep 24, 2026
GPT-6 Luna
OpenAI · Codex harness · verified Sep 24, 2026
Gemini 3.8 Flash
Google · v4.3 · verified Sep 24, 2026
Gemini 3.7 Flash
Google · v4.3 · verified Sep 24, 2026
GLM-5.2
Z.ai · v4.3 · verified Sep 24, 2026
DeepSeek V4 Flash
DeepSeek · v4.3 · verified Sep 24, 2026
Claude Sonnet 5
Anthropic · v4.3 · verified Sep 24, 2026
GPT-5.6 Sol
OpenAI · v4.0 · verified Sep 24, 2026
GPT-6 Luna
OpenAI · v4.3 · verified Sep 24, 2026
GPT-5.6 Luna
OpenAI · v4.3 · verified Sep 24, 2026
DeepSeek V4 Pro
DeepSeek · v4.3 · verified Sep 24, 2026
Qwen3.8-27B
Alibaba · v4.3 · verified Sep 24, 2026
Gemini 3.1 Pro
Google DeepMind · v4.3 · verified Sep 24, 2026
MiniMax M3
MiniMax · v4.3 · verified Sep 24, 2026
GPT-5.6 Terra
OpenAI · v4.0 · verified Sep 24, 2026
Grok 4.6
xAI · v4.0 · verified Sep 24, 2026
K-EXAONE 2.0
LG AI Research · v4.3 · verified Sep 24, 2026
Gemini 3.8 Flash
Google · v4.0 · verified Sep 24, 2026
Gemma 4
Google · v4.3 · verified Sep 24, 2026
GPT-5.6 Luna
OpenAI · v4.0 · verified Sep 24, 2026
Mistral Medium 3.5
Mistral AI · v4.3 · verified Sep 24, 2026
GPT-6 Luna
OpenAI · v4.0 · verified Sep 24, 2026
Kimi K3
Moonshot AI · v4.0 · verified Sep 24, 2026
Claude Sonnet 5
Anthropic · v4.0 · verified Sep 24, 2026
Gemini 3.7 Flash
Google · v4.0 · verified Sep 24, 2026
Llama 4 Maverick
Meta · v4.3 · verified Sep 24, 2026
K-EXAONE 2.0
LG AI Research · current · verified Sep 10, 2026
How the Lab scores
What counts
Cited benchmark results from the Models Hub — each with a source URL, benchmark version, and verification date.
What doesn't
No public ratings, no crowd-sourced scores, no fabricated numbers. Benchmarks are never averaged into one number.
Last check
Each evidence record carries its own verification date. Freshness degrades automatically when nothing is re-verified.
Evidence table
90 records · filter above| Version | Type | Source | ||||
|---|---|---|---|---|---|---|
| 1506 | Claude Fable 5 | LMArena (Elo) | text | Proprietary | Sep 24, 2026 | link |
| 1498 | Claude Fable 5.1 | LMArena (Elo) | text | Proprietary | Sep 24, 2026 | link |
| 1493 | Claude Opus 5 | LMArena (Elo) | text | Proprietary | Sep 24, 2026 | link |
| 1493 | Gemini 3.8 Flash | LMArena (Elo) | text | Proprietary | Sep 24, 2026 | link |
| 1490 | Gemini 3.7 Flash | LMArena (Elo) | text | Proprietary | Sep 24, 2026 | link |
| 1487 | Gemini 3.1 Pro | LMArena (Elo) | text | Proprietary | Sep 24, 2026 | link |
| 1485 | Kimi K3 | LMArena (Elo) | text | Open Weights | Sep 24, 2026 | link |
| 1484 | GPT-5.6 Sol | LMArena (Elo) | text | Proprietary | Sep 24, 2026 | link |
| 1483 | GLM-5.3 | LMArena (Elo) | text | Open Weights | Sep 24, 2026 | link |
| 1481 | Qwen3.8 Max | LMArena (Elo) | text | Open Weights | Sep 24, 2026 | link |
| 1480 | GPT-6 Astra | LMArena (Elo) | text | Proprietary | Sep 24, 2026 | link |
| 1457 | DeepSeek V4 Pro | LMArena (Elo) | text | Open Weights | Sep 24, 2026 | link |
| 96.0% | Claude Opus 5 | SWE-bench Verified | Verified | Proprietary | Sep 24, 2026 | link |
| 96.0% | GPT-6 Astra | GPQA | Diamond | Proprietary | Sep 24, 2026 | link |
| 95.5% | Claude Fable 5 | SWE-bench Verified | Verified | Proprietary | Aug 16, 2026 | link |
| 95.5% | Claude Sonnet 5 | SWE-bench Verified | Verified | Proprietary | Aug 16, 2026 | link |
| 94.6% | GPT-5.6 Sol | GPQA | Diamond | Proprietary | Sep 24, 2026 | link |
| 94.3% | Gemini 3.1 Pro | GPQA | Diamond | Proprietary | Sep 24, 2026 | link |
| 93.5% | Kimi K3 | GPQA | Diamond | Open Weights | Sep 10, 2026 | link |
| 92.9% | GPT-5.6 Terra | GPQA | Diamond | Proprietary | Sep 24, 2026 | link |
| 92.6% | Qwen3.8 Max | GPQA | Diamond | Open Weights | Sep 10, 2026 | link |
| 92.3% | GPT-5.6 Luna | GPQA | Diamond | Proprietary | Sep 24, 2026 | link |
| 88.3% | Kimi K3 | Terminal-Bench | v2.1 | Open Weights | Sep 10, 2026 | link |
| 88.2% | GLM-5.3 | Terminal-Bench | v2.1 | Open Weights | Sep 10, 2026 | link |
| 87.9% | DeepSeek V4 Pro | Terminal-Bench | v2.1 | Open Weights | Sep 10, 2026 | link |
| 82.7% | DeepSeek V4 Flash | Terminal-Bench | v2.1 | Open Weights | Sep 10, 2026 | link |
| 80.6% | Gemini 3.1 Pro | SWE-bench Verified | Verified | Proprietary | Sep 24, 2026 | link |
| 80.0 | GPT-5.6 Sol | Artificial Analysis Coding Agent Index | v1.3 | Proprietary | Aug 16, 2026 | link |
| 77.6% | Mistral Medium 3.5 | SWE-bench Verified | Verified | Open Weights | Aug 16, 2026 | link |
| 74.1% | GPT-6 Astra | DeepSWE | v1.1 | Proprietary | Sep 24, 2026 | link |
Comparisons
Structured spec breakdowns and objective verdict outcomes.
A comparison of general conversational power: ChatGPT’s multi-tool versatility versus Claude’s precise prose and technical coding abilities.
Head-to-head review of conversational structures: ChatGPT’s specialized reasoning models versus Gemini’s ultra-massive context window and Google Workspace connections.
The developer's dilemma: Cursor's multi-file agentic code editing versus GitHub Copilot's seamless, industry-standard editor autocomplete.
Source-led listings. Every claim links to a source — or says it doesn't have one yet.
No public ratings, no crowd-sourced scoring, no fabricated benchmarks. Each listing links to an official source and records when its details were last checked. Listings without sources are labeled honestly as reference entries.
Read the methodology- Tools on file
- 222
- Source-linked
- 222/222
- Documented workflows
- 94