Decision Intelligence Platform

What AI should you use for this job?

Tell us your goal and constraints. Get a sourced recommendation with verification dates, trade-offs, and the alternatives we rejected — so you can check our reasoning.

A real determination, with its sources

What a recommendation looks like

Steps, tools, and cost come straight from the catalog — with every claim linked and dated. Rejected options are shown struck-through, with the reason.

Goal · Construct a complete, multi-page SaaS marketing website featuring pricing tiers, feature lists, documentation pages, and contact captures.
  1. 01 · Site Builder

    Generates multi-page SaaS websites including pricing, features, and about grids; Pro starts at $25/month for 100 monthly credits.

  2. 02 · Hero Art

    Renders the custom hero backgrounds and product mockups; Basic starts around $10/month with no free plan.

StackPaid Site: Lovable + Midjourney
Estimated cost~$35/month (Lovable Pro $25 + Midjourney Basic $10)

Lovable builds and hosts the multi-page marketing site from prompts, and Midjourney supplies the custom hero art and mockups that templates can't.

Open the full workflow
Workflows

Documented pipelines

Three-node stacks pulled from the catalog. Every step links to a sourced tool.

All workflows
Honestack Benchmark Lab
Evidence in Models Hub

Measurement, not marketing

Every bar is one cited benchmark result with a source URL and verification date. We never mix benchmarks into a universal score, and a missing number is shown as a gap — never filled in.

Evidence records
90
sourced benchmark results
Models with evidence
28/31
from the Models Hub
Benchmarks
13
each with a real source
Sourced tools
222/222
222 verified today
Free / freemium
187
of catalog

Model evidence leaderboard

Per-benchmark results · not a universal score

LMArena (Elo)

Claude Fable 5

Anthropic · text · verified Sep 24, 2026

1506
LMArena (Elo)

Claude Fable 5.1

Anthropic · text · verified Sep 24, 2026

1498
LMArena (Elo)

Claude Opus 5

Anthropic · text · verified Sep 24, 2026

1493
LMArena (Elo)

Gemini 3.8 Flash

Google · text · verified Sep 24, 2026

1493
LMArena (Elo)

Gemini 3.7 Flash

Google · text · verified Sep 24, 2026

1490
LMArena (Elo)

Gemini 3.1 Pro

Google DeepMind · text · verified Sep 24, 2026

1487
LMArena (Elo)

Kimi K3

Moonshot AI · text · verified Sep 24, 2026

1485
LMArena (Elo)

GPT-5.6 Sol

OpenAI · text · verified Sep 24, 2026

1484
LMArena (Elo)

GLM-5.3

Z.ai · text · verified Sep 24, 2026

1483
LMArena (Elo)

Qwen3.8 Max

Alibaba · text · verified Sep 24, 2026

1481
LMArena (Elo)

GPT-6 Astra

OpenAI · text · verified Sep 24, 2026

1480
LMArena (Elo)

DeepSeek V4 Pro

DeepSeek · text · verified Sep 24, 2026

1457
SWE-bench Verified

Claude Opus 5

Anthropic · Verified · verified Sep 24, 2026

96.0%
GPQA

GPT-6 Astra

OpenAI · Diamond · verified Sep 24, 2026

96.0%
SWE-bench Verified

Claude Fable 5

Anthropic · Verified · verified Aug 16, 2026

95.5%
SWE-bench Verified

Claude Sonnet 5

Anthropic · Verified · verified Aug 16, 2026

95.5%
GPQA

GPT-5.6 Sol

OpenAI · Diamond · verified Sep 24, 2026

94.6%
GPQA

Gemini 3.1 Pro

Google DeepMind · Diamond · verified Sep 24, 2026

94.3%
GPQA

Kimi K3

Moonshot AI · Diamond · verified Sep 10, 2026

93.5%
GPQA

GPT-5.6 Terra

OpenAI · Diamond · verified Sep 24, 2026

92.9%
GPQA

Qwen3.8 Max

Alibaba · Diamond · verified Sep 10, 2026

92.6%
GPQA

GPT-5.6 Luna

OpenAI · Diamond · verified Sep 24, 2026

92.3%
Terminal-Bench

Kimi K3

Moonshot AI · v2.1 · verified Sep 10, 2026

88.3%
Terminal-Bench

GLM-5.3

Z.ai · v2.1 · verified Sep 10, 2026

88.2%
Terminal-Bench

DeepSeek V4 Pro

DeepSeek · v2.1 · verified Sep 10, 2026

87.9%
Terminal-Bench

DeepSeek V4 Flash

DeepSeek · v2.1 · verified Sep 10, 2026

82.7%
SWE-bench Verified

Gemini 3.1 Pro

Google DeepMind · Verified · verified Sep 24, 2026

80.6%
Artificial Analysis Coding Agent Index

GPT-5.6 Sol

OpenAI · v1.3 · verified Aug 16, 2026

80.0
SWE-bench Verified

Mistral Medium 3.5

Mistral AI · Verified · verified Aug 16, 2026

77.6%
DeepSWE

GPT-6 Astra

OpenAI · v1.1 · verified Sep 24, 2026

74.1%
Terminal-Bench

Qwen3.8-27B

Alibaba · v2.1 · verified Sep 10, 2026

73.0%
DeepSWE

GPT-5.6 Sol

OpenAI · v1.1 · verified Sep 24, 2026

72.7%
Artificial Analysis Coding Index

Claude Sonnet 5

Anthropic · current · verified Aug 16, 2026

71.3
Humanity's Last Exam

Claude Opus 5.5

Anthropic · with tools · verified Sep 24, 2026

67.7%
Terminal-Bench

Claude Opus 5.5

Anthropic · v4.0 · verified Sep 24, 2026

66.4%
DeepSWE

Grok 4.6

xAI · v1.1 · verified Sep 10, 2026

65.9%
Humanity's Last Exam

Claude Fable 5.1

Anthropic · with tools · verified Sep 24, 2026

65.6%
Humanity's Last Exam

Claude Opus 5

Anthropic · with tools · verified Sep 24, 2026

63.6%
DeepSWE

DeepSeek V4 Pro

DeepSeek · 113 tasks · verified Sep 10, 2026

62.7%
Terminal-Bench

Claude Opus 5.5

Anthropic · v4.0 · verified Sep 24, 2026

59.6%
Artificial Analysis Coding Index

Qwen3-Coder-Next

Alibaba · current · verified Aug 16, 2026

58.2
Artificial Analysis Intelligence Index

Claude Opus 5.5

Anthropic · v4.3 · verified Sep 24, 2026

58
Terminal-Bench

GPT-6 Astra

OpenAI · v4.0 · verified Sep 24, 2026

57.9%
Humanity's Last Exam

GPT-6 Astra

OpenAI · with tools · verified Sep 24, 2026

57.2%
Artificial Analysis Coding Agent Index

GPT-6 Sol

OpenAI · Codex harness · verified Sep 24, 2026

57
Terminal-Bench

Claude Fable 5.1

Anthropic · v4.0 · verified Sep 24, 2026

55.8%
Humanity's Last Exam

Gemini 3.8 Flash

Google · HLE-Verified · verified Sep 10, 2026

54.9%
DeepSWE

DeepSeek V4 Flash

DeepSeek · 113 tasks · verified Sep 10, 2026

54.4%
Artificial Analysis Intelligence Index

Claude Fable 5.1

Anthropic · v4.3 · verified Sep 24, 2026

53
Artificial Analysis Intelligence Index

GPT-6 Astra

OpenAI · v4.3 · verified Sep 24, 2026

53
Terminal-Bench

Claude Opus 5

Anthropic · v4.0 · verified Sep 24, 2026

51.8%
Artificial Analysis Intelligence Index

Claude Opus 5

Anthropic · v4.3 · verified Sep 24, 2026

51
Artificial Analysis Intelligence Index

Claude Fable 5

Anthropic · v4.3 · verified Sep 24, 2026

50
Artificial Analysis Intelligence Index

GPT-6 Sol

OpenAI · v4.3 · verified Sep 24, 2026

48
Humanity's Last Exam

GPT-6 Sol

OpenAI · AA run · verified Sep 24, 2026

48%
Artificial Analysis Intelligence Index

GPT-5.6 Sol

OpenAI · v4.3 · verified Sep 24, 2026

47
Artificial Analysis Intelligence Index

Qwen3.8 Max

Alibaba · v4.3 · verified Sep 24, 2026

45
Artificial Analysis Intelligence Index

GLM-5.3

Z.ai · v4.3 · verified Sep 24, 2026

45
Terminal-Bench

Claude Fable 5

Anthropic · v4.0 · verified Sep 24, 2026

44.5%
Terminal-Bench

GPT-6 Sol

OpenAI · v4.0 · verified Sep 24, 2026

44%
Artificial Analysis Intelligence Index

Grok 4.6

xAI · v4.3 · verified Sep 24, 2026

44
Artificial Analysis Intelligence Index

Kimi K3

Moonshot AI · v4.3 · verified Sep 24, 2026

44
Artificial Analysis Intelligence Index

GPT-5.6 Terra

OpenAI · v4.3 · verified Sep 24, 2026

42
Terminal-Bench

GLM-5.3

Z.ai · v4.0 · verified Sep 24, 2026

41.8%
Artificial Analysis Coding Agent Index

GPT-6 Luna

OpenAI · Codex harness · verified Sep 24, 2026

41
Artificial Analysis Intelligence Index

Gemini 3.8 Flash

Google · v4.3 · verified Sep 24, 2026

41
Artificial Analysis Intelligence Index

Gemini 3.7 Flash

Google · v4.3 · verified Sep 24, 2026

40
Artificial Analysis Intelligence Index

GLM-5.2

Z.ai · v4.3 · verified Sep 24, 2026

39
Artificial Analysis Intelligence Index

DeepSeek V4 Flash

DeepSeek · v4.3 · verified Sep 24, 2026

39
Artificial Analysis Intelligence Index

Claude Sonnet 5

Anthropic · v4.3 · verified Sep 24, 2026

38
Terminal-Bench

GPT-5.6 Sol

OpenAI · v4.0 · verified Sep 24, 2026

37.3%
Artificial Analysis Intelligence Index

GPT-6 Luna

OpenAI · v4.3 · verified Sep 24, 2026

37
Artificial Analysis Intelligence Index

GPT-5.6 Luna

OpenAI · v4.3 · verified Sep 24, 2026

37
Artificial Analysis Intelligence Index

DeepSeek V4 Pro

DeepSeek · v4.3 · verified Sep 24, 2026

36
Artificial Analysis Intelligence Index

Qwen3.8-27B

Alibaba · v4.3 · verified Sep 24, 2026

34
Artificial Analysis Intelligence Index

Gemini 3.1 Pro

Google DeepMind · v4.3 · verified Sep 24, 2026

30
Artificial Analysis Intelligence Index

MiniMax M3

MiniMax · v4.3 · verified Sep 24, 2026

29
Terminal-Bench

GPT-5.6 Terra

OpenAI · v4.0 · verified Sep 24, 2026

21.5%
Terminal-Bench

Grok 4.6

xAI · v4.0 · verified Sep 24, 2026

20.3%
Artificial Analysis Intelligence Index

K-EXAONE 2.0

LG AI Research · v4.3 · verified Sep 24, 2026

20
Terminal-Bench

Gemini 3.8 Flash

Google · v4.0 · verified Sep 24, 2026

19.1%
Artificial Analysis Intelligence Index

Gemma 4

Google · v4.3 · verified Sep 24, 2026

19
Terminal-Bench

GPT-5.6 Luna

OpenAI · v4.0 · verified Sep 24, 2026

17.3%
Artificial Analysis Intelligence Index

Mistral Medium 3.5

Mistral AI · v4.3 · verified Sep 24, 2026

14
Terminal-Bench

GPT-6 Luna

OpenAI · v4.0 · verified Sep 24, 2026

13%
Terminal-Bench

Kimi K3

Moonshot AI · v4.0 · verified Sep 24, 2026

12.6%
Terminal-Bench

Claude Sonnet 5

Anthropic · v4.0 · verified Sep 24, 2026

12.4%
Terminal-Bench

Gemini 3.7 Flash

Google · v4.0 · verified Sep 24, 2026

11.2%
Artificial Analysis Intelligence Index

Llama 4 Maverick

Meta · v4.3 · verified Sep 24, 2026

10
Artificial Analysis Intelligence Index

K-EXAONE 2.0

LG AI Research · current · verified Sep 10, 2026

listed
ProprietaryOpen WeightsOpen Source90 sourced results

How the Lab scores

What counts

Cited benchmark results from the Models Hub — each with a source URL, benchmark version, and verification date.

What doesn't

No public ratings, no crowd-sourced scores, no fabricated numbers. Benchmarks are never averaged into one number.

Last check

Each evidence record carries its own verification date. Freshness degrades automatically when nothing is re-verified.

Read the methodology

Evidence table

90 records · filter above
VersionTypeSource
1506Claude Fable 5LMArena (Elo)textProprietarySep 24, 2026link
1498Claude Fable 5.1LMArena (Elo)textProprietarySep 24, 2026link
1493Claude Opus 5LMArena (Elo)textProprietarySep 24, 2026link
1493Gemini 3.8 FlashLMArena (Elo)textProprietarySep 24, 2026link
1490Gemini 3.7 FlashLMArena (Elo)textProprietarySep 24, 2026link
1487Gemini 3.1 ProLMArena (Elo)textProprietarySep 24, 2026link
1485Kimi K3LMArena (Elo)textOpen WeightsSep 24, 2026link
1484GPT-5.6 SolLMArena (Elo)textProprietarySep 24, 2026link
1483GLM-5.3LMArena (Elo)textOpen WeightsSep 24, 2026link
1481Qwen3.8 MaxLMArena (Elo)textOpen WeightsSep 24, 2026link
1480GPT-6 AstraLMArena (Elo)textProprietarySep 24, 2026link
1457DeepSeek V4 ProLMArena (Elo)textOpen WeightsSep 24, 2026link
96.0%Claude Opus 5SWE-bench VerifiedVerifiedProprietarySep 24, 2026link
96.0%GPT-6 AstraGPQADiamondProprietarySep 24, 2026link
95.5%Claude Fable 5SWE-bench VerifiedVerifiedProprietaryAug 16, 2026link
95.5%Claude Sonnet 5SWE-bench VerifiedVerifiedProprietaryAug 16, 2026link
94.6%GPT-5.6 SolGPQADiamondProprietarySep 24, 2026link
94.3%Gemini 3.1 ProGPQADiamondProprietarySep 24, 2026link
93.5%Kimi K3GPQADiamondOpen WeightsSep 10, 2026link
92.9%GPT-5.6 TerraGPQADiamondProprietarySep 24, 2026link
92.6%Qwen3.8 MaxGPQADiamondOpen WeightsSep 10, 2026link
92.3%GPT-5.6 LunaGPQADiamondProprietarySep 24, 2026link
88.3%Kimi K3Terminal-Benchv2.1Open WeightsSep 10, 2026link
88.2%GLM-5.3Terminal-Benchv2.1Open WeightsSep 10, 2026link
87.9%DeepSeek V4 ProTerminal-Benchv2.1Open WeightsSep 10, 2026link
82.7%DeepSeek V4 FlashTerminal-Benchv2.1Open WeightsSep 10, 2026link
80.6%Gemini 3.1 ProSWE-bench VerifiedVerifiedProprietarySep 24, 2026link
80.0GPT-5.6 SolArtificial Analysis Coding Agent Indexv1.3ProprietaryAug 16, 2026link
77.6%Mistral Medium 3.5SWE-bench VerifiedVerifiedOpen WeightsAug 16, 2026link
74.1%GPT-6 AstraDeepSWEv1.1ProprietarySep 24, 2026link
Showing top 30 of 90 records — full table lives in the Models Hub.
The Honestack rule

Source-led listings. Every claim links to a source — or says it doesn't have one yet.

No public ratings, no crowd-sourced scoring, no fabricated benchmarks. Each listing links to an official source and records when its details were last checked. Listings without sources are labeled honestly as reference entries.

Read the methodology
Tools on file
222
Source-linked
222/222
Documented workflows
94