Head to head
Compare any two models
Accuracy, cost, and latency on the getEvals Index. The better result on each row is highlighted. Use it to pick a model for a task — not to crown a winner.
Model A
Claude Opus 5 · Anthropic Claude Fable 5 · Anthropic GPT-5.6 Sol · OpenAI Claude Opus 4.8 · Anthropic GPT-5.6 Luna · OpenAI Claude Sonnet 5 · Anthropic Grok 4.6 · SpaceXAI Kimi K3 · Moonshot GPT 5.5 · OpenAI Muse Spark 1.2 · Meta GPT-5.6 Terra · OpenAI Claude Opus 4.7 · Anthropic Gemini 3.6 Flash · Google Muse Spark 1.1 · Meta DeepSeek V4 Flash 0731 · DeepSeek GLM 5.2 · zAI Gemini 3.5 Flash · Google DeepSeek V4 Pro 0813 · DeepSeek Qwen 3.8 Max · Alibaba Grok 4.5 · SpaceXAI Claude Sonnet 4.6 · Anthropic Qwen 3.7 Max · Alibaba Kimi K2.6 · Moonshot DeepSeek V4 · DeepSeek MiniMax-M3 · MiniMax Inkling · Thinking Nemotron 3 Ultra · NVIDIA Gemma 4 31B IT · Google Mistral Medium 3.5 · Mistral
Model B
Claude Opus 5 · Anthropic Claude Fable 5 · Anthropic GPT-5.6 Sol · OpenAI Claude Opus 4.8 · Anthropic GPT-5.6 Luna · OpenAI Claude Sonnet 5 · Anthropic Grok 4.6 · SpaceXAI Kimi K3 · Moonshot GPT 5.5 · OpenAI Muse Spark 1.2 · Meta GPT-5.6 Terra · OpenAI Claude Opus 4.7 · Anthropic Gemini 3.6 Flash · Google Muse Spark 1.1 · Meta DeepSeek V4 Flash 0731 · DeepSeek GLM 5.2 · zAI Gemini 3.5 Flash · Google DeepSeek V4 Pro 0813 · DeepSeek Qwen 3.8 Max · Alibaba Grok 4.5 · SpaceXAI Claude Sonnet 4.6 · Anthropic Qwen 3.7 Max · Alibaba Kimi K2.6 · Moonshot DeepSeek V4 · DeepSeek MiniMax-M3 · MiniMax Inkling · Thinking Nemotron 3 Ultra · NVIDIA Gemma 4 31B IT · Google Mistral Medium 3.5 · Mistral
getEvals Index
Cost / test
Latency