Score against what it costs to get it
Drawn the way the official board draws it: colour is the lab, and each line joins one model’s effort levels from low to max. The models you want are top right: higher up solves more of the 113 tasks, further right costs less per task. Shape and line style say who measured it. Hover a point for its source; hover a row in the table to pick out its line.
Every reading, ranked
The official leaderboard’s form: bars on a 0–80% scale coloured by lab, with the 95% interval as a whisker where the source gives one. The number after each score is its source in the list below; a cost marked est. is estimated from list prices. Select a heading to sort.
| Bar | |||||||
|---|---|---|---|---|---|---|---|
| Gemini 4 ArgonGoogle · mini-swe-agent · LAB | LABGoogle DeepMind | 77.9%28 | not given | — | — | 24 Sept 2026 | |
| Muse Spark 1.3 [max]Meta · mini-swe-agent · LAB | LABMeta | 75.4%17 | not given | — | — | 2 Sept 2026 | |
| GPT-6.1 Sol [high]OpenAI · lab-internal · LAB | LABOpenAI | 75.2%30 | $0.65 | — | — | 29 Sept 2026 | |
| Ember-1 [thinking]Fireworks AI · mini-swe-agent · LAB | LABFireworks AI | 75.2%27 | $3.62 | — | — | 23 Sept 2026 | |
| Claude Opus 5.5 [max]Anthropic · lab-internal · LAB | LABAnthropic | 74.2%25 | not given | — | — | 22 Sept 2026 | |
| DeepSeek V4.1 Flash [max]DeepSeek · mini-swe-agent · LAB | LABDeepSeek | 74.2%20 | not given | — | — | 10 Sept 2026 | |
| GPT-6 Astra [xhigh]OpenAI · mini-swe-agent · DC | DCDatacurve | 74.1% ±2.871 | $4.43 | 30k | 29 | 3 Sept 2026 | |
| Gemini 3.8 Flash [high]Google · mini-swe-agent · DC | DCDatacurve | 73.8% ±1.421 | $2.36 | 143k | 166 | 1 Sept 2026 | |
| Gemini 3.8 Flash [high]Google · mini-swe-agent · LAB | LABGoogle DeepMind | 73.7%21 | $2.36 | 143k | — | 17 Sept 2026 | |
| Claude Opus 5 [max]Anthropic · mini-swe-agent · DC | DCDatacurve | 73.7% ±3.871 | $12 | 118k | 99 | 25 Jul 2026 | |
| GPT-6 Astra [high]OpenAI · mini-swe-agent · DC | DCDatacurve | 73.2% ±3.421 | $3.92 | 27k | 27 | 3 Sept 2026 | |
| GPT-6 Astra [max]OpenAI · mini-swe-agent · DC | DCDatacurve | 73.2% ±0.831 | $7.50 | 61k | 29 | 3 Sept 2026 | |
| Claude Opus 5 [xhigh]Anthropic · mini-swe-agent · DC | DCDatacurve | 73.2% ±3.061 | $9.07 | 92k | 89 | 25 Jul 2026 | |
| GPT-6.1 Sol [medium]OpenAI · lab-internal · LAB | LABOpenAI | 73%30 | $0.42 | — | — | 29 Sept 2026 | |
| Grok 4.7 [xhigh]xAI · Grok Build · IND | INDArtificial Analysis | 73%35 | not given | — | — | 21 Sept 2026 | |
| SWE-2Cognition · Devin CLI · LAB | LABCognition | 73%19 | not given | — | — | 10 Sept 2026 | |
| Claude Opus 5 [high]Anthropic · mini-swe-agent · DC | DCDatacurve | 72.8% ±1.951 | $6.08 | 64k | 73 | 25 Jul 2026 | |
| GPT-6 Astra [medium]OpenAI · mini-swe-agent · DC | DCDatacurve | 72.8% ±2.591 | $3.08 | 20k | 26 | 3 Sept 2026 | |
| GPT-5.6 Sol [max]OpenAI · mini-swe-agent · DC | DCDatacurve | 72.7% ±2.831 | $6.46 | 60k | 61 | 10 Jul 2026 | |
| DeepSeek V4.1 Flash [max]DeepSeek · lab-internal · LAB | LABDeepSeek | 72.6%20 | not given | — | — | 10 Sept 2026 | |
| Claude Opus 5.5 [max]Anthropic · mini-swe-agent · IND | INDMercor | 72.3% ±7.336 | not given | — | — | not given | |
| GPT-6.1 Sol [max]OpenAI · mini-swe-agent · IND | INDMercor | 72.3% ±7.436 | not given | — | — | not given | |
| GPT-6 Astra [max]OpenAI · mini-swe-agent · IND | INDMercor | 72% ±7.536 | not given | — | — | not given | |
| GPT-6.1 Sol [max]OpenAI · lab-internal · LAB | LABOpenAI | 71.9%30 | $1.57 | — | — | 29 Sept 2026 | |
| GPT-6.1 Sol [xhigh]OpenAI · lab-internal · LAB | LABOpenAI | 71.9%30 | $0.79 | — | — | 29 Sept 2026 | |
| MiMo-V2.6-Pro [max]Xiaomi · lab-internal · LAB | LABXiaomi | 71.9%24 | not given | — | — | 21 Sept 2026 | |
| DeepSeek V4.1 Flash [max]DeepSeek · mini-swe-agent · IND | INDMercor | 71.7% ±6.336 | not given | — | — | not given | |
| Claude Opus 5 [max]Anthropic · mini-swe-agent · IND | INDMercor | 71.4% ±6.836 | not given | — | — | not given | |
| Gemini 3.8 Flash [high]Google · mini-swe-agent · IND | INDMercor | 71.4% ±6.836 | not given | — | — | not given | |
| Gemini 3.8 Flash [medium]Google · mini-swe-agent · DC | DCDatacurve | 71% ±2.281 | $1.97 | 125k | 147 | 1 Sept 2026 | |
| Claude Sonnet 5.5 [max]Anthropic · lab-internal · LAB | LABAnthropic | 71%29 | not given | — | — | 28 Sept 2026 | |
| Grok 4.7 [high]xAI · mini-swe-agent · LAB | LABxAI | 71%23 | not given | — | — | 21 Sept 2026 | |
| GPT-5.6 Sol [xhigh]OpenAI · mini-swe-agent · DC | DCDatacurve | 70.7% ±0.821 | $3.60 | 41k | 44 | 10 Jul 2026 | |
| Claude Sonnet 5.5 [max]Anthropic · mini-swe-agent · IND | INDMercor | 70.5% ±7.736 | not given | — | — | not given | |
| DeepSeek V4.1 Flash [max]DeepSeek · lab-internal · LAB | LABDeepSeek | 70.5%20 | not given | — | — | 10 Sept 2026 | |
| GLM-5.3 [max]Z.ai · mini-swe-agent · IND | INDMercor | 70.5% ±6.336 | not given | — | — | not given | |
| GPT-5.6 Sol [max]OpenAI · mini-swe-agent · IND | INDMercor | 70.5% ±6.936 | not given | — | — | not given | |
| GPT-6 Sol [max]OpenAI · mini-swe-agent · IND | INDMercor | 70.5% ±7.236 | not given | — | — | not given | |
| Claude Opus 5 [xhigh]Anthropic · mini-swe-agent · IND | INDMercor | 70.2% ±6.936 | not given | — | — | not given | |
| GPT-5.6 Terra [max]OpenAI · mini-swe-agent · IND | INDMercor | 70.2% ±7.236 | not given | — | — | not given | |
| Claude Fable 5 [xhigh]Anthropic · mini-swe-agent · DC | DCDatacurve | 69.9% ±3.241 | $13 | 80k | 68 | 15 Jun 2026 | |
| DeepSeek V4.1 Flash [max]DeepSeek · Claude Code · LAB | LABDeepSeek | 69.8%20 | not given | — | — | 10 Sept 2026 | |
| Claude Fable 5 [max]Anthropic · mini-swe-agent · DC | DCDatacurve | 69.7% ±4.031 | $22 | 119k | 88 | 15 Jun 2026 | |
| GPT-5.6 Terra [max]OpenAI · mini-swe-agent · DC | DCDatacurve | 69.6% ±2.561 | $3.96 | 72k | 76 | 10 Jul 2026 | |
| GPT-5.6 Sol [high]OpenAI · mini-swe-agent · DC | DCDatacurve | 69.4% ±1.431 | $2.66 | 28k | 37 | 10 Jul 2026 | |
| GLM-5.3 [max]Z.ai · mini-swe-agent · DC | DCDatacurve | 69% ±3.021 | $3.99 | 80k | 125 | 20 Aug 2026 | |
| Claude Opus 5 [medium]Anthropic · mini-swe-agent · DC | DCDatacurve | 68.9% ±1.171 | $3.29 | 37k | 52 | 25 Jul 2026 | |
| GPT-6 Sol [max]OpenAI · lab-internal · LAB | LABOpenAI | 68.8%26 | $2.74 | — | — | 22 Sept 2026 | |
| Claude Opus 5 [max]Anthropic · lab-internal · LAB | LABAnthropic | 68.8%9 | not given | — | — | 24 Jul 2026 | |
| Claude Fable 5 [high]Anthropic · mini-swe-agent · DC | DCDatacurve | 68.6% ±1.121 | $9.18 | 57k | 59 | 15 Jun 2026 | |
| Kimi K3 [max]Moonshot AI · mini-swe-agent · DC | DCDatacurve | 68.5% ±4.541 | $4.65 | 82k | 98 | 18 Jul 2026 | |
| MiMo-V2.6-Flash [max]Xiaomi · lab-internal · LAB | LABXiaomi | 67.9%24 | not given | — | — | 21 Sept 2026 | |
| Step 5 Preview [high]StepFun · mini-swe-agent · LAB | LABStepFun | 67.7%22 | not given | — | — | 20 Sept 2026 | |
| DeepSeek V4.1 Flash [max]DeepSeek · lab-internal · LAB | LABDeepSeek | 67.6%20 | not given | — | — | 10 Sept 2026 | |
| Kimi K3 [max]Moonshot AI · lab-internal · LAB | LABMoonshot AI | 67.5%8 | not given | — | — | 23 Jul 2026 | |
| Grok 4.6 [medium]xAI · mini-swe-agent · DC | DCDatacurve | 67.5% ±2.281 | $3.45 | 50k | 70 | 12 Aug 2026 | |
| Claude Fable 5.1 [max]Anthropic · lab-internal · LAB | LABAnthropic | 67.4%16 | not given | — | — | 1 Sept 2026 | |
| Claude Fable 5 [max]Anthropic · mini-swe-agent · IND | INDMercor | 67.3% ±7.236 | not given | — | — | not given | |
| Claude Fable 5.1 [high]Anthropic · mini-swe-agent · IND | INDMercor | 67.3% ±6.936 | not given | — | — | not given | |
| GPT-5.6 Luna [max]OpenAI · mini-swe-agent · DC | DCDatacurve | 67.2% ±3.991 | $0.61 | 73k | 102 | 10 Jul 2026 | |
| GPT-5.5 [xhigh]OpenAI · mini-swe-agent · DC | DCDatacurve | 67% ±6.471 | $7.23 | 46k | 82 | 15 Jun 2026 | |
| GPT-6 Astra [low]OpenAI · mini-swe-agent · DC | DCDatacurve | 67% ±1.31 | $1.60 | 11k | 20 | 3 Sept 2026 | |
| GPT-6 Luna [max]OpenAI · mini-swe-agent · IND | INDMercor | 67% ±6.836 | not given | — | — | not given | |
| GLM-5.3 [max]Z.ai · mini-swe-agent · LAB | LABZ.ai | 66.9%14 | not given | — | — | 25 Aug 2026 | |
| Grok 4.6 [xhigh]xAI · mini-swe-agent · DC | DCDatacurve | 66.7% ±2.181 | $5.50 | 71k | 87 | 12 Aug 2026 | |
| GPT-6 Luna [max]OpenAI · lab-internal · LAB | LABOpenAI | 66.6%26 | $0.22 | — | — | 22 Sept 2026 | |
| GPT-6 Sol [xhigh]OpenAI · lab-internal · LAB | LABOpenAI | 66.6%26 | $1.00 | — | — | 22 Sept 2026 | |
| Kimi K3 [max]Moonshot AI · mini-swe-agent · IND | INDFireworks AI | 66.4%27 | $4.74 | — | — | 23 Sept 2026 | |
| DeepSeek V4.1 Flash [max]DeepSeek · Pi · LAB | LABDeepSeek | 66.2%20 | not given | — | — | 10 Sept 2026 | |
| GPT-5.5 [xhigh]OpenAI · mini-swe-agent · IND | INDMercor | 66.1% ±6.936 | not given | — | — | not given | |
| Grok 4.6 [high]xAI · mini-swe-agent · LAB | LABxAI | 65.9%12 | not given | — | — | 18 Aug 2026 | |
| DeepSeek V4.1 Flash [max]DeepSeek · Codex · LAB | LABDeepSeek | 65.6%20 | not given | — | — | 10 Sept 2026 | |
| Nex-N2.5-MaxNex-AGI · NexAU · LAB | LABNex-AGI | 65.6%18 | not given | — | — | 8 Sept 2026 | |
| DeepSeek V4.1 Flash [max]DeepSeek · OpenCode · LAB | LABDeepSeek | 65.5%20 | not given | — | — | 10 Sept 2026 | |
| Gemini 3.7 Flash [high]Google · mini-swe-agent · IND | INDMercor | 65.5% ±7.136 | not given | — | — | not given | |
| Gemini 3.7 Flash [medium]Google · mini-swe-agent · DC | DCDatacurve | 65.5% ±3.091 | $2.03 | 94k | 117 | 13 Aug 2026 | |
| Claude Fable 5 [medium]Anthropic · mini-swe-agent · DC | DCDatacurve | 65.4% ±4.421 | $6.09 | 40k | 48 | 15 Jun 2026 | |
| Gemini 3.7 Flash [high]Google · mini-swe-agent · DC | DCDatacurve | 65.3% ±1.791 | $2.18 | 107k | 125 | 13 Aug 2026 | |
| GPT-6 Sol [high]OpenAI · lab-internal · LAB | LABOpenAI | 65.3%26 | $0.64 | — | — | 22 Sept 2026 | |
| Grok 4.6 [high]xAI · mini-swe-agent · DC | DCDatacurve | 65.2% ±1.531 | $4.38 | 61k | 79 | 12 Aug 2026 | |
| Grok 4.6 [high]xAI · Grok Build · IND | INDArtificial Analysis | 65%35 | not given | — | — | 21 Sept 2026 | |
| Claude Fable 5.1 [max]Anthropic · mini-swe-agent · IND | INDMercor | 64.9% ±7.236 | not given | — | — | not given | |
| GPT-5.5 [high]OpenAI · mini-swe-agent · DC | DCDatacurve | 64.4% ±3.121 | $5.10 | 31k | 62 | 15 Jun 2026 | |
| GPT-6.1 Sol [low]OpenAI · lab-internal · LAB | LABOpenAI | 64.4%30 | $0.17 | — | — | 29 Sept 2026 | |
| GPT-5.6 Luna [max]OpenAI · mini-swe-agent · IND | INDMercor | 64.3% ±7.436 | not given | — | — | not given | |
| Hy4 PreviewTencent · lab-internal · LAB | LABTencent | 64.3%15 | not given | — | — | 27 Aug 2026 | |
| Grok 4.6 [high]xAI · mini-swe-agent · IND | INDMercor | 63.7% ±7.436 | not given | — | — | not given | |
| GLM-5.3 Flash [max]Z.ai · mini-swe-agent · DC | DCDatacurve | 63.4% ±4.381 | $0.24 | 73k | 123 | 26 Aug 2026 | |
| DeepSeek V4 Pro [max]DeepSeek · mini-swe-agent · DC | DCDatacurve | 62.8% ±6.331 | $1.67 | 106k | 155 | 12 Aug 2026 | |
| Kimi K3 [high]Moonshot AI · mini-swe-agent · IND | INDFireworks AI | 62.8%27 | not given | — | — | 23 Sept 2026 | |
| DeepSeek V4 Pro [max]DeepSeek · mini-swe-agent · LAB | LABDeepSeek | 62.7%20 | not given | — | — | 10 Sept 2026 | |
| GPT-5.6 Luna [max]OpenAI · lab-internal · LAB | LABOpenAI | 62.2%26 | $0.53 | — | — | 22 Sept 2026 | |
| Mistral Large 4 [thinking]Mistral · partner eval (Artificial Analysis, Surge AI) · LAB | LABMistral AI | 61.7%32 | not given | — | — | 6 Oct 2026 | |
| GPT-6 Luna [xhigh]OpenAI · lab-internal · LAB | LABOpenAI | 61.3%26 | $0.11 | — | — | 22 Sept 2026 | |
| GPT-5.6 Sol [medium]OpenAI · mini-swe-agent · DC | DCDatacurve | 61.1% ±1.581 | $1.42 | 18k | 31 | 10 Jul 2026 | |
| GPT-5.6 Terra [xhigh]OpenAI · mini-swe-agent · DC | DCDatacurve | 60.2% ±2.121 | $1.70 | 40k | 43 | 10 Jul 2026 | |
| Claude Haiku 5.5 [max]Anthropic · mini-swe-agent · IND | INDMercor | 59.9% ±7.736 | not given | — | — | not given | |
| Claude Opus 4.8 [max]Anthropic · mini-swe-agent · IND | INDMercor | 59.6% ±7.136 | not given | — | — | not given | |
| Claude Fable 5 [low]Anthropic · mini-swe-agent · DC | DCDatacurve | 59.6% ±2.791 | $3.76 | 25k | 38 | 15 Jun 2026 | |
| GPT-6 Luna [high]OpenAI · lab-internal · LAB | LABOpenAI | 59.3%26 | $0.084 | — | — | 22 Sept 2026 | |
| Claude Opus 4.8 [max]Anthropic · mini-swe-agent · DC | DCDatacurve | 59% ±1.761 | $13 | 135k | 120 | 15 Jun 2026 | |
| Qwen3.8-Flash-NextAlibaba · mini-swe-agent · LAB | LABAlibaba (Qwen) | 58.7%13 | not given | — | — | 24 Aug 2026 | |
| Claude Opus 5 [low]Anthropic · mini-swe-agent · DC | DCDatacurve | 58.1% ±2.331 | $1.66 | 20k | 36 | 25 Jul 2026 | |
| Qwen3.8 Max [xhigh]Alibaba · mini-swe-agent · DC | DCDatacurve | 57.5% ±2.661 | $3.73 | 95k | 111 | 4 Aug 2026 | |
| GPT-5.6 Luna [xhigh]OpenAI · mini-swe-agent · DC | DCDatacurve | 56.9% ±2.171 | $0.31 | 45k | 71 | 10 Jul 2026 | |
| GPT-6 Sol [medium]OpenAI · lab-internal · LAB | LABOpenAI | 56.6%26 | $0.38 | — | — | 22 Sept 2026 | |
| DeepSeek V4 Flash [max]DeepSeek · mini-swe-agent · IND | INDMercor | 56.6% ±8.836 | not given | — | — | not given | |
| Qwen3.8 MaxAlibaba · Claude Code · LAB | LABAlibaba (Qwen) | 56.6%11 | not given | — | — | 8 Aug 2026 | |
| DeepSeek V4 Pro 0813 [max]DeepSeek · mini-swe-agent · IND | INDMercor | 56.3% ±7.236 | not given | — | — | not given | |
| GPT-5.6 Luna [xhigh]OpenAI · lab-internal · LAB | LABOpenAI | 56.2%26 | $0.27 | — | — | 22 Sept 2026 | |
| Kimi K3 [low]Moonshot AI · mini-swe-agent · IND | INDFireworks AI | 55.8%27 | not given | — | — | 23 Sept 2026 | |
| Nex-N2.5-ProNex-AGI · NexAU · LAB | LABNex-AGI | 55.8%18 | not given | — | — | 8 Sept 2026 | |
| Muse Spark 1.2 [xhigh]Meta · mini-swe-agent · DC | DCDatacurve | 54.9% ±2.121 | $3.70 | 99k | 101 | 7 Aug 2026 | |
| DeepSeek V4 Flash [max]DeepSeek · mini-swe-agent · LAB | LABDeepSeek | 54.4%20 | not given | — | — | 10 Sept 2026 | |
| Claude Opus 4.8 [xhigh]Anthropic · mini-swe-agent · DC | DCDatacurve | 54.4% ±3.711 | $8.01 | 86k | 95 | 15 Jun 2026 | |
| Claude Opus 4.7 [max]Anthropic · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 54.2% ±4.722 | $18 | 103k | 203 | 15 Jun 2026 | |
| GPT-5.5 [medium]OpenAI · mini-swe-agent · DC | DCDatacurve | 54% ±2.551 | $2.75 | 20k | 46 | 15 Jun 2026 | |
| Claude Sonnet 5 [max]Anthropic · mini-swe-agent · DC | DCDatacurve | 53.9% ±4.241 | $26 | 214k | 269 | 2 Jul 2026 | |
| Gemini 3.7 Flash [low]Google · mini-swe-agent · DC | DCDatacurve | 53.8% ±2.591 | $1.83 | 73k | 130 | 13 Aug 2026 | |
| GPT-5.6 Terra [high]OpenAI · mini-swe-agent · DC | DCDatacurve | 53.8% ±4.331 | $0.91 | 22k | 34 | 10 Jul 2026 | |
| Grok 4.5 [high]xAI · mini-swe-agent · DC | DCDatacurve | 53.8% ±2.281 | $2.42 | 36k | 61 | 16 Jul 2026 | |
| DeepSeek V4 Flash [max]DeepSeek · mini-swe-agent · DC | DCDatacurve | 53.3% ±3.571 | $0.46 | 108k | 153 | 6 Aug 2026 | |
| Muse Spark 1.1 [xhigh]Meta · mini-swe-agent · DC | DCDatacurve | 53.3% ±3.041 | $2.36 | 74k | 96 | 14 Jul 2026 | |
| Qwen3.8 Max [xhigh]Alibaba · mini-swe-agent · IND | INDMercor | 53.1% ±8.836 | not given | — | — | not given | |
| Claude Opus 4.8 [high]Anthropic · mini-swe-agent · DC | DCDatacurve | 51.8% ±4.561 | $4.28 | 50k | 73 | 15 Jun 2026 | |
| GPT-5.4 [xhigh]OpenAI · mini-swe-agent · DC | DCDatacurve | 51.8% ±1.51 | $5.65 | 71k | 71 | 15 Jun 2026 | |
| Claude Sonnet 5 [xhigh]Anthropic · mini-swe-agent · DC | DCDatacurve | 49.7% ±3.451 | $12 | 121k | 186 | 2 Jul 2026 | |
| Claude Sonnet 5 [max]Anthropic · mini-swe-agent · IND | INDMercor | 49.6% ±7.236 | not given | — | — | not given | |
| Claude Opus 4.7 [max]Anthropic · mini-swe-agent · IND | INDMercor | 49.3% ±6.936 | not given | — | — | not given | |
| Gemini 3.6 Flash [high]Google · mini-swe-agent · IND | INDMercor | 49.3% ±7.136 | not given | — | — | not given | |
| Grok 4.5 [high]xAI · mini-swe-agent · IND | INDMercor | 49.3% ±7.436 | not given | — | — | not given | |
| GPT-5.4 [xhigh]OpenAI · mini-swe-agent · IND | INDMercor | 49% ±6.836 | not given | — | — | not given | |
| Claude Opus 4.8 [medium]Anthropic · mini-swe-agent · DC | DCDatacurve | 48.7% ±2.241 | $3.44 | 41k | 66 | 15 Jun 2026 | |
| Claude Sonnet 5 [high]Anthropic · mini-swe-agent · DC | DCDatacurve | 48.2% ±4.511 | $7.43 | 87k | 147 | 2 Jul 2026 | |
| Gemini 3.6 Flash [high]Google · mini-swe-agent · DC | DCDatacurve | 46.7% ±3.71 | $2.21 | 96k | 117 | 22 Jul 2026 | |
| GLM-5.2 [max]Z.ai · mini-swe-agent · LAB | LABZ.ai | 46.2%6 | not given | — | — | 16 Jun 2026 | |
| GPT-5.6 Sol [low]OpenAI · mini-swe-agent · DC | DCDatacurve | 45.4% ±2.391 | $0.82 | 11k | 23 | 10 Jul 2026 | |
| Claude Opus 4.7 [xhigh]Anthropic · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 44.7% ±2.882 | $8.58 | 66k | 126 | 15 Jun 2026 | |
| GPT-6 Luna [medium]OpenAI · lab-internal · LAB | LABOpenAI | 44.5%26 | $0.052 | — | — | 22 Sept 2026 | |
| BeamReflection AI · LAB | LABReflection AI | 44.4%31 | not given | — | — | 5 Oct 2026 | |
| GPT-5.6 Luna [high]OpenAI · mini-swe-agent · DC | DCDatacurve | 44.3% ±2.921 | $0.16 | 26k | 49 | 10 Jul 2026 | |
| GLM-5.2 [max]Z.ai · mini-swe-agent · DC | DCDatacurve | 43.8% ±1.731 | $3.92 | 78k | 129 | 21 Jun 2026 | |
| GLM-5.3 Flash [max]Z.ai · mini-swe-agent · IND | INDMercor | 42.5% ±6.536 | not given | — | — | not given | |
| GPT-5.6 Luna [high]OpenAI · lab-internal · LAB | LABOpenAI | 42.4%26 | $0.13 | — | — | 22 Sept 2026 | |
| Qwen3.8-27BAlibaba · Claude Code · LAB | LABAlibaba (Qwen) | 42.2%10 | not given | — | — | 5 Aug 2026 | |
| Grok 4.6 [low]xAI · mini-swe-agent · DC | DCDatacurve | 41.6% ±2.321 | $1.04 | 16k | 44 | 12 Aug 2026 | |
| Claude Opus 4.8 [low]Anthropic · mini-swe-agent · DC | DCDatacurve | 40.8% ±1.461 | $2.29 | 29k | 54 | 15 Jun 2026 | |
| Laguna S 2.1 [max]Poolside · pool · LAB | LABPoolside | 40.4%7 | ≈$0.022–0.045 est.40 | 249k total | — | 21 Jul 2026 | |
| Muse Spark 1.1 [xhigh]Meta · mini-swe-agent · IND | INDMercor | 40.4% ±6.536 | not given | — | — | not given | |
| Claude Opus 4.7 [high]Anthropic · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 40.3% ±3.52 | $4.83 | 46k | 87 | 15 Jun 2026 | |
| Claude Sonnet 5 [medium]Anthropic · mini-swe-agent · DC | DCDatacurve | 39.8% ±3.131 | $4.08 | 57k | 108 | 2 Jul 2026 | |
| GLM-5.2 [max]Z.ai · mini-swe-agent · IND | INDMercor | 38.9% ±6.536 | not given | — | — | not given | |
| SWE-1.7Cognition · Devin CLI · LAB | LABCognition | 37.7%19 | not given | — | — | 10 Sept 2026 | |
| GPT-6 Sol [low]OpenAI · lab-internal · LAB | LABOpenAI | 37.2%26 | $0.16 | — | — | 22 Sept 2026 | |
| Gemini 3.5 Flash [high]Google · mini-swe-agent · IND | INDMercor | 36.6% ±6.936 | not given | — | — | not given | |
| GLM-5.2 [high]Z.ai · mini-swe-agent · DC | DCDatacurve | 36.3% ±4.751 | $2.84 | 54k | 122 | 21 Jun 2026 | |
| Nex-N2.5-MiniNex-AGI · NexAU · LAB | LABNex-AGI | 36.1%18 | not given | — | — | 8 Sept 2026 | |
| Gemini 3.5 Flash [high]Google · mini-swe-agent · DC | DCDatacurve | 36.1% ±3.971 | $3.45 | 76k | 105 | 15 Jun 2026 | |
| GPT-5.6 Terra [medium]OpenAI · mini-swe-agent · DC | DCDatacurve | 35.1% ±3.381 | $0.47 | 12k | 25 | 10 Jul 2026 | |
| Claude Sonnet 4.6 [high]Anthropic · mini-swe-agent · IND | INDMercor | 34.2% ±6.336 | not given | — | — | not given | |
| Grok 4.7 [xhigh]xAI · mini-swe-agent · IND | INDMercor | 33.3% ±636 | not given | — | — | not given | |
| Claude Opus 4.6 [max]Anthropic · mini-swe-agent · IND | INDMercor | 31.9% ±6.336 | not given | — | — | not given | |
| Claude Opus 4.7 [medium]Anthropic · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 31.6% ±4.032 | $2.34 | 27k | 55 | 15 Jun 2026 | |
| Kimi K2.7 CodeMoonshot AI · mini-swe-agent · DC | DCDatacurve | 30.5% ±0.51 | $2.82 | 59k | 149 | 15 Jun 2026 | |
| Claude Sonnet 5 [low]Anthropic · mini-swe-agent · DC | DCDatacurve | 30.5% ±1.131 | $2.19 | 36k | 77 | 2 Jul 2026 | |
| Kimi K2.7 Code [high]Moonshot AI · mini-swe-agent · IND | INDMercor | 30.1% ±8.436 | not given | — | — | not given | |
| Claude Sonnet 4.6 [high]Anthropic · mini-swe-agent · DC | DCDatacurve | 29.9% ±4.091 | $5.52 | 76k | 134 | 15 Jun 2026 | |
| Gemini 3.5 Flash [medium]Google · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 28.3% ±3.542 | $7.42 | 189k | 75 | 15 Jun 2026 | |
| Hy3Tencent · lab-internal · LAB | LABTencent | 28%15 | not given | — | — | 27 Aug 2026 | |
| Claude Opus 4.6 [max]Anthropic · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 27.6% ±3.682 | $5.39 | 44k | 103 | 15 Jun 2026 | |
| GPT-5.5 [low]OpenAI · mini-swe-agent · DC | DCDatacurve | 27% ±2.291 | $1.20 | 9.4k | 28 | 15 Jun 2026 | |
| Grok 4.6 [xhigh]xAI · mini-swe-agent · IND | INDMercor | 24.8% ±6.836 | not given | — | — | not given | |
| GPT-5.4 mini [xhigh]OpenAI · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 24.3% ±2.962 | $2.08 | 135k | 86 | 15 Jun 2026 | |
| GPT-5.6 Terra [low]OpenAI · mini-swe-agent · DC | DCDatacurve | 24.1% ±0.781 | $0.34 | 8.6k | 22 | 10 Jul 2026 | |
| Kimi K2.6Moonshot AI · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 23.9% ±2.452 | $3.16 | 84k | 147 | 15 Jun 2026 | |
| Qwen3.7 MaxAlibaba · Claude Code · LAB | LABAlibaba (Qwen) | 21.6%11 | not given | — | — | 8 Aug 2026 | |
| MiniMax M3MiniMax · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 20.4% ±3.732 | $5.57 | 98k | 314 | 15 Jun 2026 | |
| MiMo-V2.5-ProXiaomi · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 19.5% ±1.872 | $1.99 | 49k | 122 | 15 Jun 2026 | |
| GLM-5.1Z.ai · mini-swe-agent · IND | INDMercor | 18.9% ±5.536 | not given | — | — | not given | |
| Qwen3.7 MaxAlibaba · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 17.7% ±1.422 | $2.12 | 42k | 110 | 15 Jun 2026 | |
| GLM-5.1Z.ai · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 17.5% ±0.792 | $7.46 | 49k | 177 | 15 Jun 2026 | |
| MiniMax M3 [thinking]MiniMax · mini-swe-agent · IND | INDentrpi (independent) | 16.8%34 | not given | — | — | 2 Jun 2026 | |
| Qwen3.7 PlusAlibaba · mini-swe-agent · LAB | LABAlibaba (Qwen) | 16.5%13 | not given | — | — | 24 Aug 2026 | |
| MiniMax M3 [thinking]MiniMax · mini-swe-agent · IND | INDentrpi (independent) | 13.3%34 | $7.48 | 80k | 325 | 2 Jun 2026 | |
| Qwen3.6-27BAlibaba · Claude Code · LAB | LABAlibaba (Qwen) | 13.3%10 | not given | — | — | 5 Aug 2026 | |
| Grok Build 0.1xAI · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 13.1% ±2.442 | $6.60 | 52k | 176 | 15 Jun 2026 | |
| Gemini 3.1 Pro (preview) [high]Google · mini-swe-agent · DC | DCDatacurve | 11.7% ±1.481 | $2.14 | 28k | 76 | 15 Jun 2026 | |
| GPT-5.6 Luna [medium]OpenAI · mini-swe-agent · DC | DCDatacurve | 11.3% ±0.831 | $0.043 | 8.2k | 24 | 10 Jul 2026 | |
| Gemini 3.1 Pro (preview)Google · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 9.7% ±2.832 | $1.84 | 53k | 74 | 15 Jun 2026 | |
| GPT-5.6 Luna [medium]OpenAI · lab-internal · LAB | LABOpenAI | 9.3%26 | $0.031 | — | — | 22 Sept 2026 | |
| DeepSeek V4 Pro [max]DeepSeek · mini-swe-agent · IND | INDMercor | 9.1% ±3.836 | not given | — | — | not given | |
| MiniMax M3 [high]MiniMax · mini-swe-agent · IND | INDMercor | 9.1% ±3.136 | not given | — | — | not given | |
| Gemini 3.1 Pro [high]Google · mini-swe-agent · IND | INDMercor | 7.7% ±3.136 | not given | — | — | not given | |
| DeepSeek V4 ProDeepSeek · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 7.5% ±2.72 | $4.22 | 50k | 111 | 15 Jun 2026 | |
| Gemini 3 Flash (preview)Google · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 5.1% ±2.482 | $1.53 | 233k | 71 | 15 Jun 2026 | |
| Inkling [high]Thinking Machines · mini-swe-agent · IND | INDMercor | 5% ±2.836 | not given | — | — | not given | |
| Qwen3.6 PlusAlibaba · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 2.6% ±1.232 | $4.25 | 67k | 164 | 15 Jun 2026 | |
| GPT-6 Luna [low]OpenAI · lab-internal · LAB | LABOpenAI | 2.4%26 | $0.006 | — | — | 22 Sept 2026 | |
| GLM-5Z.ai · mini-swe-agent · IND | INDMercor | 2.1% ±1.936 | not given | — | — | not given | |
| GPT-5.6 Luna [low]OpenAI · mini-swe-agent · DC | DCDatacurve | 1.6% ±0.831 | $0.015 | 3.1k | 13 | 10 Jul 2026 | |
| GPT-5.6 Luna [low]OpenAI · lab-internal · LAB | LABOpenAI | 1.2%26 | $0.011 | — | — | 22 Sept 2026 | |
| DeepSeek V3.2DeepSeek · mini-swe-agent · IND | INDMercor | 0.6% ±0.936 | not given | — | — | not given | |
| Nemotron 3 Ultra [high]NVIDIA · mini-swe-agent · IND | INDMercor | 0.6% ±0.736 | not given | — | — | not given | |
| Claude Haiku 4.5Anthropic · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 0.2% ±0.432 | $0.84 | 39k | 109 | 15 Jun 2026 | |
| MiniMax M2.7MiniMax · mini-swe-agent · DC | DCInternet Archive (capture of the Datacurve board) | 0.2% ±0.432 | $0.70 | 60k | 136 | 15 Jun 2026 | |
| Gemma 4 31BGoogle · mini-swe-agent · IND | INDMercor | 0% ±036 | not given | — | — | not given | |
| gpt-oss-120b [high]OpenAI · mini-swe-agent · IND | INDMercor | 0% ±036 | not given | — | — | not given | |
| Kimi K2 [high]Moonshot AI · mini-swe-agent · IND | INDMercor | 0% ±036 | not given | — | — | not given | |
| MiniMax M2.7 [high]MiniMax · mini-swe-agent · IND | INDMercor | 0% ±036 | not given | — | — | not given | |
| Qwen3.5Alibaba · mini-swe-agent · IND | INDMercor | 0% ±036 | not given | — | — | not given | |
Same model, different measurer
The official leaderboard’s bar form, one bar per source under each model: its best official result, its lab’s claim and any independent re-run. Only models measured by at least two sources are shown.
How to read these numbers
- The official board has added no model since 3 Sep 2026 (GPT-6 Astra; its runs finished on 1 Sep).3 Its data file was regenerated on 22 Sep with no new runs.5 A GitHub issue on the benchmark's repository quotes Datacurve's CEO saying the team is building the next version.39
- Epoch AI's review of 7 Sep 2026 rated DeepSWE 1.1 'flawed', finding grading defects in 23 of 113 tasks, mostly hidden tests colliding with tests the agent wrote.37 Tokenless separately documented 70 ways a submission can rewrite test outcomes.38 Treat differences of a few points with care.
- Lab claims are run by the lab, on its own harness, effort settings and trial count, so they are not directly comparable with the official runs (mini-swe-agent, four passes over 113 tasks).4 The harness is listed on every row. xAI says Datacurve ran the Grok 4.7 evaluation for it, though the result never appeared on the public board.23
- Independent runs come from Mercor (mini-swe-agent, 500 steps, 2-hour limit, 3 passes),36 Artificial Analysis (Grok Build harness),35 Fireworks (Kimi K3 at three effort levels)27 and entrpi (MiniMax M3).34
- Costs are per task in USD. The official page and its data file disagree for 19 of 70 configurations; this tracker shows the page's figure and keeps the file's in the row notes.5 Lab costs are the lab's own, shown only where the lab published them.
- Mercor shows no date for each model, so its results appear in the charts and table but not on the timeline.36
- A number is included only when it was read at the source that published it. Figures that appear only on aggregator sites are left out; research/ingest-log.txt in the repository lists each one and why.
- Anthropic has published no DeepSWE 1.1 score for Claude Haiku 5.5;33 Mercor's independent run is the only one. There is no GPT-6.1 Luna: OpenAI released GPT-6.1 Sol only.30
- Where a source published token counts but no cost, the cost is estimated from Vercel AI Gateway’s list prices and shown as a range marked est.40 With only a total token count, the range runs from every token priced as input to every token priced as output; cached input is cheaper still, so the true cost may sit below it. Only Laguna S 2.1 qualifies today.
How this was made
- What was gathered
- On 9 Oct 2026, Claude (Anthropic’s model) ran Dossier’s free local research loop with its own web search: six search tasks covering the official board, the US labs, Chinese and open-weight labs and Mistral, aggregator sites, and community and press coverage. A separate fact-check pass then settled the conflicts between them. No paid research API was used.
- What was kept
- The sweep returned 605 raw figures. 210 readings of 80 models survived the rule that a number is read at the source that published it, and the page cites 40 sources in the list below. Official rows are re-read from the board’s own page every day.
- What was verified
- Every lab figure was re-read on the lab’s own page or PDF. Three conflicts were settled against the primary document: GPT-6.1 Sol’s xhigh and max scores, who published Grok 4.7’s 73%, and the harness behind Muse Spark 1.3. Of the 22 sources the research report cites, none were fabricated; two OpenAI pages block automated reads and one StepFun page did not load.
- What could not be established
- When each model joined Mercor’s board; cost per task for Anthropic, Google and Meta models, which those labs do not publish; a dated StepFun page for Step 5 Preview; and why Datacurve ran Grok 4.7 without posting the result. No person has yet re-checked every row; corrections are welcome as issues on the repository.
Every original source, numbered
40 sources. Each entry says what it establishes and links back to the readings that cite it. The same list is in SOURCES.md, and every row of the dataset carries its source URL and the line quoted from it.
- Datacurve DeepSWE 1.1 leaderboardDatacurve · Official leaderboard · published 15 Jun 2026 · read 9 Oct 2026Pass@1 for GPT-6 Astra 67%–74.1% across 5 settings; Gemini 3.8 Flash 71%–73.8% across 2 settings; Claude Opus 5 58.1%–73.7% across 5 settings; GPT-5.6 Sol 45.4%–72.7% across 5 settings; Claude Fable 5 59.6%–69.9% across 5 settings; GPT-5.6 Terra 24.1%–69.6% across 5 settings; and 22 more models, with cost per task, tokens, confidence intervals.Cited by 70 readings: GPT-6 Astra [xhigh], Gemini 3.8 Flash [high], Claude Opus 5 [max], GPT-6 Astra [high], GPT-6 Astra [max], Claude Opus 5 [xhigh], Claude Opus 5 [high], GPT-6 Astra [medium], GPT-5.6 Sol [max], Gemini 3.8 Flash [medium], GPT-5.6 Sol [xhigh], Claude Fable 5 [xhigh],
and 58 more
Claude Fable 5 [max], GPT-5.6 Terra [max], GPT-5.6 Sol [high], GLM-5.3 [max], Claude Opus 5 [medium], Claude Fable 5 [high], Kimi K3 [max], Grok 4.6 [medium], GPT-5.6 Luna [max], GPT-5.5 [xhigh], GPT-6 Astra [low], Grok 4.6 [xhigh], Gemini 3.7 Flash [medium], Claude Fable 5 [medium], Gemini 3.7 Flash [high], Grok 4.6 [high], GPT-5.5 [high], GLM-5.3 Flash [max], DeepSeek V4 Pro [max], GPT-5.6 Sol [medium], GPT-5.6 Terra [xhigh], Claude Fable 5 [low], Claude Opus 4.8 [max], Claude Opus 5 [low], Qwen3.8 Max [xhigh], GPT-5.6 Luna [xhigh], Muse Spark 1.2 [xhigh], Claude Opus 4.8 [xhigh], GPT-5.5 [medium], Claude Sonnet 5 [max], Gemini 3.7 Flash [low], GPT-5.6 Terra [high], Grok 4.5 [high], DeepSeek V4 Flash [max], Muse Spark 1.1 [xhigh], Claude Opus 4.8 [high], GPT-5.4 [xhigh], Claude Sonnet 5 [xhigh], Claude Opus 4.8 [medium], Claude Sonnet 5 [high], Gemini 3.6 Flash [high], GPT-5.6 Sol [low], GPT-5.6 Luna [high], GLM-5.2 [max], Grok 4.6 [low], Claude Opus 4.8 [low], Claude Sonnet 5 [medium], GLM-5.2 [high], Gemini 3.5 Flash [high], GPT-5.6 Terra [medium], Kimi K2.7 Code, Claude Sonnet 5 [low], Claude Sonnet 4.6 [high], GPT-5.5 [low], GPT-5.6 Terra [low], Gemini 3.1 Pro (preview) [high], GPT-5.6 Luna [medium], GPT-5.6 Luna [low] - Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)Internet Archive (capture of the Datacurve board) · Official leaderboard · published 15 Jun 2026 · read 9 Oct 2026Pass@1 for Claude Opus 4.7 31.6%–54.2% across 4 settings; Gemini 3.5 Flash 28.3% [medium]; Claude Opus 4.6 27.6% [max]; GPT-5.4 mini 24.3% [xhigh]; Kimi K2.6 23.9%; MiniMax M3 20.4%; and 10 more models, with cost per task, tokens, confidence intervals.Cited by 19 readings: Claude Opus 4.7 [max], Claude Opus 4.7 [xhigh], Claude Opus 4.7 [high], Claude Opus 4.7 [medium], Gemini 3.5 Flash [medium], Claude Opus 4.6 [max], GPT-5.4 mini [xhigh], Kimi K2.6, MiniMax M3, MiMo-V2.5-Pro, Qwen3.7 Max, GLM-5.1,
- DeepSWE changelogDatacurve · Official changelog · published 3 Sept 2026 · read 9 Oct 2026Lists each model addition by date; the most recent is GPT-6 Astra (all efforts) on 3 Sep 2026.Cited in the notes above.
- Run DeepSWEDatacurve · Official documentation · published 15 Jun 2026 · read 9 Oct 2026Official scores are produced with Pier running mini-swe-agent on Modal, in isolated containers.Cited in the notes above.
- leaderboard-live.json (v1.1 data file)Datacurve · Official data file · published 22 Sept 2026 · read 9 Oct 2026The board's data file: generated 22 Sep 2026, latest job finished 1 Sep; its costs differ from the page's for 19 of 70 configurations.Cited in the notes above.
- GLM-5.2 HuggingFace Model CardZ.ai · Lab self-report · published 16 Jun 2026 · read 9 Oct 2026Pass@1 for GLM-5.2 46.2% [max].Cited by 1 reading: GLM-5.2 [max]
- Poolside Laguna S 2.1 launch postPoolside · Lab self-report · published 21 Jul 2026 · read 9 Oct 2026Pass@1 for Laguna S 2.1 40.4% [max].Cited by 1 reading: Laguna S 2.1 [max]
- Kimi-K3 HuggingFace Model CardMoonshot AI · Lab self-report · published 23 Jul 2026 · read 9 Oct 2026Pass@1 for Kimi K3 67.5% [max].Cited by 1 reading: Kimi K3 [max]
- Claude Opus 5 System CardAnthropic · Lab self-report · published 24 Jul 2026 · read 9 Oct 2026Pass@1 for Claude Opus 5 68.8% [max].Cited by 1 reading: Claude Opus 5 [max]
- Qwen3.8-27B HuggingFace Model CardAlibaba (Qwen) · Lab self-report · published 5 Aug 2026 · read 9 Oct 2026Pass@1 for Qwen3.8-27B 42.2%; Qwen3.6-27B 13.3%.Cited by 2 readings: Qwen3.8-27B, Qwen3.6-27B
- Qwen3.8-2.4T-A95B (Qwen3.8-Max) HuggingFace Model CardAlibaba (Qwen) · Lab self-report · published 8 Aug 2026 · read 9 Oct 2026Pass@1 for Qwen3.8 Max 56.6%; Qwen3.7 Max 21.6%.Cited by 2 readings: Qwen3.8 Max, Qwen3.7 Max
- xAI Grok 4.6 Model CardxAI · Lab self-report · published 18 Aug 2026 · read 9 Oct 2026Pass@1 for Grok 4.6 65.9% [high].Cited by 1 reading: Grok 4.6 [high]
- Qwen3.8-Flash-Next HuggingFace Model CardAlibaba (Qwen) · Lab self-report · published 24 Aug 2026 · read 9 Oct 2026Pass@1 for Qwen3.8-Flash-Next 58.7%; Qwen3.7 Plus 16.5%.Cited by 2 readings: Qwen3.8-Flash-Next, Qwen3.7 Plus
- GLM-5.3 HuggingFace Model CardZ.ai · Lab self-report · published 25 Aug 2026 · read 9 Oct 2026Pass@1 for GLM-5.3 66.9% [max].Cited by 1 reading: GLM-5.3 [max]
- Tencent Hy4-preview Technical Report & Model CardTencent · Lab self-report · published 27 Aug 2026 · read 9 Oct 2026Pass@1 for Hy4 Preview 64.3%; Hy3 28%.Cited by 2 readings: Hy4 Preview, Hy3
- Claude Fable 5.1 & Claude Mythos 5.1 System CardAnthropic · Lab self-report · published 1 Sept 2026 · read 9 Oct 2026Pass@1 for Claude Fable 5.1 67.4% [max].Cited by 1 reading: Claude Fable 5.1 [max]
- Meta AI Research Muse Spark 1.3 Evaluation MethodologyMeta · Lab self-report · published 2 Sept 2026 · read 9 Oct 2026Pass@1 for Muse Spark 1.3 75.4% [max].Cited by 1 reading: Muse Spark 1.3 [max]
- Nex-AGI Nex-N2.5 GitHub README / Model CardNex-AGI · Lab self-report · published 8 Sept 2026 · read 9 Oct 2026Pass@1 for Nex-N2.5-Max 65.6%; Nex-N2.5-Pro 55.8%; Nex-N2.5-Mini 36.1%.Cited by 3 readings: Nex-N2.5-Max, Nex-N2.5-Pro, Nex-N2.5-Mini
- Cognition SWE-2 launch postCognition · Lab self-report · published 10 Sept 2026 · read 9 Oct 2026Pass@1 for SWE-2 73%; SWE-1.7 37.7%.Cited by 2 readings: SWE-2, SWE-1.7
- DeepSeek-V4.1-Flash HuggingFace Model CardDeepSeek · Lab self-report · published 10 Sept 2026 · read 9 Oct 2026Pass@1 for DeepSeek V4.1 Flash 65.5%–74.2% across 8 settings; DeepSeek V4 Pro 62.7% [max]; DeepSeek V4 Flash 54.4% [max].Cited by 10 readings: DeepSeek V4.1 Flash [max], DeepSeek V4.1 Flash [max], DeepSeek V4.1 Flash [max], DeepSeek V4.1 Flash [max], DeepSeek V4.1 Flash [max], DeepSeek V4.1 Flash [max], DeepSeek V4.1 Flash [max], DeepSeek V4.1 Flash [max], DeepSeek V4 Pro [max], DeepSeek V4 Flash [max]
- Google Gemini 3.8 Flash launch blog and Model Evaluation ReportGoogle DeepMind · Lab self-report · published 17 Sept 2026 · read 9 Oct 2026Pass@1 for Gemini 3.8 Flash 73.7% [high], with cost per task, tokens.Cited by 1 reading: Gemini 3.8 Flash [high]
- StepFun Step 5 Preview AnnouncementStepFun · Lab self-report · published 20 Sept 2026 · read 9 Oct 2026Pass@1 for Step 5 Preview 67.7% [high].Cited by 1 reading: Step 5 Preview [high]
- xAI Grok 4.7 Model CardxAI · Lab self-report · published 21 Sept 2026 · read 9 Oct 2026Pass@1 for Grok 4.7 71% [high].Cited by 1 reading: Grok 4.7 [high]
- Xiaomi MiMo-V2.6 Launch Announcement and AppendixXiaomi · Lab self-report · published 21 Sept 2026 · read 9 Oct 2026Pass@1 for MiMo-V2.6-Pro 71.9% [max]; MiMo-V2.6-Flash 67.9% [max].Cited by 2 readings: MiMo-V2.6-Pro [max], MiMo-V2.6-Flash [max]
- Claude Opus 5.5 System CardAnthropic · Lab self-report · published 22 Sept 2026 · read 9 Oct 2026Pass@1 for Claude Opus 5.5 74.2% [max].Cited by 1 reading: Claude Opus 5.5 [max]
- OpenAI GPT-6 Sol and Luna launch postOpenAI · Lab self-report · published 22 Sept 2026 · read 9 Oct 2026Pass@1 for GPT-6 Sol 37.2%–68.8% across 5 settings; GPT-6 Luna 2.4%–66.6% across 5 settings; GPT-5.6 Luna 1.2%–62.2% across 5 settings, with cost per task.Cited by 15 readings: GPT-6 Sol [max], GPT-6 Luna [max], GPT-6 Sol [xhigh], GPT-6 Sol [high], GPT-5.6 Luna [max], GPT-6 Luna [xhigh], GPT-6 Luna [high], GPT-6 Sol [medium], GPT-5.6 Luna [xhigh], GPT-6 Luna [medium], GPT-5.6 Luna [high], GPT-6 Sol [low],
- Fireworks AI Ember-1 AnnouncementFireworks AI · Lab self-report · published 23 Sept 2026 · read 9 Oct 2026Pass@1 for Ember-1 75.2% [thinking]; Kimi K3 55.8%–66.4% across 3 settings, with cost per task.Cited by 4 readings: Ember-1 [thinking], Kimi K3 [max], Kimi K3 [high], Kimi K3 [low]
- Gemini 4 Argon Model Evaluation ReportGoogle DeepMind · Lab self-report · published 24 Sept 2026 · read 9 Oct 2026Pass@1 for Gemini 4 Argon 77.9%.Cited by 1 reading: Gemini 4 Argon
- Claude Sonnet 5.5 System CardAnthropic · Lab self-report · published 28 Sept 2026 · read 9 Oct 2026Pass@1 for Claude Sonnet 5.5 71% [max].Cited by 1 reading: Claude Sonnet 5.5 [max]
- OpenAI GPT-6.1 Sol launch postOpenAI · Lab self-report · published 29 Sept 2026 · read 9 Oct 2026Pass@1 for GPT-6.1 Sol 64.4%–75.2% across 5 settings, with cost per task.Cited by 5 readings: GPT-6.1 Sol [high], GPT-6.1 Sol [medium], GPT-6.1 Sol [max], GPT-6.1 Sol [xhigh], GPT-6.1 Sol [low]
- Reflection AI Beam launch postReflection AI · Lab self-report · published 5 Oct 2026 · read 9 Oct 2026Pass@1 for Beam 44.4%.Cited by 1 reading: Beam
- Mistral Large 4 AnnouncementMistral AI · Lab self-report · published 6 Oct 2026 · read 9 Oct 2026Pass@1 for Mistral Large 4 61.7% [thinking].Cited by 1 reading: Mistral Large 4 [thinking]
- Claude Haiku 5.5 System CardAnthropic · Lab system card · published 29 Sept 2026 · read 9 Oct 2026Reports SWE-bench Pro, FrontierSWE v2 and ProgramBench for Haiku 5.5, and no DeepSWE 1.1 result.Cited in the notes above.
- Independent DeepSWE Audit by entrpientrpi (independent) · Independent run · published 2 Jun 2026 · read 9 Oct 2026Pass@1 for MiniMax M3 13.3%–16.8% across 2 settings, with cost per task, tokens.Cited by 2 readings: MiniMax M3 [thinking], MiniMax M3 [thinking]
- Artificial Analysis (Benchmarking Grok 4.7)Artificial Analysis · Independent run · published 21 Sept 2026 · read 9 Oct 2026Pass@1 for Grok 4.7 73% [xhigh]; Grok 4.6 65% [high].Cited by 2 readings: Grok 4.7 [xhigh], Grok 4.6 [high]
- Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)Mercor · Independent run · published undated · read 9 Oct 2026Pass@1 for Claude Opus 5.5 72.3% [max]; GPT-6.1 Sol 72.3% [max]; GPT-6 Astra 72% [max]; DeepSeek V4.1 Flash 71.7% [max]; Claude Opus 5 70.2%–71.4% across 2 settings; Gemini 3.8 Flash 71.4% [high]; and 43 more models, with confidence intervals.Cited by 52 readings: Claude Opus 5.5 [max], GPT-6.1 Sol [max], GPT-6 Astra [max], DeepSeek V4.1 Flash [max], Claude Opus 5 [max], Gemini 3.8 Flash [high], Claude Sonnet 5.5 [max], GLM-5.3 [max], GPT-5.6 Sol [max], GPT-6 Sol [max], Claude Opus 5 [xhigh], GPT-5.6 Terra [max],
and 40 more
Claude Fable 5 [max], Claude Fable 5.1 [high], GPT-6 Luna [max], GPT-5.5 [xhigh], Gemini 3.7 Flash [high], Claude Fable 5.1 [max], GPT-5.6 Luna [max], Grok 4.6 [high], Claude Haiku 5.5 [max], Claude Opus 4.8 [max], DeepSeek V4 Flash [max], DeepSeek V4 Pro 0813 [max], Qwen3.8 Max [xhigh], Claude Sonnet 5 [max], Claude Opus 4.7 [max], Gemini 3.6 Flash [high], Grok 4.5 [high], GPT-5.4 [xhigh], GLM-5.3 Flash [max], Muse Spark 1.1 [xhigh], GLM-5.2 [max], Gemini 3.5 Flash [high], Claude Sonnet 4.6 [high], Grok 4.7 [xhigh], Claude Opus 4.6 [max], Kimi K2.7 Code [high], Grok 4.6 [xhigh], GLM-5.1, DeepSeek V4 Pro [max], MiniMax M3 [high], Gemini 3.1 Pro [high], Inkling [high], GLM-5, DeepSeek V3.2, Nemotron 3 Ultra [high], Gemma 4 31B, gpt-oss-120b [high], Kimi K2 [high], MiniMax M2.7 [high], Qwen3.5 - Benchmark review: DeepSWE v1.1Epoch AI · Independent audit · published 7 Sept 2026 · read 9 Oct 2026Rates DeepSWE v1.1 flawed: grading defects in at least 23 of 113 tasks, most from hidden tests colliding with tests the agent wrote.Cited in the notes above.
- envcheck: auditing DeepSWETokenless · Independent audit · published 30 Sept 2026 · read 9 Oct 2026Documents 70 ways a submission can rewrite test outcomes, because the test harness sits inside the repository the agent edits.Cited in the notes above.
- datacurve-ai/deep-swe issue #103GitHub (community discussion) · Discussion · published 30 Sept 2026 · read 9 Oct 2026Notes the board has published no new model since 3 Sep and quotes Datacurve's CEO saying the team is building the next generation of DeepSWE.Cited in the notes above.
- AI Gateway model list and pricesVercel · Price list · published undated · read 9 Oct 2026Per-token list prices used to estimate cost where a source published tokens but no cost: Laguna S 2.1 at $0.09 per million input tokens and $0.18 per million output tokens.Cited in the notes above.