DeepSWE 1.1 Tracker

Official board last added a model on 3, days ago.

Independent · not affiliated with Datacurve · source and data
Fig. 1 · Readings against cost

Score against what it costs to get it

Drawn the way the official board draws it: colour is the lab, and each line joins one model’s effort levels from low to max. The models you want are top right: higher up solves more of the 113 tasks, further right costs less per task. Shape and line style say who measured it. Hover a point for its source; hover a row in the table to pick out its line.

DeepSWE 1.1 score against cost per taskScatter of 113 results: pass@1 on the vertical axis, avg cost per task on the horizontal from highest at left to lowest at right; each model’s effort levels are joined in order from low to max.DeepSWE score ↑ more of the 113 tasks solvedBest: higher score, cheaper ↗0%10%20%30%40%50%60%70%80%$0.005$0.025$0.10$0.50$2.50$10Avg cost per task (log scale) · cheaper →No other result here scores higher for lessEfficient frontierGPT-6.1 Sol [high]: 75.2%GPT-6.1 Sol [medium]: 73%GPT-6.1 Sol [max]: 71.9%GPT-6.1 Sol [xhigh]: 71.9%GPT-6.1 Sol [low]: 64.4%Ember-1 [thinking]: 75.2%GPT-6 Astra [xhigh]: 74.1%GPT-6 Astra [high]: 73.2%GPT-6 Astra [max]: 73.2%GPT-6 Astra [medium]: 72.8%GPT-6 Astra [low]: 67%Gemini 3.8 Flash [high]: 73.8%Gemini 3.8 Flash [medium]: 71%Gemini 3.8 Flash [high]: 73.7%Claude Opus 5 [max]: 73.7%Claude Opus 5 [xhigh]: 73.2%Claude Opus 5 [high]: 72.8%Claude Opus 5 [medium]: 68.9%Claude Opus 5 [low]: 58.1%GPT-5.6 Sol [max]: 72.7%GPT-5.6 Sol [xhigh]: 70.7%GPT-5.6 Sol [high]: 69.4%GPT-5.6 Sol [medium]: 61.1%GPT-5.6 Sol [low]: 45.4%Claude Fable 5 [xhigh]: 69.9%Claude Fable 5 [max]: 69.7%Claude Fable 5 [high]: 68.6%Claude Fable 5 [medium]: 65.4%Claude Fable 5 [low]: 59.6%GPT-5.6 Terra [max]: 69.6%GPT-5.6 Terra [xhigh]: 60.2%GPT-5.6 Terra [high]: 53.8%GPT-5.6 Terra [medium]: 35.1%GPT-5.6 Terra [low]: 24.1%GLM-5.3 [max]: 69%GPT-6 Sol [max]: 68.8%GPT-6 Sol [xhigh]: 66.6%GPT-6 Sol [high]: 65.3%GPT-6 Sol [medium]: 56.6%GPT-6 Sol [low]: 37.2%Kimi K3 [max]: 68.5%Grok 4.6 [medium]: 67.5%Grok 4.6 [xhigh]: 66.7%Grok 4.6 [high]: 65.2%Grok 4.6 [low]: 41.6%GPT-5.6 Luna [max]: 67.2%GPT-5.6 Luna [xhigh]: 56.9%GPT-5.6 Luna [high]: 44.3%GPT-5.6 Luna [medium]: 11.3%GPT-5.6 Luna [low]: 1.6%GPT-5.5 [xhigh]: 67%GPT-5.5 [high]: 64.4%GPT-5.5 [medium]: 54%GPT-5.5 [low]: 27%GPT-6 Luna [max]: 66.6%GPT-6 Luna [xhigh]: 61.3%GPT-6 Luna [high]: 59.3%GPT-6 Luna [medium]: 44.5%GPT-6 Luna [low]: 2.4%Kimi K3 [max]: 66.4%Gemini 3.7 Flash [medium]: 65.5%Gemini 3.7 Flash [high]: 65.3%Gemini 3.7 Flash [low]: 53.8%GLM-5.3 Flash [max]: 63.4%DeepSeek V4 Pro [max]: 62.8%GPT-5.6 Luna [max]: 62.2%GPT-5.6 Luna [xhigh]: 56.2%GPT-5.6 Luna [high]: 42.4%GPT-5.6 Luna [medium]: 9.3%GPT-5.6 Luna [low]: 1.2%Claude Opus 4.8 [max]: 59%Claude Opus 4.8 [xhigh]: 54.4%Claude Opus 4.8 [high]: 51.8%Claude Opus 4.8 [medium]: 48.7%Claude Opus 4.8 [low]: 40.8%Qwen3.8 Max [xhigh]: 57.5%Muse Spark 1.2 [xhigh]: 54.9%Claude Opus 4.7 [max]: 54.2%Claude Opus 4.7 [xhigh]: 44.7%Claude Opus 4.7 [high]: 40.3%Claude Opus 4.7 [medium]: 31.6%Claude Sonnet 5 [max]: 53.9%Claude Sonnet 5 [xhigh]: 49.7%Claude Sonnet 5 [high]: 48.2%Claude Sonnet 5 [medium]: 39.8%Claude Sonnet 5 [low]: 30.5%Grok 4.5 [high]: 53.8%DeepSeek V4 Flash [max]: 53.3%Muse Spark 1.1 [xhigh]: 53.3%GPT-5.4 [xhigh]: 51.8%Gemini 3.6 Flash [high]: 46.7%GLM-5.2 [max]: 43.8%GLM-5.2 [high]: 36.3%Gemini 3.5 Flash [high]: 36.1%Kimi K2.7 Code: 30.5%Claude Sonnet 4.6 [high]: 29.9%Gemini 3.5 Flash [medium]: 28.3%Claude Opus 4.6 [max]: 27.6%GPT-5.4 mini [xhigh]: 24.3%Kimi K2.6: 23.9%MiniMax M3: 20.4%MiMo-V2.5-Pro: 19.5%Qwen3.7 Max: 17.7%GLM-5.1: 17.5%MiniMax M3 [thinking]: 13.3%Grok Build 0.1: 13.1%Gemini 3.1 Pro (preview) [high]: 11.7%Gemini 3.1 Pro (preview): 9.7%DeepSeek V4 Pro: 7.5%Gemini 3 Flash (preview): 5.1%Qwen3.6 Plus: 2.6%Claude Haiku 4.5: 0.2%MiniMax M2.7: 0.2%GPT-6.1 SolHIGHEmber-1THINKINGGPT-6 AstraXHIGHClaude Opus 5MAXClaude Fable 5XHIGHGPT-6 SolMAXKimi K3MAXGrok 4.6MEDIUMGPT-5.6 LunaMAXGPT-6 LunaMAXGLM-5.3 FlashMAXDeepSeek V4 ProMAXClaude Opus 4.8MAXQwen3.8 MaxXHIGHMuse Spark 1.2XHIGHClaude Opus 4.7MAXDeepSeek V4 FlashMAXGPT-5.4XHIGHGemini 3.6 FlashHIGHGLM-5.2MAXGemini 3.5 FlashHIGH

Fig. 2 · Readings

Every reading, ranked

The official leaderboard’s form: bars on a 0–80% scale coloured by lab, with the 95% interval as a whisker where the source gives one. The number after each score is its source in the list below; a cost marked est. is estimated from list prices. Select a heading to sort.

DeepSWE 1.1 readings for the current filters
Bar
Gemini 4 ArgonGoogle · mini-swe-agent · LABLABGoogle DeepMind77.9%28not given——24 Sept 2026
Muse Spark 1.3 [max]Meta · mini-swe-agent · LABLABMeta75.4%17not given——2 Sept 2026
GPT-6.1 Sol [high]OpenAI · lab-internal · LABLABOpenAI75.2%30$0.65——29 Sept 2026
Ember-1 [thinking]Fireworks AI · mini-swe-agent · LABLABFireworks AI75.2%27$3.62——23 Sept 2026
Claude Opus 5.5 [max]Anthropic · lab-internal · LABLABAnthropic74.2%25not given——22 Sept 2026
DeepSeek V4.1 Flash [max]DeepSeek · mini-swe-agent · LABLABDeepSeek74.2%20not given——10 Sept 2026
GPT-6 Astra [xhigh]OpenAI · mini-swe-agent · DCDCDatacurve74.1% ±2.871$4.4330k293 Sept 2026
Gemini 3.8 Flash [high]Google · mini-swe-agent · DCDCDatacurve73.8% ±1.421$2.36143k1661 Sept 2026
Gemini 3.8 Flash [high]Google · mini-swe-agent · LABLABGoogle DeepMind73.7%21$2.36143k—17 Sept 2026
Claude Opus 5 [max]Anthropic · mini-swe-agent · DCDCDatacurve73.7% ±3.871$12118k9925 Jul 2026
GPT-6 Astra [high]OpenAI · mini-swe-agent · DCDCDatacurve73.2% ±3.421$3.9227k273 Sept 2026
GPT-6 Astra [max]OpenAI · mini-swe-agent · DCDCDatacurve73.2% ±0.831$7.5061k293 Sept 2026
Claude Opus 5 [xhigh]Anthropic · mini-swe-agent · DCDCDatacurve73.2% ±3.061$9.0792k8925 Jul 2026
GPT-6.1 Sol [medium]OpenAI · lab-internal · LABLABOpenAI73%30$0.42——29 Sept 2026
Grok 4.7 [xhigh]xAI · Grok Build · INDINDArtificial Analysis73%35not given——21 Sept 2026
SWE-2Cognition · Devin CLI · LABLABCognition73%19not given——10 Sept 2026
Claude Opus 5 [high]Anthropic · mini-swe-agent · DCDCDatacurve72.8% ±1.951$6.0864k7325 Jul 2026
GPT-6 Astra [medium]OpenAI · mini-swe-agent · DCDCDatacurve72.8% ±2.591$3.0820k263 Sept 2026
GPT-5.6 Sol [max]OpenAI · mini-swe-agent · DCDCDatacurve72.7% ±2.831$6.4660k6110 Jul 2026
DeepSeek V4.1 Flash [max]DeepSeek · lab-internal · LABLABDeepSeek72.6%20not given——10 Sept 2026
Claude Opus 5.5 [max]Anthropic · mini-swe-agent · INDINDMercor72.3% ±7.336not given——not given
GPT-6.1 Sol [max]OpenAI · mini-swe-agent · INDINDMercor72.3% ±7.436not given——not given
GPT-6 Astra [max]OpenAI · mini-swe-agent · INDINDMercor72% ±7.536not given——not given
GPT-6.1 Sol [max]OpenAI · lab-internal · LABLABOpenAI71.9%30$1.57——29 Sept 2026
GPT-6.1 Sol [xhigh]OpenAI · lab-internal · LABLABOpenAI71.9%30$0.79——29 Sept 2026
MiMo-V2.6-Pro [max]Xiaomi · lab-internal · LABLABXiaomi71.9%24not given——21 Sept 2026
DeepSeek V4.1 Flash [max]DeepSeek · mini-swe-agent · INDINDMercor71.7% ±6.336not given——not given
Claude Opus 5 [max]Anthropic · mini-swe-agent · INDINDMercor71.4% ±6.836not given——not given
Gemini 3.8 Flash [high]Google · mini-swe-agent · INDINDMercor71.4% ±6.836not given——not given
Gemini 3.8 Flash [medium]Google · mini-swe-agent · DCDCDatacurve71% ±2.281$1.97125k1471 Sept 2026
Claude Sonnet 5.5 [max]Anthropic · lab-internal · LABLABAnthropic71%29not given——28 Sept 2026
Grok 4.7 [high]xAI · mini-swe-agent · LABLABxAI71%23not given——21 Sept 2026
GPT-5.6 Sol [xhigh]OpenAI · mini-swe-agent · DCDCDatacurve70.7% ±0.821$3.6041k4410 Jul 2026
Claude Sonnet 5.5 [max]Anthropic · mini-swe-agent · INDINDMercor70.5% ±7.736not given——not given
DeepSeek V4.1 Flash [max]DeepSeek · lab-internal · LABLABDeepSeek70.5%20not given——10 Sept 2026
GLM-5.3 [max]Z.ai · mini-swe-agent · INDINDMercor70.5% ±6.336not given——not given
GPT-5.6 Sol [max]OpenAI · mini-swe-agent · INDINDMercor70.5% ±6.936not given——not given
GPT-6 Sol [max]OpenAI · mini-swe-agent · INDINDMercor70.5% ±7.236not given——not given
Claude Opus 5 [xhigh]Anthropic · mini-swe-agent · INDINDMercor70.2% ±6.936not given——not given
GPT-5.6 Terra [max]OpenAI · mini-swe-agent · INDINDMercor70.2% ±7.236not given——not given
Claude Fable 5 [xhigh]Anthropic · mini-swe-agent · DCDCDatacurve69.9% ±3.241$1380k6815 Jun 2026
DeepSeek V4.1 Flash [max]DeepSeek · Claude Code · LABLABDeepSeek69.8%20not given——10 Sept 2026
Claude Fable 5 [max]Anthropic · mini-swe-agent · DCDCDatacurve69.7% ±4.031$22119k8815 Jun 2026
GPT-5.6 Terra [max]OpenAI · mini-swe-agent · DCDCDatacurve69.6% ±2.561$3.9672k7610 Jul 2026
GPT-5.6 Sol [high]OpenAI · mini-swe-agent · DCDCDatacurve69.4% ±1.431$2.6628k3710 Jul 2026
GLM-5.3 [max]Z.ai · mini-swe-agent · DCDCDatacurve69% ±3.021$3.9980k12520 Aug 2026
Claude Opus 5 [medium]Anthropic · mini-swe-agent · DCDCDatacurve68.9% ±1.171$3.2937k5225 Jul 2026
GPT-6 Sol [max]OpenAI · lab-internal · LABLABOpenAI68.8%26$2.74——22 Sept 2026
Claude Opus 5 [max]Anthropic · lab-internal · LABLABAnthropic68.8%9not given——24 Jul 2026
Claude Fable 5 [high]Anthropic · mini-swe-agent · DCDCDatacurve68.6% ±1.121$9.1857k5915 Jun 2026
Kimi K3 [max]Moonshot AI · mini-swe-agent · DCDCDatacurve68.5% ±4.541$4.6582k9818 Jul 2026
MiMo-V2.6-Flash [max]Xiaomi · lab-internal · LABLABXiaomi67.9%24not given——21 Sept 2026
Step 5 Preview [high]StepFun · mini-swe-agent · LABLABStepFun67.7%22not given——20 Sept 2026
DeepSeek V4.1 Flash [max]DeepSeek · lab-internal · LABLABDeepSeek67.6%20not given——10 Sept 2026
Kimi K3 [max]Moonshot AI · lab-internal · LABLABMoonshot AI67.5%8not given——23 Jul 2026
Grok 4.6 [medium]xAI · mini-swe-agent · DCDCDatacurve67.5% ±2.281$3.4550k7012 Aug 2026
Claude Fable 5.1 [max]Anthropic · lab-internal · LABLABAnthropic67.4%16not given——1 Sept 2026
Claude Fable 5 [max]Anthropic · mini-swe-agent · INDINDMercor67.3% ±7.236not given——not given
Claude Fable 5.1 [high]Anthropic · mini-swe-agent · INDINDMercor67.3% ±6.936not given——not given
GPT-5.6 Luna [max]OpenAI · mini-swe-agent · DCDCDatacurve67.2% ±3.991$0.6173k10210 Jul 2026
GPT-5.5 [xhigh]OpenAI · mini-swe-agent · DCDCDatacurve67% ±6.471$7.2346k8215 Jun 2026
GPT-6 Astra [low]OpenAI · mini-swe-agent · DCDCDatacurve67% ±1.31$1.6011k203 Sept 2026
GPT-6 Luna [max]OpenAI · mini-swe-agent · INDINDMercor67% ±6.836not given——not given
GLM-5.3 [max]Z.ai · mini-swe-agent · LABLABZ.ai66.9%14not given——25 Aug 2026
Grok 4.6 [xhigh]xAI · mini-swe-agent · DCDCDatacurve66.7% ±2.181$5.5071k8712 Aug 2026
GPT-6 Luna [max]OpenAI · lab-internal · LABLABOpenAI66.6%26$0.22——22 Sept 2026
GPT-6 Sol [xhigh]OpenAI · lab-internal · LABLABOpenAI66.6%26$1.00——22 Sept 2026
Kimi K3 [max]Moonshot AI · mini-swe-agent · INDINDFireworks AI66.4%27$4.74——23 Sept 2026
DeepSeek V4.1 Flash [max]DeepSeek · Pi · LABLABDeepSeek66.2%20not given——10 Sept 2026
GPT-5.5 [xhigh]OpenAI · mini-swe-agent · INDINDMercor66.1% ±6.936not given——not given
Grok 4.6 [high]xAI · mini-swe-agent · LABLABxAI65.9%12not given——18 Aug 2026
DeepSeek V4.1 Flash [max]DeepSeek · Codex · LABLABDeepSeek65.6%20not given——10 Sept 2026
Nex-N2.5-MaxNex-AGI · NexAU · LABLABNex-AGI65.6%18not given——8 Sept 2026
DeepSeek V4.1 Flash [max]DeepSeek · OpenCode · LABLABDeepSeek65.5%20not given——10 Sept 2026
Gemini 3.7 Flash [high]Google · mini-swe-agent · INDINDMercor65.5% ±7.136not given——not given
Gemini 3.7 Flash [medium]Google · mini-swe-agent · DCDCDatacurve65.5% ±3.091$2.0394k11713 Aug 2026
Claude Fable 5 [medium]Anthropic · mini-swe-agent · DCDCDatacurve65.4% ±4.421$6.0940k4815 Jun 2026
Gemini 3.7 Flash [high]Google · mini-swe-agent · DCDCDatacurve65.3% ±1.791$2.18107k12513 Aug 2026
GPT-6 Sol [high]OpenAI · lab-internal · LABLABOpenAI65.3%26$0.64——22 Sept 2026
Grok 4.6 [high]xAI · mini-swe-agent · DCDCDatacurve65.2% ±1.531$4.3861k7912 Aug 2026
Grok 4.6 [high]xAI · Grok Build · INDINDArtificial Analysis65%35not given——21 Sept 2026
Claude Fable 5.1 [max]Anthropic · mini-swe-agent · INDINDMercor64.9% ±7.236not given——not given
GPT-5.5 [high]OpenAI · mini-swe-agent · DCDCDatacurve64.4% ±3.121$5.1031k6215 Jun 2026
GPT-6.1 Sol [low]OpenAI · lab-internal · LABLABOpenAI64.4%30$0.17——29 Sept 2026
GPT-5.6 Luna [max]OpenAI · mini-swe-agent · INDINDMercor64.3% ±7.436not given——not given
Hy4 PreviewTencent · lab-internal · LABLABTencent64.3%15not given——27 Aug 2026
Grok 4.6 [high]xAI · mini-swe-agent · INDINDMercor63.7% ±7.436not given——not given
GLM-5.3 Flash [max]Z.ai · mini-swe-agent · DCDCDatacurve63.4% ±4.381$0.2473k12326 Aug 2026
DeepSeek V4 Pro [max]DeepSeek · mini-swe-agent · DCDCDatacurve62.8% ±6.331$1.67106k15512 Aug 2026
Kimi K3 [high]Moonshot AI · mini-swe-agent · INDINDFireworks AI62.8%27not given——23 Sept 2026
DeepSeek V4 Pro [max]DeepSeek · mini-swe-agent · LABLABDeepSeek62.7%20not given——10 Sept 2026
GPT-5.6 Luna [max]OpenAI · lab-internal · LABLABOpenAI62.2%26$0.53——22 Sept 2026
Mistral Large 4 [thinking]Mistral · partner eval (Artificial Analysis, Surge AI) · LABLABMistral AI61.7%32not given——6 Oct 2026
GPT-6 Luna [xhigh]OpenAI · lab-internal · LABLABOpenAI61.3%26$0.11——22 Sept 2026
GPT-5.6 Sol [medium]OpenAI · mini-swe-agent · DCDCDatacurve61.1% ±1.581$1.4218k3110 Jul 2026
GPT-5.6 Terra [xhigh]OpenAI · mini-swe-agent · DCDCDatacurve60.2% ±2.121$1.7040k4310 Jul 2026
Claude Haiku 5.5 [max]Anthropic · mini-swe-agent · INDINDMercor59.9% ±7.736not given——not given
Claude Opus 4.8 [max]Anthropic · mini-swe-agent · INDINDMercor59.6% ±7.136not given——not given
Claude Fable 5 [low]Anthropic · mini-swe-agent · DCDCDatacurve59.6% ±2.791$3.7625k3815 Jun 2026
GPT-6 Luna [high]OpenAI · lab-internal · LABLABOpenAI59.3%26$0.084——22 Sept 2026
Claude Opus 4.8 [max]Anthropic · mini-swe-agent · DCDCDatacurve59% ±1.761$13135k12015 Jun 2026
Qwen3.8-Flash-NextAlibaba · mini-swe-agent · LABLABAlibaba (Qwen)58.7%13not given——24 Aug 2026
Claude Opus 5 [low]Anthropic · mini-swe-agent · DCDCDatacurve58.1% ±2.331$1.6620k3625 Jul 2026
Qwen3.8 Max [xhigh]Alibaba · mini-swe-agent · DCDCDatacurve57.5% ±2.661$3.7395k1114 Aug 2026
GPT-5.6 Luna [xhigh]OpenAI · mini-swe-agent · DCDCDatacurve56.9% ±2.171$0.3145k7110 Jul 2026
GPT-6 Sol [medium]OpenAI · lab-internal · LABLABOpenAI56.6%26$0.38——22 Sept 2026
DeepSeek V4 Flash [max]DeepSeek · mini-swe-agent · INDINDMercor56.6% ±8.836not given——not given
Qwen3.8 MaxAlibaba · Claude Code · LABLABAlibaba (Qwen)56.6%11not given——8 Aug 2026
DeepSeek V4 Pro 0813 [max]DeepSeek · mini-swe-agent · INDINDMercor56.3% ±7.236not given——not given
GPT-5.6 Luna [xhigh]OpenAI · lab-internal · LABLABOpenAI56.2%26$0.27——22 Sept 2026
Kimi K3 [low]Moonshot AI · mini-swe-agent · INDINDFireworks AI55.8%27not given——23 Sept 2026
Nex-N2.5-ProNex-AGI · NexAU · LABLABNex-AGI55.8%18not given——8 Sept 2026
Muse Spark 1.2 [xhigh]Meta · mini-swe-agent · DCDCDatacurve54.9% ±2.121$3.7099k1017 Aug 2026
DeepSeek V4 Flash [max]DeepSeek · mini-swe-agent · LABLABDeepSeek54.4%20not given——10 Sept 2026
Claude Opus 4.8 [xhigh]Anthropic · mini-swe-agent · DCDCDatacurve54.4% ±3.711$8.0186k9515 Jun 2026
Claude Opus 4.7 [max]Anthropic · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)54.2% ±4.722$18103k20315 Jun 2026
GPT-5.5 [medium]OpenAI · mini-swe-agent · DCDCDatacurve54% ±2.551$2.7520k4615 Jun 2026
Claude Sonnet 5 [max]Anthropic · mini-swe-agent · DCDCDatacurve53.9% ±4.241$26214k2692 Jul 2026
Gemini 3.7 Flash [low]Google · mini-swe-agent · DCDCDatacurve53.8% ±2.591$1.8373k13013 Aug 2026
GPT-5.6 Terra [high]OpenAI · mini-swe-agent · DCDCDatacurve53.8% ±4.331$0.9122k3410 Jul 2026
Grok 4.5 [high]xAI · mini-swe-agent · DCDCDatacurve53.8% ±2.281$2.4236k6116 Jul 2026
DeepSeek V4 Flash [max]DeepSeek · mini-swe-agent · DCDCDatacurve53.3% ±3.571$0.46108k1536 Aug 2026
Muse Spark 1.1 [xhigh]Meta · mini-swe-agent · DCDCDatacurve53.3% ±3.041$2.3674k9614 Jul 2026
Qwen3.8 Max [xhigh]Alibaba · mini-swe-agent · INDINDMercor53.1% ±8.836not given——not given
Claude Opus 4.8 [high]Anthropic · mini-swe-agent · DCDCDatacurve51.8% ±4.561$4.2850k7315 Jun 2026
GPT-5.4 [xhigh]OpenAI · mini-swe-agent · DCDCDatacurve51.8% ±1.51$5.6571k7115 Jun 2026
Claude Sonnet 5 [xhigh]Anthropic · mini-swe-agent · DCDCDatacurve49.7% ±3.451$12121k1862 Jul 2026
Claude Sonnet 5 [max]Anthropic · mini-swe-agent · INDINDMercor49.6% ±7.236not given——not given
Claude Opus 4.7 [max]Anthropic · mini-swe-agent · INDINDMercor49.3% ±6.936not given——not given
Gemini 3.6 Flash [high]Google · mini-swe-agent · INDINDMercor49.3% ±7.136not given——not given
Grok 4.5 [high]xAI · mini-swe-agent · INDINDMercor49.3% ±7.436not given——not given
GPT-5.4 [xhigh]OpenAI · mini-swe-agent · INDINDMercor49% ±6.836not given——not given
Claude Opus 4.8 [medium]Anthropic · mini-swe-agent · DCDCDatacurve48.7% ±2.241$3.4441k6615 Jun 2026
Claude Sonnet 5 [high]Anthropic · mini-swe-agent · DCDCDatacurve48.2% ±4.511$7.4387k1472 Jul 2026
Gemini 3.6 Flash [high]Google · mini-swe-agent · DCDCDatacurve46.7% ±3.71$2.2196k11722 Jul 2026
GLM-5.2 [max]Z.ai · mini-swe-agent · LABLABZ.ai46.2%6not given——16 Jun 2026
GPT-5.6 Sol [low]OpenAI · mini-swe-agent · DCDCDatacurve45.4% ±2.391$0.8211k2310 Jul 2026
Claude Opus 4.7 [xhigh]Anthropic · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)44.7% ±2.882$8.5866k12615 Jun 2026
GPT-6 Luna [medium]OpenAI · lab-internal · LABLABOpenAI44.5%26$0.052——22 Sept 2026
BeamReflection AI · LABLABReflection AI44.4%31not given——5 Oct 2026
GPT-5.6 Luna [high]OpenAI · mini-swe-agent · DCDCDatacurve44.3% ±2.921$0.1626k4910 Jul 2026
GLM-5.2 [max]Z.ai · mini-swe-agent · DCDCDatacurve43.8% ±1.731$3.9278k12921 Jun 2026
GLM-5.3 Flash [max]Z.ai · mini-swe-agent · INDINDMercor42.5% ±6.536not given——not given
GPT-5.6 Luna [high]OpenAI · lab-internal · LABLABOpenAI42.4%26$0.13——22 Sept 2026
Qwen3.8-27BAlibaba · Claude Code · LABLABAlibaba (Qwen)42.2%10not given——5 Aug 2026
Grok 4.6 [low]xAI · mini-swe-agent · DCDCDatacurve41.6% ±2.321$1.0416k4412 Aug 2026
Claude Opus 4.8 [low]Anthropic · mini-swe-agent · DCDCDatacurve40.8% ±1.461$2.2929k5415 Jun 2026
Laguna S 2.1 [max]Poolside · pool · LABLABPoolside40.4%7≈$0.022–0.045 est.40249k total—21 Jul 2026
Muse Spark 1.1 [xhigh]Meta · mini-swe-agent · INDINDMercor40.4% ±6.536not given——not given
Claude Opus 4.7 [high]Anthropic · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)40.3% ±3.52$4.8346k8715 Jun 2026
Claude Sonnet 5 [medium]Anthropic · mini-swe-agent · DCDCDatacurve39.8% ±3.131$4.0857k1082 Jul 2026
GLM-5.2 [max]Z.ai · mini-swe-agent · INDINDMercor38.9% ±6.536not given——not given
SWE-1.7Cognition · Devin CLI · LABLABCognition37.7%19not given——10 Sept 2026
GPT-6 Sol [low]OpenAI · lab-internal · LABLABOpenAI37.2%26$0.16——22 Sept 2026
Gemini 3.5 Flash [high]Google · mini-swe-agent · INDINDMercor36.6% ±6.936not given——not given
GLM-5.2 [high]Z.ai · mini-swe-agent · DCDCDatacurve36.3% ±4.751$2.8454k12221 Jun 2026
Nex-N2.5-MiniNex-AGI · NexAU · LABLABNex-AGI36.1%18not given——8 Sept 2026
Gemini 3.5 Flash [high]Google · mini-swe-agent · DCDCDatacurve36.1% ±3.971$3.4576k10515 Jun 2026
GPT-5.6 Terra [medium]OpenAI · mini-swe-agent · DCDCDatacurve35.1% ±3.381$0.4712k2510 Jul 2026
Claude Sonnet 4.6 [high]Anthropic · mini-swe-agent · INDINDMercor34.2% ±6.336not given——not given
Grok 4.7 [xhigh]xAI · mini-swe-agent · INDINDMercor33.3% ±636not given——not given
Claude Opus 4.6 [max]Anthropic · mini-swe-agent · INDINDMercor31.9% ±6.336not given——not given
Claude Opus 4.7 [medium]Anthropic · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)31.6% ±4.032$2.3427k5515 Jun 2026
Kimi K2.7 CodeMoonshot AI · mini-swe-agent · DCDCDatacurve30.5% ±0.51$2.8259k14915 Jun 2026
Claude Sonnet 5 [low]Anthropic · mini-swe-agent · DCDCDatacurve30.5% ±1.131$2.1936k772 Jul 2026
Kimi K2.7 Code [high]Moonshot AI · mini-swe-agent · INDINDMercor30.1% ±8.436not given——not given
Claude Sonnet 4.6 [high]Anthropic · mini-swe-agent · DCDCDatacurve29.9% ±4.091$5.5276k13415 Jun 2026
Gemini 3.5 Flash [medium]Google · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)28.3% ±3.542$7.42189k7515 Jun 2026
Hy3Tencent · lab-internal · LABLABTencent28%15not given——27 Aug 2026
Claude Opus 4.6 [max]Anthropic · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)27.6% ±3.682$5.3944k10315 Jun 2026
GPT-5.5 [low]OpenAI · mini-swe-agent · DCDCDatacurve27% ±2.291$1.209.4k2815 Jun 2026
Grok 4.6 [xhigh]xAI · mini-swe-agent · INDINDMercor24.8% ±6.836not given——not given
GPT-5.4 mini [xhigh]OpenAI · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)24.3% ±2.962$2.08135k8615 Jun 2026
GPT-5.6 Terra [low]OpenAI · mini-swe-agent · DCDCDatacurve24.1% ±0.781$0.348.6k2210 Jul 2026
Kimi K2.6Moonshot AI · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)23.9% ±2.452$3.1684k14715 Jun 2026
Qwen3.7 MaxAlibaba · Claude Code · LABLABAlibaba (Qwen)21.6%11not given——8 Aug 2026
MiniMax M3MiniMax · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)20.4% ±3.732$5.5798k31415 Jun 2026
MiMo-V2.5-ProXiaomi · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)19.5% ±1.872$1.9949k12215 Jun 2026
GLM-5.1Z.ai · mini-swe-agent · INDINDMercor18.9% ±5.536not given——not given
Qwen3.7 MaxAlibaba · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)17.7% ±1.422$2.1242k11015 Jun 2026
GLM-5.1Z.ai · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)17.5% ±0.792$7.4649k17715 Jun 2026
MiniMax M3 [thinking]MiniMax · mini-swe-agent · INDINDentrpi (independent)16.8%34not given——2 Jun 2026
Qwen3.7 PlusAlibaba · mini-swe-agent · LABLABAlibaba (Qwen)16.5%13not given——24 Aug 2026
MiniMax M3 [thinking]MiniMax · mini-swe-agent · INDINDentrpi (independent)13.3%34$7.4880k3252 Jun 2026
Qwen3.6-27BAlibaba · Claude Code · LABLABAlibaba (Qwen)13.3%10not given——5 Aug 2026
Grok Build 0.1xAI · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)13.1% ±2.442$6.6052k17615 Jun 2026
Gemini 3.1 Pro (preview) [high]Google · mini-swe-agent · DCDCDatacurve11.7% ±1.481$2.1428k7615 Jun 2026
GPT-5.6 Luna [medium]OpenAI · mini-swe-agent · DCDCDatacurve11.3% ±0.831$0.0438.2k2410 Jul 2026
Gemini 3.1 Pro (preview)Google · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)9.7% ±2.832$1.8453k7415 Jun 2026
GPT-5.6 Luna [medium]OpenAI · lab-internal · LABLABOpenAI9.3%26$0.031——22 Sept 2026
DeepSeek V4 Pro [max]DeepSeek · mini-swe-agent · INDINDMercor9.1% ±3.836not given——not given
MiniMax M3 [high]MiniMax · mini-swe-agent · INDINDMercor9.1% ±3.136not given——not given
Gemini 3.1 Pro [high]Google · mini-swe-agent · INDINDMercor7.7% ±3.136not given——not given
DeepSeek V4 ProDeepSeek · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)7.5% ±2.72$4.2250k11115 Jun 2026
Gemini 3 Flash (preview)Google · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)5.1% ±2.482$1.53233k7115 Jun 2026
Inkling [high]Thinking Machines · mini-swe-agent · INDINDMercor5% ±2.836not given——not given
Qwen3.6 PlusAlibaba · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)2.6% ±1.232$4.2567k16415 Jun 2026
GPT-6 Luna [low]OpenAI · lab-internal · LABLABOpenAI2.4%26$0.006——22 Sept 2026
GLM-5Z.ai · mini-swe-agent · INDINDMercor2.1% ±1.936not given——not given
GPT-5.6 Luna [low]OpenAI · mini-swe-agent · DCDCDatacurve1.6% ±0.831$0.0153.1k1310 Jul 2026
GPT-5.6 Luna [low]OpenAI · lab-internal · LABLABOpenAI1.2%26$0.011——22 Sept 2026
DeepSeek V3.2DeepSeek · mini-swe-agent · INDINDMercor0.6% ±0.936not given——not given
Nemotron 3 Ultra [high]NVIDIA · mini-swe-agent · INDINDMercor0.6% ±0.736not given——not given
Claude Haiku 4.5Anthropic · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)0.2% ±0.432$0.8439k10915 Jun 2026
MiniMax M2.7MiniMax · mini-swe-agent · DCDCInternet Archive (capture of the Datacurve board)0.2% ±0.432$0.7060k13615 Jun 2026
Gemma 4 31BGoogle · mini-swe-agent · INDINDMercor0% ±036not given——not given
gpt-oss-120b [high]OpenAI · mini-swe-agent · INDINDMercor0% ±036not given——not given
Kimi K2 [high]Moonshot AI · mini-swe-agent · INDINDMercor0% ±036not given——not given
MiniMax M2.7 [high]MiniMax · mini-swe-agent · INDINDMercor0% ±036not given——not given
Qwen3.5Alibaba · mini-swe-agent · INDINDMercor0% ±036not given——not given
Fig. 3 · The spread

Same model, different measurer

The official leaderboard’s bar form, one bar per source under each model: its best official result, its lab’s claim and any independent re-run. Only models measured by at least two sources are shown.

Best DeepSWE 1.1 score per model, by who measured it40 models; for each, a bar for its best official, lab-claimed and independent pass@1 at any effort level.0%20%40%60%80%GPT-6.1 SolLab claim75.2%GPT-6.1 Sol [high] · Lab claim: 75.2%Independent run72.3%GPT-6.1 Sol [max] · Independent run: 72.3%Claude Opus 5.5Lab claim74.2%Claude Opus 5.5 [max] · Lab claim: 74.2%Independent run72.3%Claude Opus 5.5 [max] · Independent run: 72.3%DeepSeek V4.1 FlashLab claim74.2%DeepSeek V4.1 Flash [max] · Lab claim: 74.2%Independent run71.7%DeepSeek V4.1 Flash [max] · Independent run: 71.7%GPT-6 AstraOfficial board74.1%GPT-6 Astra [xhigh] · Official board: 74.1%Independent run72%GPT-6 Astra [max] · Independent run: 72%Gemini 3.8 FlashOfficial board73.8%Gemini 3.8 Flash [high] · Official board: 73.8%Lab claim73.7%Gemini 3.8 Flash [high] · Lab claim: 73.7%Independent run71.4%Gemini 3.8 Flash [high] · Independent run: 71.4%Claude Opus 5Official board73.7%Claude Opus 5 [max] · Official board: 73.7%Lab claim68.8%Claude Opus 5 [max] · Lab claim: 68.8%Independent run71.4%Claude Opus 5 [max] · Independent run: 71.4%Grok 4.7Lab claim71%Grok 4.7 [high] · Lab claim: 71%Independent run73%Grok 4.7 [xhigh] · Independent run: 73%GPT-5.6 SolOfficial board72.7%GPT-5.6 Sol [max] · Official board: 72.7%Independent run70.5%GPT-5.6 Sol [max] · Independent run: 70.5%Claude Sonnet 5.5Lab claim71%Claude Sonnet 5.5 [max] · Lab claim: 71%Independent run70.5%Claude Sonnet 5.5 [max] · Independent run: 70.5%GLM-5.3Official board69%GLM-5.3 [max] · Official board: 69%Lab claim66.9%GLM-5.3 [max] · Lab claim: 66.9%Independent run70.5%GLM-5.3 [max] · Independent run: 70.5%GPT-6 SolLab claim68.8%GPT-6 Sol [max] · Lab claim: 68.8%Independent run70.5%GPT-6 Sol [max] · Independent run: 70.5%GPT-5.6 TerraOfficial board69.6%GPT-5.6 Terra [max] · Official board: 69.6%Independent run70.2%GPT-5.6 Terra [max] · Independent run: 70.2%Claude Fable 5Official board69.9%Claude Fable 5 [xhigh] · Official board: 69.9%Independent run67.3%Claude Fable 5 [max] · Independent run: 67.3%Kimi K3Official board68.5%Kimi K3 [max] · Official board: 68.5%Lab claim67.5%Kimi K3 [max] · Lab claim: 67.5%Independent run66.4%Kimi K3 [max] · Independent run: 66.4%Grok 4.6Official board67.5%Grok 4.6 [medium] · Official board: 67.5%Lab claim65.9%Grok 4.6 [high] · Lab claim: 65.9%Independent run65%Grok 4.6 [high] · Independent run: 65%Claude Fable 5.1Lab claim67.4%Claude Fable 5.1 [max] · Lab claim: 67.4%Independent run67.3%Claude Fable 5.1 [high] · Independent run: 67.3%GPT-5.6 LunaOfficial board67.2%GPT-5.6 Luna [max] · Official board: 67.2%Lab claim62.2%GPT-5.6 Luna [max] · Lab claim: 62.2%Independent run64.3%GPT-5.6 Luna [max] · Independent run: 64.3%GPT-5.5Official board67%GPT-5.5 [xhigh] · Official board: 67%Independent run66.1%GPT-5.5 [xhigh] · Independent run: 66.1%GPT-6 LunaLab claim66.6%GPT-6 Luna [max] · Lab claim: 66.6%Independent run67%GPT-6 Luna [max] · Independent run: 67%Gemini 3.7 FlashOfficial board65.5%Gemini 3.7 Flash [medium] · Official board: 65.5%Independent run65.5%Gemini 3.7 Flash [high] · Independent run: 65.5%GLM-5.3 FlashOfficial board63.4%GLM-5.3 Flash [max] · Official board: 63.4%Independent run42.5%GLM-5.3 Flash [max] · Independent run: 42.5%DeepSeek V4 ProOfficial board62.8%DeepSeek V4 Pro [max] · Official board: 62.8%Lab claim62.7%DeepSeek V4 Pro [max] · Lab claim: 62.7%Independent run9.1%DeepSeek V4 Pro [max] · Independent run: 9.1%Claude Opus 4.8Official board59%Claude Opus 4.8 [max] · Official board: 59%Independent run59.6%Claude Opus 4.8 [max] · Independent run: 59.6%Qwen3.8 MaxOfficial board57.5%Qwen3.8 Max [xhigh] · Official board: 57.5%Lab claim56.6%Qwen3.8 Max · Lab claim: 56.6%Independent run53.1%Qwen3.8 Max [xhigh] · Independent run: 53.1%DeepSeek V4 FlashOfficial board53.3%DeepSeek V4 Flash [max] · Official board: 53.3%Lab claim54.4%DeepSeek V4 Flash [max] · Lab claim: 54.4%Independent run56.6%DeepSeek V4 Flash [max] · Independent run: 56.6%Claude Opus 4.7Official board54.2%Claude Opus 4.7 [max] · Official board: 54.2%Independent run49.3%Claude Opus 4.7 [max] · Independent run: 49.3%Claude Sonnet 5Official board53.9%Claude Sonnet 5 [max] · Official board: 53.9%Independent run49.6%Claude Sonnet 5 [max] · Independent run: 49.6%Grok 4.5Official board53.8%Grok 4.5 [high] · Official board: 53.8%Independent run49.3%Grok 4.5 [high] · Independent run: 49.3%Muse Spark 1.1Official board53.3%Muse Spark 1.1 [xhigh] · Official board: 53.3%Independent run40.4%Muse Spark 1.1 [xhigh] · Independent run: 40.4%GPT-5.4Official board51.8%GPT-5.4 [xhigh] · Official board: 51.8%Independent run49%GPT-5.4 [xhigh] · Independent run: 49%Gemini 3.6 FlashOfficial board46.7%Gemini 3.6 Flash [high] · Official board: 46.7%Independent run49.3%Gemini 3.6 Flash [high] · Independent run: 49.3%GLM-5.2Official board43.8%GLM-5.2 [max] · Official board: 43.8%Lab claim46.2%GLM-5.2 [max] · Lab claim: 46.2%Independent run38.9%GLM-5.2 [max] · Independent run: 38.9%Gemini 3.5 FlashOfficial board36.1%Gemini 3.5 Flash [high] · Official board: 36.1%Independent run36.6%Gemini 3.5 Flash [high] · Independent run: 36.6%Claude Sonnet 4.6Official board29.9%Claude Sonnet 4.6 [high] · Official board: 29.9%Independent run34.2%Claude Sonnet 4.6 [high] · Independent run: 34.2%Claude Opus 4.6Official board27.6%Claude Opus 4.6 [max] · Official board: 27.6%Independent run31.9%Claude Opus 4.6 [max] · Independent run: 31.9%Kimi K2.7 CodeOfficial board30.5%Kimi K2.7 Code · Official board: 30.5%Independent run30.1%Kimi K2.7 Code [high] · Independent run: 30.1%Qwen3.7 MaxOfficial board17.7%Qwen3.7 Max · Official board: 17.7%Lab claim21.6%Qwen3.7 Max · Lab claim: 21.6%MiniMax M3Official board20.4%MiniMax M3 · Official board: 20.4%Independent run16.8%MiniMax M3 [thinking] · Independent run: 16.8%GLM-5.1Official board17.5%GLM-5.1 · Official board: 17.5%Independent run18.9%GLM-5.1 · Independent run: 18.9%MiniMax M2.7Official board0.2%MiniMax M2.7 · Official board: 0.2%Independent run0%MiniMax M2.7 [high] · Independent run: 0%
Notes on the readings

How to read these numbers

  1. The official board has added no model since 3 Sep 2026 (GPT-6 Astra; its runs finished on 1 Sep).3 Its data file was regenerated on 22 Sep with no new runs.5 A GitHub issue on the benchmark's repository quotes Datacurve's CEO saying the team is building the next version.39
  2. Epoch AI's review of 7 Sep 2026 rated DeepSWE 1.1 'flawed', finding grading defects in 23 of 113 tasks, mostly hidden tests colliding with tests the agent wrote.37 Tokenless separately documented 70 ways a submission can rewrite test outcomes.38 Treat differences of a few points with care.
  3. Lab claims are run by the lab, on its own harness, effort settings and trial count, so they are not directly comparable with the official runs (mini-swe-agent, four passes over 113 tasks).4 The harness is listed on every row. xAI says Datacurve ran the Grok 4.7 evaluation for it, though the result never appeared on the public board.23
  4. Independent runs come from Mercor (mini-swe-agent, 500 steps, 2-hour limit, 3 passes),36 Artificial Analysis (Grok Build harness),35 Fireworks (Kimi K3 at three effort levels)27 and entrpi (MiniMax M3).34
  5. Costs are per task in USD. The official page and its data file disagree for 19 of 70 configurations; this tracker shows the page's figure and keeps the file's in the row notes.5 Lab costs are the lab's own, shown only where the lab published them.
  6. Mercor shows no date for each model, so its results appear in the charts and table but not on the timeline.36
  7. A number is included only when it was read at the source that published it. Figures that appear only on aggregator sites are left out; research/ingest-log.txt in the repository lists each one and why.
  8. Anthropic has published no DeepSWE 1.1 score for Claude Haiku 5.5;33 Mercor's independent run is the only one. There is no GPT-6.1 Luna: OpenAI released GPT-6.1 Sol only.30
  9. Where a source published token counts but no cost, the cost is estimated from Vercel AI Gateway’s list prices and shown as a range marked est.40 With only a total token count, the range runs from every token priced as input to every token priced as output; cached input is cheaper still, so the true cost may sit below it. Only Laguna S 2.1 qualifies today.
Methods

How this was made

What was gathered
On 9 Oct 2026, Claude (Anthropic’s model) ran Dossier’s free local research loop with its own web search: six search tasks covering the official board, the US labs, Chinese and open-weight labs and Mistral, aggregator sites, and community and press coverage. A separate fact-check pass then settled the conflicts between them. No paid research API was used.
What was kept
The sweep returned 605 raw figures. 210 readings of 80 models survived the rule that a number is read at the source that published it, and the page cites 40 sources in the list below. Official rows are re-read from the board’s own page every day.
What was verified
Every lab figure was re-read on the lab’s own page or PDF. Three conflicts were settled against the primary document: GPT-6.1 Sol’s xhigh and max scores, who published Grok 4.7’s 73%, and the harness behind Muse Spark 1.3. Of the 22 sources the research report cites, none were fabricated; two OpenAI pages block automated reads and one StepFun page did not load.
What could not be established
When each model joined Mercor’s board; cost per task for Anthropic, Google and Meta models, which those labs do not publish; a dated StepFun page for Step 5 Preview; and why Datacurve ran Grok 4.7 without posting the result. No person has yet re-checked every row; corrections are welcome as issues on the repository.
Sources

Every original source, numbered

40 sources. Each entry says what it establishes and links back to the readings that cite it. The same list is in SOURCES.md, and every row of the dataset carries its source URL and the line quoted from it.

  1. Datacurve DeepSWE 1.1 leaderboardDatacurve · Official leaderboard · published 15 Jun 2026 · read 9 Oct 2026Pass@1 for GPT-6 Astra 67%–74.1% across 5 settings; Gemini 3.8 Flash 71%–73.8% across 2 settings; Claude Opus 5 58.1%–73.7% across 5 settings; GPT-5.6 Sol 45.4%–72.7% across 5 settings; Claude Fable 5 59.6%–69.9% across 5 settings; GPT-5.6 Terra 24.1%–69.6% across 5 settings; and 22 more models, with cost per task, tokens, confidence intervals.Cited by 70 readings: GPT-6 Astra [xhigh], Gemini 3.8 Flash [high], Claude Opus 5 [max], GPT-6 Astra [high], GPT-6 Astra [max], Claude Opus 5 [xhigh], Claude Opus 5 [high], GPT-6 Astra [medium], GPT-5.6 Sol [max], Gemini 3.8 Flash [medium], GPT-5.6 Sol [xhigh], Claude Fable 5 [xhigh],
    and 58 moreClaude Fable 5 [max], GPT-5.6 Terra [max], GPT-5.6 Sol [high], GLM-5.3 [max], Claude Opus 5 [medium], Claude Fable 5 [high], Kimi K3 [max], Grok 4.6 [medium], GPT-5.6 Luna [max], GPT-5.5 [xhigh], GPT-6 Astra [low], Grok 4.6 [xhigh], Gemini 3.7 Flash [medium], Claude Fable 5 [medium], Gemini 3.7 Flash [high], Grok 4.6 [high], GPT-5.5 [high], GLM-5.3 Flash [max], DeepSeek V4 Pro [max], GPT-5.6 Sol [medium], GPT-5.6 Terra [xhigh], Claude Fable 5 [low], Claude Opus 4.8 [max], Claude Opus 5 [low], Qwen3.8 Max [xhigh], GPT-5.6 Luna [xhigh], Muse Spark 1.2 [xhigh], Claude Opus 4.8 [xhigh], GPT-5.5 [medium], Claude Sonnet 5 [max], Gemini 3.7 Flash [low], GPT-5.6 Terra [high], Grok 4.5 [high], DeepSeek V4 Flash [max], Muse Spark 1.1 [xhigh], Claude Opus 4.8 [high], GPT-5.4 [xhigh], Claude Sonnet 5 [xhigh], Claude Opus 4.8 [medium], Claude Sonnet 5 [high], Gemini 3.6 Flash [high], GPT-5.6 Sol [low], GPT-5.6 Luna [high], GLM-5.2 [max], Grok 4.6 [low], Claude Opus 4.8 [low], Claude Sonnet 5 [medium], GLM-5.2 [high], Gemini 3.5 Flash [high], GPT-5.6 Terra [medium], Kimi K2.7 Code, Claude Sonnet 5 [low], Claude Sonnet 4.6 [high], GPT-5.5 [low], GPT-5.6 Terra [low], Gemini 3.1 Pro (preview) [high], GPT-5.6 Luna [medium], GPT-5.6 Luna [low]
  2. Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)Internet Archive (capture of the Datacurve board) · Official leaderboard · published 15 Jun 2026 · read 9 Oct 2026Pass@1 for Claude Opus 4.7 31.6%–54.2% across 4 settings; Gemini 3.5 Flash 28.3% [medium]; Claude Opus 4.6 27.6% [max]; GPT-5.4 mini 24.3% [xhigh]; Kimi K2.6 23.9%; MiniMax M3 20.4%; and 10 more models, with cost per task, tokens, confidence intervals.Cited by 19 readings: Claude Opus 4.7 [max], Claude Opus 4.7 [xhigh], Claude Opus 4.7 [high], Claude Opus 4.7 [medium], Gemini 3.5 Flash [medium], Claude Opus 4.6 [max], GPT-5.4 mini [xhigh], Kimi K2.6, MiniMax M3, MiMo-V2.5-Pro, Qwen3.7 Max, GLM-5.1,
    and 7 moreGrok Build 0.1, Gemini 3.1 Pro (preview), DeepSeek V4 Pro, Gemini 3 Flash (preview), Qwen3.6 Plus, Claude Haiku 4.5, MiniMax M2.7
  3. DeepSWE changelogDatacurve · Official changelog · published 3 Sept 2026 · read 9 Oct 2026Lists each model addition by date; the most recent is GPT-6 Astra (all efforts) on 3 Sep 2026.Cited in the notes above.
  4. Run DeepSWEDatacurve · Official documentation · published 15 Jun 2026 · read 9 Oct 2026Official scores are produced with Pier running mini-swe-agent on Modal, in isolated containers.Cited in the notes above.
  5. leaderboard-live.json (v1.1 data file)Datacurve · Official data file · published 22 Sept 2026 · read 9 Oct 2026The board's data file: generated 22 Sep 2026, latest job finished 1 Sep; its costs differ from the page's for 19 of 70 configurations.Cited in the notes above.
  6. GLM-5.2 HuggingFace Model CardZ.ai · Lab self-report · published 16 Jun 2026 · read 9 Oct 2026Pass@1 for GLM-5.2 46.2% [max].Cited by 1 reading: GLM-5.2 [max]
  7. Poolside Laguna S 2.1 launch postPoolside · Lab self-report · published 21 Jul 2026 · read 9 Oct 2026Pass@1 for Laguna S 2.1 40.4% [max].Cited by 1 reading: Laguna S 2.1 [max]
  8. Kimi-K3 HuggingFace Model CardMoonshot AI · Lab self-report · published 23 Jul 2026 · read 9 Oct 2026Pass@1 for Kimi K3 67.5% [max].Cited by 1 reading: Kimi K3 [max]
  9. Claude Opus 5 System CardAnthropic · Lab self-report · published 24 Jul 2026 · read 9 Oct 2026Pass@1 for Claude Opus 5 68.8% [max].Cited by 1 reading: Claude Opus 5 [max]
  10. Qwen3.8-27B HuggingFace Model CardAlibaba (Qwen) · Lab self-report · published 5 Aug 2026 · read 9 Oct 2026Pass@1 for Qwen3.8-27B 42.2%; Qwen3.6-27B 13.3%.Cited by 2 readings: Qwen3.8-27B, Qwen3.6-27B
  11. Qwen3.8-2.4T-A95B (Qwen3.8-Max) HuggingFace Model CardAlibaba (Qwen) · Lab self-report · published 8 Aug 2026 · read 9 Oct 2026Pass@1 for Qwen3.8 Max 56.6%; Qwen3.7 Max 21.6%.Cited by 2 readings: Qwen3.8 Max, Qwen3.7 Max
  12. xAI Grok 4.6 Model CardxAI · Lab self-report · published 18 Aug 2026 · read 9 Oct 2026Pass@1 for Grok 4.6 65.9% [high].Cited by 1 reading: Grok 4.6 [high]
  13. Qwen3.8-Flash-Next HuggingFace Model CardAlibaba (Qwen) · Lab self-report · published 24 Aug 2026 · read 9 Oct 2026Pass@1 for Qwen3.8-Flash-Next 58.7%; Qwen3.7 Plus 16.5%.Cited by 2 readings: Qwen3.8-Flash-Next, Qwen3.7 Plus
  14. GLM-5.3 HuggingFace Model CardZ.ai · Lab self-report · published 25 Aug 2026 · read 9 Oct 2026Pass@1 for GLM-5.3 66.9% [max].Cited by 1 reading: GLM-5.3 [max]
  15. Tencent Hy4-preview Technical Report & Model CardTencent · Lab self-report · published 27 Aug 2026 · read 9 Oct 2026Pass@1 for Hy4 Preview 64.3%; Hy3 28%.Cited by 2 readings: Hy4 Preview, Hy3
  16. Claude Fable 5.1 & Claude Mythos 5.1 System CardAnthropic · Lab self-report · published 1 Sept 2026 · read 9 Oct 2026Pass@1 for Claude Fable 5.1 67.4% [max].Cited by 1 reading: Claude Fable 5.1 [max]
  17. Meta AI Research Muse Spark 1.3 Evaluation MethodologyMeta · Lab self-report · published 2 Sept 2026 · read 9 Oct 2026Pass@1 for Muse Spark 1.3 75.4% [max].Cited by 1 reading: Muse Spark 1.3 [max]
  18. Nex-AGI Nex-N2.5 GitHub README / Model CardNex-AGI · Lab self-report · published 8 Sept 2026 · read 9 Oct 2026Pass@1 for Nex-N2.5-Max 65.6%; Nex-N2.5-Pro 55.8%; Nex-N2.5-Mini 36.1%.Cited by 3 readings: Nex-N2.5-Max, Nex-N2.5-Pro, Nex-N2.5-Mini
  19. Cognition SWE-2 launch postCognition · Lab self-report · published 10 Sept 2026 · read 9 Oct 2026Pass@1 for SWE-2 73%; SWE-1.7 37.7%.Cited by 2 readings: SWE-2, SWE-1.7
  20. DeepSeek-V4.1-Flash HuggingFace Model CardDeepSeek · Lab self-report · published 10 Sept 2026 · read 9 Oct 2026Pass@1 for DeepSeek V4.1 Flash 65.5%–74.2% across 8 settings; DeepSeek V4 Pro 62.7% [max]; DeepSeek V4 Flash 54.4% [max].Cited by 10 readings: DeepSeek V4.1 Flash [max], DeepSeek V4.1 Flash [max], DeepSeek V4.1 Flash [max], DeepSeek V4.1 Flash [max], DeepSeek V4.1 Flash [max], DeepSeek V4.1 Flash [max], DeepSeek V4.1 Flash [max], DeepSeek V4.1 Flash [max], DeepSeek V4 Pro [max], DeepSeek V4 Flash [max]
  21. Google Gemini 3.8 Flash launch blog and Model Evaluation ReportGoogle DeepMind · Lab self-report · published 17 Sept 2026 · read 9 Oct 2026Pass@1 for Gemini 3.8 Flash 73.7% [high], with cost per task, tokens.Cited by 1 reading: Gemini 3.8 Flash [high]
  22. StepFun Step 5 Preview AnnouncementStepFun · Lab self-report · published 20 Sept 2026 · read 9 Oct 2026Pass@1 for Step 5 Preview 67.7% [high].Cited by 1 reading: Step 5 Preview [high]
  23. xAI Grok 4.7 Model CardxAI · Lab self-report · published 21 Sept 2026 · read 9 Oct 2026Pass@1 for Grok 4.7 71% [high].Cited by 1 reading: Grok 4.7 [high]
  24. Xiaomi MiMo-V2.6 Launch Announcement and AppendixXiaomi · Lab self-report · published 21 Sept 2026 · read 9 Oct 2026Pass@1 for MiMo-V2.6-Pro 71.9% [max]; MiMo-V2.6-Flash 67.9% [max].Cited by 2 readings: MiMo-V2.6-Pro [max], MiMo-V2.6-Flash [max]
  25. Claude Opus 5.5 System CardAnthropic · Lab self-report · published 22 Sept 2026 · read 9 Oct 2026Pass@1 for Claude Opus 5.5 74.2% [max].Cited by 1 reading: Claude Opus 5.5 [max]
  26. OpenAI GPT-6 Sol and Luna launch postOpenAI · Lab self-report · published 22 Sept 2026 · read 9 Oct 2026Pass@1 for GPT-6 Sol 37.2%–68.8% across 5 settings; GPT-6 Luna 2.4%–66.6% across 5 settings; GPT-5.6 Luna 1.2%–62.2% across 5 settings, with cost per task.Cited by 15 readings: GPT-6 Sol [max], GPT-6 Luna [max], GPT-6 Sol [xhigh], GPT-6 Sol [high], GPT-5.6 Luna [max], GPT-6 Luna [xhigh], GPT-6 Luna [high], GPT-6 Sol [medium], GPT-5.6 Luna [xhigh], GPT-6 Luna [medium], GPT-5.6 Luna [high], GPT-6 Sol [low],
    and 3 moreGPT-5.6 Luna [medium], GPT-6 Luna [low], GPT-5.6 Luna [low]
  27. Fireworks AI Ember-1 AnnouncementFireworks AI · Lab self-report · published 23 Sept 2026 · read 9 Oct 2026Pass@1 for Ember-1 75.2% [thinking]; Kimi K3 55.8%–66.4% across 3 settings, with cost per task.Cited by 4 readings: Ember-1 [thinking], Kimi K3 [max], Kimi K3 [high], Kimi K3 [low]
  28. Gemini 4 Argon Model Evaluation ReportGoogle DeepMind · Lab self-report · published 24 Sept 2026 · read 9 Oct 2026Pass@1 for Gemini 4 Argon 77.9%.Cited by 1 reading: Gemini 4 Argon
  29. Claude Sonnet 5.5 System CardAnthropic · Lab self-report · published 28 Sept 2026 · read 9 Oct 2026Pass@1 for Claude Sonnet 5.5 71% [max].Cited by 1 reading: Claude Sonnet 5.5 [max]
  30. OpenAI GPT-6.1 Sol launch postOpenAI · Lab self-report · published 29 Sept 2026 · read 9 Oct 2026Pass@1 for GPT-6.1 Sol 64.4%–75.2% across 5 settings, with cost per task.Cited by 5 readings: GPT-6.1 Sol [high], GPT-6.1 Sol [medium], GPT-6.1 Sol [max], GPT-6.1 Sol [xhigh], GPT-6.1 Sol [low]
  31. Reflection AI Beam launch postReflection AI · Lab self-report · published 5 Oct 2026 · read 9 Oct 2026Pass@1 for Beam 44.4%.Cited by 1 reading: Beam
  32. Mistral Large 4 AnnouncementMistral AI · Lab self-report · published 6 Oct 2026 · read 9 Oct 2026Pass@1 for Mistral Large 4 61.7% [thinking].Cited by 1 reading: Mistral Large 4 [thinking]
  33. Claude Haiku 5.5 System CardAnthropic · Lab system card · published 29 Sept 2026 · read 9 Oct 2026Reports SWE-bench Pro, FrontierSWE v2 and ProgramBench for Haiku 5.5, and no DeepSWE 1.1 result.Cited in the notes above.
  34. Independent DeepSWE Audit by entrpientrpi (independent) · Independent run · published 2 Jun 2026 · read 9 Oct 2026Pass@1 for MiniMax M3 13.3%–16.8% across 2 settings, with cost per task, tokens.Cited by 2 readings: MiniMax M3 [thinking], MiniMax M3 [thinking]
  35. Artificial Analysis (Benchmarking Grok 4.7)Artificial Analysis · Independent run · published 21 Sept 2026 · read 9 Oct 2026Pass@1 for Grok 4.7 73% [xhigh]; Grok 4.6 65% [high].Cited by 2 readings: Grok 4.7 [xhigh], Grok 4.6 [high]
  36. Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)Mercor · Independent run · published undated · read 9 Oct 2026Pass@1 for Claude Opus 5.5 72.3% [max]; GPT-6.1 Sol 72.3% [max]; GPT-6 Astra 72% [max]; DeepSeek V4.1 Flash 71.7% [max]; Claude Opus 5 70.2%–71.4% across 2 settings; Gemini 3.8 Flash 71.4% [high]; and 43 more models, with confidence intervals.Cited by 52 readings: Claude Opus 5.5 [max], GPT-6.1 Sol [max], GPT-6 Astra [max], DeepSeek V4.1 Flash [max], Claude Opus 5 [max], Gemini 3.8 Flash [high], Claude Sonnet 5.5 [max], GLM-5.3 [max], GPT-5.6 Sol [max], GPT-6 Sol [max], Claude Opus 5 [xhigh], GPT-5.6 Terra [max],
    and 40 moreClaude Fable 5 [max], Claude Fable 5.1 [high], GPT-6 Luna [max], GPT-5.5 [xhigh], Gemini 3.7 Flash [high], Claude Fable 5.1 [max], GPT-5.6 Luna [max], Grok 4.6 [high], Claude Haiku 5.5 [max], Claude Opus 4.8 [max], DeepSeek V4 Flash [max], DeepSeek V4 Pro 0813 [max], Qwen3.8 Max [xhigh], Claude Sonnet 5 [max], Claude Opus 4.7 [max], Gemini 3.6 Flash [high], Grok 4.5 [high], GPT-5.4 [xhigh], GLM-5.3 Flash [max], Muse Spark 1.1 [xhigh], GLM-5.2 [max], Gemini 3.5 Flash [high], Claude Sonnet 4.6 [high], Grok 4.7 [xhigh], Claude Opus 4.6 [max], Kimi K2.7 Code [high], Grok 4.6 [xhigh], GLM-5.1, DeepSeek V4 Pro [max], MiniMax M3 [high], Gemini 3.1 Pro [high], Inkling [high], GLM-5, DeepSeek V3.2, Nemotron 3 Ultra [high], Gemma 4 31B, gpt-oss-120b [high], Kimi K2 [high], MiniMax M2.7 [high], Qwen3.5
  37. Benchmark review: DeepSWE v1.1Epoch AI · Independent audit · published 7 Sept 2026 · read 9 Oct 2026Rates DeepSWE v1.1 flawed: grading defects in at least 23 of 113 tasks, most from hidden tests colliding with tests the agent wrote.Cited in the notes above.
  38. envcheck: auditing DeepSWETokenless · Independent audit · published 30 Sept 2026 · read 9 Oct 2026Documents 70 ways a submission can rewrite test outcomes, because the test harness sits inside the repository the agent edits.Cited in the notes above.
  39. datacurve-ai/deep-swe issue #103GitHub (community discussion) · Discussion · published 30 Sept 2026 · read 9 Oct 2026Notes the board has published no new model since 3 Sep and quotes Datacurve's CEO saying the team is building the next generation of DeepSWE.Cited in the notes above.
  40. AI Gateway model list and pricesVercel · Price list · published undated · read 9 Oct 2026Per-token list prices used to estimate cost where a source published tokens but no cost: Laguna S 2.1 at $0.09 per million input tokens and $0.18 per million output tokens.Cited in the notes above.