{
  "benchmark": "DeepSWE 1.1",
  "description": "Pass@1 on Datacurve DeepSWE 1.1 (113 long-horizon engineering tasks), from the official leaderboard, lab reports and independent runs.",
  "generated_from": [
    "data/models.csv",
    "data/deepswe-1.1.csv",
    "data/meta.json"
  ],
  "meta": {
    "as_of": "2026-10-09",
    "official_url": "https://deepswe.datacurve.ai/",
    "official_last_added": "2026-09-03",
    "official_last_added_basis": "Datacurve changelog: GPT-6 Astra (all efforts) added 2026-09-03; its trials finished 2026-09-01. The 2026-09-22 generated_at is a site re-export with no new runs.",
    "notes": [
      "The official board has added no model since 3 Sep 2026 (GPT-6 Astra; its runs finished on 1 Sep).{{datacurve-changelog}} Its data file was regenerated on 22 Sep with no new runs.{{datacurve-artifact}} A GitHub issue on the benchmark's repository quotes Datacurve's CEO saying the team is building the next version.{{gh-issue-103}}",
      "Epoch AI's review of 7 Sep 2026 rated DeepSWE 1.1 'flawed', finding grading defects in 23 of 113 tasks, mostly hidden tests colliding with tests the agent wrote.{{epoch-review}} Tokenless separately documented 70 ways a submission can rewrite test outcomes.{{tokenless-envcheck}} Treat differences of a few points with care.",
      "Lab claims are run by the lab, on its own harness, effort settings and trial count, so they are not directly comparable with the official runs (mini-swe-agent, four passes over 113 tasks).{{datacurve-run}} The harness is listed on every row. xAI says Datacurve ran the Grok 4.7 evaluation for it, though the result never appeared on the public board.{{url:https://media.x.ai/v1/website/card4p7-3a96f40b.pdf}}",
      "Independent runs come from Mercor (mini-swe-agent, 500 steps, 2-hour limit, 3 passes),{{url:https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/}} Artificial Analysis (Grok Build harness),{{url:https://artificialanalysis.ai/articles/benchmarking-grok-4-7}} Fireworks (Kimi K3 at three effort levels){{url:https://fireworks.ai/blog/ember-1}} and entrpi (MiniMax M3).{{url:https://entrpi.github.io/misc/deep-swe-minimax-m3/}}",
      "Costs are per task in USD. The official page and its data file disagree for 19 of 70 configurations; this tracker shows the page's figure and keeps the file's in the row notes.{{datacurve-artifact}} Lab costs are the lab's own, shown only where the lab published them.",
      "Mercor shows no date for each model, so its results appear in the charts and table but not on the timeline.{{url:https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/}}",
      "A number is included only when it was read at the source that published it. Figures that appear only on aggregator sites are left out; research/ingest-log.txt in the repository lists each one and why.",
      "Anthropic has published no DeepSWE 1.1 score for Claude Haiku 5.5;{{anthropic-haiku-55}} Mercor's independent run is the only one. There is no GPT-6.1 Luna: OpenAI released GPT-6.1 Sol only.{{url:https://openai.com/index/introducing-gpt-6-1-sol/}}",
      "Where a source published token counts but no cost, the cost is estimated from Vercel AI Gateway’s list prices and shown as a range marked est.{{ai-gateway-pricing}} With only a total token count, the range runs from every token priced as input to every token priced as output; cached input is cheaper still, so the true cost may sit below it. Only Laguna S 2.1 qualifies today."
    ],
    "official_generated_at": "2026-09-22T06:27:15.860279+00:00",
    "official_latest_job": {
      "name": "20260901-deep-swe-1-1-gpt-6-astra",
      "finished_at": "2026-09-01T07:35:13Z"
    },
    "official_configs": 70,
    "official_checked": "2026-10-09",
    "official_rows_from": "page",
    "context_sources": [
      {
        "id": "datacurve-changelog",
        "title": "DeepSWE changelog",
        "publisher": "Datacurve",
        "type": "Official changelog",
        "url": "https://deepswe.datacurve.ai/changelog",
        "published": "2026-09-03",
        "accessed": "2026-10-09",
        "establishes": "Lists each model addition by date; the most recent is GPT-6 Astra (all efforts) on 3 Sep 2026."
      },
      {
        "id": "datacurve-run",
        "title": "Run DeepSWE",
        "publisher": "Datacurve",
        "type": "Official documentation",
        "url": "https://deepswe.datacurve.ai/run",
        "published": "2026-06-15",
        "accessed": "2026-10-09",
        "establishes": "Official scores are produced with Pier running mini-swe-agent on Modal, in isolated containers."
      },
      {
        "id": "datacurve-artifact",
        "title": "leaderboard-live.json (v1.1 data file)",
        "publisher": "Datacurve",
        "type": "Official data file",
        "url": "https://deepswe.datacurve.ai/artifacts/v1.1/leaderboard-live.json",
        "published": "2026-09-22",
        "accessed": "2026-10-09",
        "establishes": "The board's data file: generated 22 Sep 2026, latest job finished 1 Sep; its costs differ from the page's for 19 of 70 configurations."
      },
      {
        "id": "epoch-review",
        "title": "Benchmark review: DeepSWE v1.1",
        "publisher": "Epoch AI",
        "type": "Independent audit",
        "url": "https://epoch.ai/benchmarks/deepswe/review",
        "published": "2026-09-07",
        "accessed": "2026-10-09",
        "establishes": "Rates DeepSWE v1.1 flawed: grading defects in at least 23 of 113 tasks, most from hidden tests colliding with tests the agent wrote."
      },
      {
        "id": "tokenless-envcheck",
        "title": "envcheck: auditing DeepSWE",
        "publisher": "Tokenless",
        "type": "Independent audit",
        "url": "https://usetokenless.com/blog/envcheck",
        "published": "2026-09-30",
        "accessed": "2026-10-09",
        "establishes": "Documents 70 ways a submission can rewrite test outcomes, because the test harness sits inside the repository the agent edits."
      },
      {
        "id": "gh-issue-103",
        "title": "datacurve-ai/deep-swe issue #103",
        "publisher": "GitHub (community discussion)",
        "type": "Discussion",
        "url": "https://github.com/datacurve-ai/deep-swe/issues/103",
        "published": "2026-09-30",
        "accessed": "2026-10-09",
        "establishes": "Notes the board has published no new model since 3 Sep and quotes Datacurve's CEO saying the team is building the next generation of DeepSWE."
      },
      {
        "id": "anthropic-haiku-55",
        "title": "Claude Haiku 5.5 System Card",
        "publisher": "Anthropic",
        "type": "Lab system card",
        "url": "https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf",
        "published": "2026-09-29",
        "accessed": "2026-10-09",
        "establishes": "Reports SWE-bench Pro, FrontierSWE v2 and ProgramBench for Haiku 5.5, and no DeepSWE 1.1 result."
      },
      {
        "id": "ai-gateway-pricing",
        "title": "AI Gateway model list and prices",
        "publisher": "Vercel",
        "type": "Price list",
        "url": "https://ai-gateway.vercel.sh/v1/models",
        "published": "",
        "accessed": "2026-10-09",
        "establishes": "Per-token list prices used to estimate cost where a source published tokens but no cost: Laguna S 2.1 at $0.09 per million input tokens and $0.18 per million output tokens."
      }
    ]
  },
  "source_types": {
    "official_leaderboard": "Official Datacurve leaderboard",
    "lab_self_reported": "Reported by the model’s lab",
    "third_party_run": "Independent third-party run"
  },
  "models": 80,
  "observations": 210,
  "sources": [
    {
      "id": "r-16wwn5l",
      "url": "https://deepswe.datacurve.ai/",
      "title": "Datacurve DeepSWE 1.1 leaderboard",
      "publisher": "Datacurve",
      "type": "Official leaderboard",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for GPT-6 Astra 67%–74.1% across 5 settings; Gemini 3.8 Flash 71%–73.8% across 2 settings; Claude Opus 5 58.1%–73.7% across 5 settings; GPT-5.6 Sol 45.4%–72.7% across 5 settings; Claude Fable 5 59.6%–69.9% across 5 settings; GPT-5.6 Terra 24.1%–69.6% across 5 settings; and 22 more models, with cost per task, tokens, confidence intervals.",
      "readings": 70,
      "n": 1
    },
    {
      "id": "r-1fnts5d",
      "url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "title": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "publisher": "Internet Archive (capture of the Datacurve board)",
      "type": "Official leaderboard",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Claude Opus 4.7 31.6%–54.2% across 4 settings; Gemini 3.5 Flash 28.3% [medium]; Claude Opus 4.6 27.6% [max]; GPT-5.4 mini 24.3% [xhigh]; Kimi K2.6 23.9%; MiniMax M3 20.4%; and 10 more models, with cost per task, tokens, confidence intervals.",
      "readings": 19,
      "n": 2
    },
    {
      "id": "r-glshid",
      "title": "DeepSWE changelog",
      "publisher": "Datacurve",
      "type": "Official changelog",
      "url": "https://deepswe.datacurve.ai/changelog",
      "published": "2026-09-03",
      "accessed": "2026-10-09",
      "establishes": "Lists each model addition by date; the most recent is GPT-6 Astra (all efforts) on 3 Sep 2026.",
      "key": "datacurve-changelog",
      "readings": 0,
      "n": 3
    },
    {
      "id": "r-1dj68fa",
      "title": "Run DeepSWE",
      "publisher": "Datacurve",
      "type": "Official documentation",
      "url": "https://deepswe.datacurve.ai/run",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "establishes": "Official scores are produced with Pier running mini-swe-agent on Modal, in isolated containers.",
      "key": "datacurve-run",
      "readings": 0,
      "n": 4
    },
    {
      "id": "r-qec9fu",
      "title": "leaderboard-live.json (v1.1 data file)",
      "publisher": "Datacurve",
      "type": "Official data file",
      "url": "https://deepswe.datacurve.ai/artifacts/v1.1/leaderboard-live.json",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "establishes": "The board's data file: generated 22 Sep 2026, latest job finished 1 Sep; its costs differ from the page's for 19 of 70 configurations.",
      "key": "datacurve-artifact",
      "readings": 0,
      "n": 5
    },
    {
      "id": "r-11h0zp7",
      "url": "https://huggingface.co/zai-org/GLM-5.2",
      "title": "GLM-5.2 HuggingFace Model Card",
      "publisher": "Z.ai",
      "type": "Lab self-report",
      "published": "2026-06-16",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for GLM-5.2 46.2% [max].",
      "readings": 1,
      "n": 6
    },
    {
      "id": "r-12e8ru1",
      "url": "https://poolside.ai/blog/introducing-laguna-s-2-1",
      "title": "Poolside Laguna S 2.1 launch post",
      "publisher": "Poolside",
      "type": "Lab self-report",
      "published": "2026-07-21",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Laguna S 2.1 40.4% [max].",
      "readings": 1,
      "n": 7
    },
    {
      "id": "r-kyj4bu",
      "url": "https://huggingface.co/moonshotai/Kimi-K3",
      "title": "Kimi-K3 HuggingFace Model Card",
      "publisher": "Moonshot AI",
      "type": "Lab self-report",
      "published": "2026-07-23",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Kimi K3 67.5% [max].",
      "readings": 1,
      "n": 8
    },
    {
      "id": "r-l1zffg",
      "url": "https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf",
      "title": "Claude Opus 5 System Card",
      "publisher": "Anthropic",
      "type": "Lab self-report",
      "published": "2026-07-24",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Claude Opus 5 68.8% [max].",
      "readings": 1,
      "n": 9
    },
    {
      "id": "r-190dm7p",
      "url": "https://huggingface.co/Qwen/Qwen3.8-27B",
      "title": "Qwen3.8-27B HuggingFace Model Card",
      "publisher": "Alibaba (Qwen)",
      "type": "Lab self-report",
      "published": "2026-08-05",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Qwen3.8-27B 42.2%; Qwen3.6-27B 13.3%.",
      "readings": 2,
      "n": 10
    },
    {
      "id": "r-ouk7wk",
      "url": "https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B",
      "title": "Qwen3.8-2.4T-A95B (Qwen3.8-Max) HuggingFace Model Card",
      "publisher": "Alibaba (Qwen)",
      "type": "Lab self-report",
      "published": "2026-08-08",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Qwen3.8 Max 56.6%; Qwen3.7 Max 21.6%.",
      "readings": 2,
      "n": 11
    },
    {
      "id": "r-tt3io3",
      "url": "https://media.x.ai/v1/website/card-4p6-4cd2dc57.pdf",
      "title": "xAI Grok 4.6 Model Card",
      "publisher": "xAI",
      "type": "Lab self-report",
      "published": "2026-08-18",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Grok 4.6 65.9% [high].",
      "readings": 1,
      "n": 12
    },
    {
      "id": "r-qxom7y",
      "url": "https://huggingface.co/Qwen/Qwen3.8-Flash-Next",
      "title": "Qwen3.8-Flash-Next HuggingFace Model Card",
      "publisher": "Alibaba (Qwen)",
      "type": "Lab self-report",
      "published": "2026-08-24",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Qwen3.8-Flash-Next 58.7%; Qwen3.7 Plus 16.5%.",
      "readings": 2,
      "n": 13
    },
    {
      "id": "r-1171e08",
      "url": "https://huggingface.co/zai-org/GLM-5.3",
      "title": "GLM-5.3 HuggingFace Model Card",
      "publisher": "Z.ai",
      "type": "Lab self-report",
      "published": "2026-08-25",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for GLM-5.3 66.9% [max].",
      "readings": 1,
      "n": 14
    },
    {
      "id": "r-p3aj15",
      "url": "https://huggingface.co/tencent/Hy4-preview",
      "title": "Tencent Hy4-preview Technical Report & Model Card",
      "publisher": "Tencent",
      "type": "Lab self-report",
      "published": "2026-08-27",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Hy4 Preview 64.3%; Hy3 28%.",
      "readings": 2,
      "n": 15
    },
    {
      "id": "r-1b2qf5e",
      "url": "https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32c215d14235/Claude%20Fable%205.1%20&%20Claude%20Mythos%205.1%20System%20Card.pdf",
      "title": "Claude Fable 5.1 & Claude Mythos 5.1 System Card",
      "publisher": "Anthropic",
      "type": "Lab self-report",
      "published": "2026-09-01",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Claude Fable 5.1 67.4% [max].",
      "readings": 1,
      "n": 16
    },
    {
      "id": "r-1wqp1zm",
      "url": "https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology",
      "title": "Meta AI Research Muse Spark 1.3 Evaluation Methodology",
      "publisher": "Meta",
      "type": "Lab self-report",
      "published": "2026-09-02",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Muse Spark 1.3 75.4% [max].",
      "readings": 1,
      "n": 17
    },
    {
      "id": "r-pixn6e",
      "url": "https://github.com/nex-agi/Nex-N2.5",
      "title": "Nex-AGI Nex-N2.5 GitHub README / Model Card",
      "publisher": "Nex-AGI",
      "type": "Lab self-report",
      "published": "2026-09-08",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Nex-N2.5-Max 65.6%; Nex-N2.5-Pro 55.8%; Nex-N2.5-Mini 36.1%.",
      "readings": 3,
      "n": 18
    },
    {
      "id": "r-prwi23",
      "url": "https://cognition.com/blog/swe-2",
      "title": "Cognition SWE-2 launch post",
      "publisher": "Cognition",
      "type": "Lab self-report",
      "published": "2026-09-10",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for SWE-2 73%; SWE-1.7 37.7%.",
      "readings": 2,
      "n": 19
    },
    {
      "id": "r-1ftq9xa",
      "url": "https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash",
      "title": "DeepSeek-V4.1-Flash HuggingFace Model Card",
      "publisher": "DeepSeek",
      "type": "Lab self-report",
      "published": "2026-09-10",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for DeepSeek V4.1 Flash 65.5%–74.2% across 8 settings; DeepSeek V4 Pro 62.7% [max]; DeepSeek V4 Flash 54.4% [max].",
      "readings": 10,
      "n": 20
    },
    {
      "id": "r-s5arsj",
      "url": "https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/",
      "title": "Google Gemini 3.8 Flash launch blog and Model Evaluation Report",
      "publisher": "Google DeepMind",
      "type": "Lab self-report",
      "published": "2026-09-17",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Gemini 3.8 Flash 73.7% [high], with cost per task, tokens.",
      "readings": 1,
      "n": 21
    },
    {
      "id": "r-1o5c47s",
      "url": "https://www.stepfun.com/",
      "title": "StepFun Step 5 Preview Announcement",
      "publisher": "StepFun",
      "type": "Lab self-report",
      "published": "2026-09-20",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Step 5 Preview 67.7% [high].",
      "readings": 1,
      "n": 22
    },
    {
      "id": "r-1g5kt5u",
      "url": "https://media.x.ai/v1/website/card4p7-3a96f40b.pdf",
      "title": "xAI Grok 4.7 Model Card",
      "publisher": "xAI",
      "type": "Lab self-report",
      "published": "2026-09-21",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Grok 4.7 71% [high].",
      "readings": 1,
      "n": 23
    },
    {
      "id": "r-1mttwp9",
      "url": "https://mimo.xiaomi.com/mimo-v2-6",
      "title": "Xiaomi MiMo-V2.6 Launch Announcement and Appendix",
      "publisher": "Xiaomi",
      "type": "Lab self-report",
      "published": "2026-09-21",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for MiMo-V2.6-Pro 71.9% [max]; MiMo-V2.6-Flash 67.9% [max].",
      "readings": 2,
      "n": 24
    },
    {
      "id": "r-1jseipb",
      "url": "https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf",
      "title": "Claude Opus 5.5 System Card",
      "publisher": "Anthropic",
      "type": "Lab self-report",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Claude Opus 5.5 74.2% [max].",
      "readings": 1,
      "n": 25
    },
    {
      "id": "r-1m7ewdf",
      "url": "https://openai.com/index/introducing-gpt-6-sol-and-luna/",
      "title": "OpenAI GPT-6 Sol and Luna launch post",
      "publisher": "OpenAI",
      "type": "Lab self-report",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for GPT-6 Sol 37.2%–68.8% across 5 settings; GPT-6 Luna 2.4%–66.6% across 5 settings; GPT-5.6 Luna 1.2%–62.2% across 5 settings, with cost per task.",
      "readings": 15,
      "n": 26
    },
    {
      "id": "r-1vlr43v",
      "url": "https://fireworks.ai/blog/ember-1",
      "title": "Fireworks AI Ember-1 Announcement",
      "publisher": "Fireworks AI",
      "type": "Lab self-report",
      "published": "2026-09-23",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Ember-1 75.2% [thinking]; Kimi K3 55.8%–66.4% across 3 settings, with cost per task.",
      "readings": 4,
      "n": 27
    },
    {
      "id": "r-1ctctri",
      "url": "https://storage.googleapis.com/deepmind-media/gemini/gemini_4_argon_model_evaluation.pdf",
      "title": "Gemini 4 Argon Model Evaluation Report",
      "publisher": "Google DeepMind",
      "type": "Lab self-report",
      "published": "2026-09-24",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Gemini 4 Argon 77.9%.",
      "readings": 1,
      "n": 28
    },
    {
      "id": "r-1tiaoyx",
      "url": "https://www-cdn.anthropic.com/870c8f525702625d2c62fc6dd04c857e3250bec1/Claude%20Sonnet%205.5%20System%20Card.pdf",
      "title": "Claude Sonnet 5.5 System Card",
      "publisher": "Anthropic",
      "type": "Lab self-report",
      "published": "2026-09-28",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Claude Sonnet 5.5 71% [max].",
      "readings": 1,
      "n": 29
    },
    {
      "id": "r-1ntqzs0",
      "url": "https://openai.com/index/introducing-gpt-6-1-sol/",
      "title": "OpenAI GPT-6.1 Sol launch post",
      "publisher": "OpenAI",
      "type": "Lab self-report",
      "published": "2026-09-29",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for GPT-6.1 Sol 64.4%–75.2% across 5 settings, with cost per task.",
      "readings": 5,
      "n": 30
    },
    {
      "id": "r-1x8ip3p",
      "url": "https://reflection.ai/blog/introducing-beam",
      "title": "Reflection AI Beam launch post",
      "publisher": "Reflection AI",
      "type": "Lab self-report",
      "published": "2026-10-05",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Beam 44.4%.",
      "readings": 1,
      "n": 31
    },
    {
      "id": "r-d2tk0p",
      "url": "https://mistral.ai/news/mistral-large-4/",
      "title": "Mistral Large 4 Announcement",
      "publisher": "Mistral AI",
      "type": "Lab self-report",
      "published": "2026-10-06",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Mistral Large 4 61.7% [thinking].",
      "readings": 1,
      "n": 32
    },
    {
      "id": "r-1kqtfhm",
      "title": "Claude Haiku 5.5 System Card",
      "publisher": "Anthropic",
      "type": "Lab system card",
      "url": "https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf",
      "published": "2026-09-29",
      "accessed": "2026-10-09",
      "establishes": "Reports SWE-bench Pro, FrontierSWE v2 and ProgramBench for Haiku 5.5, and no DeepSWE 1.1 result.",
      "key": "anthropic-haiku-55",
      "readings": 0,
      "n": 33
    },
    {
      "id": "r-hzsdm7",
      "url": "https://entrpi.github.io/misc/deep-swe-minimax-m3/",
      "title": "Independent DeepSWE Audit by entrpi",
      "publisher": "entrpi (independent)",
      "type": "Independent run",
      "published": "2026-06-02",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for MiniMax M3 13.3%–16.8% across 2 settings, with cost per task, tokens.",
      "readings": 2,
      "n": 34
    },
    {
      "id": "r-17v6sj",
      "url": "https://artificialanalysis.ai/articles/benchmarking-grok-4-7",
      "title": "Artificial Analysis (Benchmarking Grok 4.7)",
      "publisher": "Artificial Analysis",
      "type": "Independent run",
      "published": "2026-09-21",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Grok 4.7 73% [xhigh]; Grok 4.6 65% [high].",
      "readings": 2,
      "n": 35
    },
    {
      "id": "r-14zikky",
      "url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "title": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "publisher": "Mercor",
      "type": "Independent run",
      "published": "",
      "accessed": "2026-10-09",
      "establishes": "Pass@1 for Claude Opus 5.5 72.3% [max]; GPT-6.1 Sol 72.3% [max]; GPT-6 Astra 72% [max]; DeepSeek V4.1 Flash 71.7% [max]; Claude Opus 5 70.2%–71.4% across 2 settings; Gemini 3.8 Flash 71.4% [high]; and 43 more models, with confidence intervals.",
      "readings": 52,
      "n": 36
    },
    {
      "id": "r-1tuhx3r",
      "title": "Benchmark review: DeepSWE v1.1",
      "publisher": "Epoch AI",
      "type": "Independent audit",
      "url": "https://epoch.ai/benchmarks/deepswe/review",
      "published": "2026-09-07",
      "accessed": "2026-10-09",
      "establishes": "Rates DeepSWE v1.1 flawed: grading defects in at least 23 of 113 tasks, most from hidden tests colliding with tests the agent wrote.",
      "key": "epoch-review",
      "readings": 0,
      "n": 37
    },
    {
      "id": "r-18kp64f",
      "title": "envcheck: auditing DeepSWE",
      "publisher": "Tokenless",
      "type": "Independent audit",
      "url": "https://usetokenless.com/blog/envcheck",
      "published": "2026-09-30",
      "accessed": "2026-10-09",
      "establishes": "Documents 70 ways a submission can rewrite test outcomes, because the test harness sits inside the repository the agent edits.",
      "key": "tokenless-envcheck",
      "readings": 0,
      "n": 38
    },
    {
      "id": "r-i0raoa",
      "title": "datacurve-ai/deep-swe issue #103",
      "publisher": "GitHub (community discussion)",
      "type": "Discussion",
      "url": "https://github.com/datacurve-ai/deep-swe/issues/103",
      "published": "2026-09-30",
      "accessed": "2026-10-09",
      "establishes": "Notes the board has published no new model since 3 Sep and quotes Datacurve's CEO saying the team is building the next generation of DeepSWE.",
      "key": "gh-issue-103",
      "readings": 0,
      "n": 39
    },
    {
      "id": "r-isoxsa",
      "title": "AI Gateway model list and prices",
      "publisher": "Vercel",
      "type": "Price list",
      "url": "https://ai-gateway.vercel.sh/v1/models",
      "published": "",
      "accessed": "2026-10-09",
      "establishes": "Per-token list prices used to estimate cost where a source published tokens but no cost: Laguna S 2.1 at $0.09 per million input tokens and $0.18 per million output tokens.",
      "key": "ai-gateway-pricing",
      "readings": 0,
      "n": 40
    }
  ],
  "rows": [
    {
      "model_key": "gemini-4-argon",
      "model_as_reported": "Gemini 4 Argon",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 77.9,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Gemini 4 Argon Model Evaluation Report",
      "source_url": "https://storage.googleapis.com/deepmind-media/gemini/gemini_4_argon_model_evaluation.pdf",
      "published": "2026-09-24",
      "accessed": "2026-10-09",
      "quote": "Gemini 4 Argon achieved 77.9% on DeepSWE v1.1 with mini-swe agent harness under highest thinking setting.",
      "notes": "Google eval PDF: self-computed with mini-swe-agent at the highest thinking setting; no level name given. Only Gemini 4 tier released. Evaluation report specifies highest thinking setting, pass@1, mini-swe-agent. No task cost or token counts given.",
      "display_name": "Gemini 4 Argon",
      "lab": "Google",
      "open_weights": false,
      "row_id": "row-1xty4qe",
      "source_id": "r-1ctctri",
      "source_n": 28
    },
    {
      "model_key": "muse-spark-1.3",
      "model_as_reported": "Muse Spark 1.3",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 75.4,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Meta AI Research Muse Spark 1.3 Evaluation Methodology",
      "source_url": "https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology",
      "published": "2026-09-02",
      "accessed": "2026-10-09",
      "quote": "We run Muse Spark 1.3 max with a mini-swe agent and obtained the benchmarks for all other models on the official leaderboard on Datacurve.",
      "notes": "Corrects t5-community.json and t4-aggregators.json: Meta's primary methodology document explicitly states mini-swe-agent (not Muse Code) and gives no task cost ($0.55 was an external AA estimate).",
      "display_name": "Muse Spark 1.3",
      "lab": "Meta",
      "open_weights": false,
      "row_id": "row-47v276",
      "source_id": "r-1wqp1zm",
      "source_n": 17
    },
    {
      "model_key": "gpt-6.1-sol",
      "model_as_reported": "GPT-6.1 Sol",
      "effort": "high",
      "harness": "lab-internal",
      "score_pct": 75.22,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 0.6461,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6.1 Sol launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-1-sol/",
      "published": "2026-09-29",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-6.1 Sol\",\"modelLabel\":\"GPT-6.1 Sol\",\"cost\":0.6461,\"score\":0.7522,\"effortLabel\":\"High\"}",
      "notes": "Reported in embedded Vega-Lite spec on launch page.",
      "display_name": "GPT-6.1 Sol",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1j0dft9",
      "source_id": "r-1ntqzs0",
      "source_n": 30
    },
    {
      "model_key": "ember-1",
      "model_as_reported": "Ember-1",
      "effort": "thinking",
      "harness": "mini-swe-agent",
      "score_pct": 75.2,
      "ci_pct": null,
      "n_samples": 113,
      "cost_per_task_usd": 3.62,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Fireworks AI Ember-1 Announcement",
      "source_url": "https://fireworks.ai/blog/ember-1",
      "published": "2026-09-23",
      "accessed": "2026-10-09",
      "quote": "Ember-1: 75.2% | Comparison (Ember-1 vs. K3 Max): -23.7% / -126.9 USD",
      "notes": "Post-trained Kimi K3 variant reducing reasoning tokens by 35-50%. Evaluated across N=113 DeepSWE 1.1 tasks. Total cost was $408.54 ($3.62/task), a 23.7% savings vs K3 Max.",
      "display_name": "Ember-1",
      "lab": "Fireworks AI",
      "open_weights": false,
      "row_id": "row-p2qq0j",
      "source_id": "r-1vlr43v",
      "source_n": 27
    },
    {
      "model_key": "claude-opus-5.5",
      "model_as_reported": "Claude Opus 5.5",
      "effort": "max",
      "harness": "lab-internal",
      "score_pct": 74.2,
      "ci_pct": null,
      "n_samples": 5,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Claude Opus 5.5 System Card",
      "source_url": "https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "quote": "Claude Opus 5.5 scored an average of 74.2% over five trials.",
      "notes": "Section 8.3 reports 74.2% over 5 trials with adaptive thinking/max reasoning budget in internal evaluation. No cost or token figures given.",
      "display_name": "Claude Opus 5.5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-k5hr20",
      "source_id": "r-1jseipb",
      "source_n": 25
    },
    {
      "model_key": "deepseek-v4.1-flash",
      "model_as_reported": "DS-V4.1-Flash",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 74.2,
      "ci_pct": null,
      "n_samples": 8,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "DeepSeek-V4.1-Flash HuggingFace Model Card",
      "source_url": "https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash",
      "published": "2026-09-10",
      "accessed": "2026-10-09",
      "quote": "| DeepSWE v1.1 (Resolved) | 74.0 | 73.0 | 67.5 | 66.9 | 62.7 | 54.4 | **74.2** |",
      "notes": "Table: Comparison with frontier models (Max reasoning effort). All instruct results use reasoning_effort=100, temp=1.0, top_p=0.95. N=8 samples per task on DeepSWE v1.1, 1M-token context limit, max_steps=500 per agent.",
      "display_name": "DeepSeek V4.1 Flash",
      "lab": "DeepSeek",
      "open_weights": true,
      "row_id": "row-o4taeh",
      "source_id": "r-1ftq9xa",
      "source_n": 20
    },
    {
      "model_key": "gpt-6-astra",
      "model_as_reported": "gpt-6-astra",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 74.12,
      "ci_pct": 2.87,
      "n_samples": 4,
      "cost_per_task_usd": 4.4291,
      "tokens_per_task": 1486484,
      "input_tokens_per_task": 1456927,
      "output_tokens_per_task": 29557,
      "steps_per_task": 28.8,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-09-03",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_6_astra_xhigh: pass_at_1 0.7411504424778761, ci_half 0.02865396174075141, n_runs 4, mean_cost_usd 4.429117203539823",
      "notes": "pass@4 80.5%; mean 19 min per attempt; cost basis: Current pricing at all context lengths: $10/M uncached input, $12.50/M cache writes, $1/M cache reads, $50/M output; no separate compute-unit fee.; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-6 Astra",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-owdu43",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gemini-3.8-flash",
      "model_as_reported": "gemini-3-8-flash",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 73.83,
      "ci_pct": 1.42,
      "n_samples": 4,
      "cost_per_task_usd": 2.3623,
      "tokens_per_task": 21877692,
      "input_tokens_per_task": 21734449,
      "output_tokens_per_task": 143243,
      "steps_per_task": 166.3,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-09-01",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gemini_3_8_flash_high: pass_at_1 0.738255033557047, ci_half 0.014173132066358783, n_runs 4, mean_cost_usd 2.362349413758389",
      "notes": "pass@4 85.8%; mean 11 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Gemini 3.8 Flash",
      "lab": "Google",
      "open_weights": false,
      "row_id": "row-1w5t75g",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gemini-3.8-flash",
      "model_as_reported": "Gemini 3.8 Flash",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 73.7,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 2.36,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": 143243,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Google Gemini 3.8 Flash launch blog and Model Evaluation Report",
      "source_url": "https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/",
      "published": "2026-09-17",
      "accessed": "2026-10-09",
      "quote": "On DeepSWE v1.1 (Long-Horizon Software Engineering) 3.8 Flash outperforms most larger frontier models in autonomously solving complex engineering problems end to end, only at a fraction of the cost.",
      "notes": "Score reported as 73.7% in evaluation table image and methodology PDF (https://storage.googleapis.com/deepmind-media/gemini/gemini_3-8_flash_model_evaluation.pdf): 'Results for Gemini 3.8 Flash are self computed, and use a mini-swe agent harness with high thinking.' Also reported as 73.8% in blog overview text and OpenAI Astra comparison chart ($2.36 / 143,243 output tokens at high effort; $1.97 / 124,684 tokens at medium effort).",
      "display_name": "Gemini 3.8 Flash",
      "lab": "Google",
      "open_weights": false,
      "row_id": "row-14zbox",
      "source_id": "r-s5arsj",
      "source_n": 21
    },
    {
      "model_key": "claude-opus-5",
      "model_as_reported": "claude-opus-5",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 73.65,
      "ci_pct": 3.87,
      "n_samples": 4,
      "cost_per_task_usd": 11.8376,
      "tokens_per_task": 15143400,
      "input_tokens_per_task": 15025834,
      "output_tokens_per_task": 117566,
      "steps_per_task": 99,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-25",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_opus_5_max: pass_at_1 0.7364864864864865, ci_half 0.03872310426371729, n_runs 4, mean_cost_usd 11.837583271396396",
      "notes": "pass@4 88.5%; mean 32 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Opus 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-346qj2",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-6-astra",
      "model_as_reported": "gpt-6-astra",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 73.23,
      "ci_pct": 3.42,
      "n_samples": 4,
      "cost_per_task_usd": 3.9237,
      "tokens_per_task": 1309777,
      "input_tokens_per_task": 1283271,
      "output_tokens_per_task": 26506,
      "steps_per_task": 27.4,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-09-03",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_6_astra_high: pass_at_1 0.7323008849557522, ci_half 0.03423496057873005, n_runs 4, mean_cost_usd 3.9237237909292038",
      "notes": "pass@4 82.3%; mean 17 min per attempt; cost basis: Current pricing at all context lengths: $10/M uncached input, $12.50/M cache writes, $1/M cache reads, $50/M output; no separate compute-unit fee.; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-6 Astra",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-vqmucw",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-6-astra",
      "model_as_reported": "gpt-6-astra",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 73.23,
      "ci_pct": 0.83,
      "n_samples": 4,
      "cost_per_task_usd": 7.4978,
      "tokens_per_task": 2244182,
      "input_tokens_per_task": 2183034,
      "output_tokens_per_task": 61149,
      "steps_per_task": 28.5,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-09-03",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_6_astra_max: pass_at_1 0.7323008849557522, ci_half 0.008303197562056526, n_runs 4, mean_cost_usd 7.497827996681416",
      "notes": "pass@4 79.6%; mean 33 min per attempt; cost basis: Current pricing at all context lengths: $10/M uncached input, $12.50/M cache writes, $1/M cache reads, $50/M output; no separate compute-unit fee.; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-6 Astra",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-5dvg8e",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "claude-opus-5",
      "model_as_reported": "claude-opus-5",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 73.15,
      "ci_pct": 3.06,
      "n_samples": 4,
      "cost_per_task_usd": 9.0722,
      "tokens_per_task": 11384401,
      "input_tokens_per_task": 11292729,
      "output_tokens_per_task": 91672,
      "steps_per_task": 88.7,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-25",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_opus_5_xhigh: pass_at_1 0.7315436241610739, ci_half 0.030609818216384643, n_runs 4, mean_cost_usd 9.07220232326622",
      "notes": "pass@4 85.8%; mean 26 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Opus 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1oeh1kj",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-6.1-sol",
      "model_as_reported": "GPT-6.1 Sol",
      "effort": "medium",
      "harness": "lab-internal",
      "score_pct": 73.01,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 0.4196,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6.1 Sol launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-1-sol/",
      "published": "2026-09-29",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-6.1 Sol\",\"modelLabel\":\"GPT-6.1 Sol\",\"cost\":0.4196,\"score\":0.7301,\"effortLabel\":\"Medium\"}",
      "notes": "Reported in embedded Vega-Lite spec on launch page.",
      "display_name": "GPT-6.1 Sol",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-16dfael",
      "source_id": "r-1ntqzs0",
      "source_n": 30
    },
    {
      "model_key": "grok-4.7",
      "model_as_reported": "Grok 4.7",
      "effort": "xhigh",
      "harness": "Grok Build",
      "score_pct": 73,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Artificial Analysis (Benchmarking Grok 4.7)",
      "source_url": "https://artificialanalysis.ai/articles/benchmarking-grok-4-7",
      "published": "2026-09-21",
      "accessed": "2026-10-09",
      "quote": "DeepSWE v1.1: 73% (improved from 65% on Grok 4.6 xhigh)",
      "notes": "Corrects t5-community.json (which attributed 73% with Grok Build to xAI news). 73.0% was evaluated and published by Artificial Analysis as part of Coding Agent Index.",
      "display_name": "Grok 4.7",
      "lab": "xAI",
      "open_weights": false,
      "row_id": "row-3bvbtr",
      "source_id": "r-17v6sj",
      "source_n": 35
    },
    {
      "model_key": "swe-2",
      "model_as_reported": "SWE-2",
      "effort": "",
      "harness": "Devin CLI",
      "score_pct": 73,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Cognition SWE-2 launch post",
      "source_url": "https://cognition.com/blog/swe-2",
      "published": "2026-09-10",
      "accessed": "2026-10-09",
      "quote": "SWE-2 | 73.0%",
      "notes": "Evaluated using Devin CLI ('Devin CLI for open-weight models'). Best score across reasoning-effort settings. Post states: 'On FrontierCode 1.1 Main and DeepSWE 1.1, SWE-2 beats SWE-1.7 and Grok 4.6 on both score and cost'.",
      "display_name": "SWE-2",
      "lab": "Cognition",
      "open_weights": false,
      "row_id": "row-1o9vzqd",
      "source_id": "r-prwi23",
      "source_n": 19
    },
    {
      "model_key": "claude-opus-5",
      "model_as_reported": "claude-opus-5",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 72.83,
      "ci_pct": 1.95,
      "n_samples": 4,
      "cost_per_task_usd": 6.0761,
      "tokens_per_task": 7298087,
      "input_tokens_per_task": 7233879,
      "output_tokens_per_task": 64207,
      "steps_per_task": 72.9,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-25",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_opus_5_high: pass_at_1 0.7282850779510023, ci_half 0.019454321571601495, n_runs 4, mean_cost_usd 6.076084489977728",
      "notes": "pass@4 87.6%; mean 19 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Opus 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-lbxd05",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-6-astra",
      "model_as_reported": "gpt-6-astra",
      "effort": "medium",
      "harness": "mini-swe-agent",
      "score_pct": 72.79,
      "ci_pct": 2.59,
      "n_samples": 4,
      "cost_per_task_usd": 3.0755,
      "tokens_per_task": 1028600,
      "input_tokens_per_task": 1008238,
      "output_tokens_per_task": 20362,
      "steps_per_task": 26,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-09-03",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_6_astra_medium: pass_at_1 0.7278761061946902, ci_half 0.025896490818318695, n_runs 4, mean_cost_usd 3.0754742975663714",
      "notes": "pass@4 82.3%; mean 15 min per attempt; cost basis: Current pricing at all context lengths: $10/M uncached input, $12.50/M cache writes, $1/M cache reads, $50/M output; no separate compute-unit fee.; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-6 Astra",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1wxhh73",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-5.6-sol",
      "model_as_reported": "gpt-5-6-sol",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 72.67,
      "ci_pct": 2.83,
      "n_samples": 4,
      "cost_per_task_usd": 6.456,
      "tokens_per_task": 7967666,
      "input_tokens_per_task": 7907652,
      "output_tokens_per_task": 60014,
      "steps_per_task": 61.3,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-10",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_6_sol_max: pass_at_1 0.7266666666666667, ci_half 0.02829822249837175, n_runs 4, mean_cost_usd 6.455954682467823",
      "notes": "pass@4 85.8%; mean 19 min per attempt; leaderboard-live.json gives $8.39 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.6 Sol",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-y7m969",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "deepseek-v4.1-flash",
      "model_as_reported": "DS-V4.1-Flash (DSH Minimal)",
      "effort": "max",
      "harness": "lab-internal",
      "score_pct": 72.6,
      "ci_pct": null,
      "n_samples": 8,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "DeepSeek-V4.1-Flash HuggingFace Model Card",
      "source_url": "https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash",
      "published": "2026-09-10",
      "accessed": "2026-10-09",
      "quote": "| DeepSWE v1.1 (Resolved) | 69.8 | 65.6 | 65.5 | 66.2 | 74.2 | 72.6 | 70.5 | 67.6 |",
      "notes": "Evaluated with DeepSeek Harness Minimal (DSH Minimal), reasoning_effort=100, N=8, temp=1.0, top_p=0.95, 1M context, max_steps=500.",
      "display_name": "DeepSeek V4.1 Flash",
      "lab": "DeepSeek",
      "open_weights": true,
      "row_id": "row-16p8paj",
      "source_id": "r-1ftq9xa",
      "source_n": 20
    },
    {
      "model_key": "claude-opus-5.5",
      "model_as_reported": "Opus 5.5",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 72.3,
      "ci_pct": 7.3,
      "n_samples": 336,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Opus 5.5 | effort: max | score: 72.3% ± 7.3% | n_samples: 336 | provider: Anthropic",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Claude Opus 5.5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1sj5dau",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gpt-6.1-sol",
      "model_as_reported": "GPT 6.1 Sol",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 72.3,
      "ci_pct": 7.4,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "GPT 6.1 Sol | effort: max | score: 72.3% ± 7.4% | n_samples: 339 | provider: OpenAI",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "GPT-6.1 Sol",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1unx8j9",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gpt-6-astra",
      "model_as_reported": "GPT 6 Astra",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 72,
      "ci_pct": 7.5,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "GPT 6 Astra | effort: max | score: 72% ± 7.5% | n_samples: 339 | provider: OpenAI",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "GPT-6 Astra",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1mtbg4a",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gpt-6.1-sol",
      "model_as_reported": "GPT-6.1 Sol",
      "effort": "max",
      "harness": "lab-internal",
      "score_pct": 71.9,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 1.5711,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6.1 Sol launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-1-sol/",
      "published": "2026-09-29",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-6.1 Sol\",\"modelLabel\":\"GPT-6.1 Sol\",\"cost\":1.5711,\"score\":0.7190,\"effortLabel\":\"Max\"}",
      "notes": "Corrects t5-community.json (71.0% via GitHub issue #102). Exact score is 71.90% with cost $1.5711.",
      "display_name": "GPT-6.1 Sol",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-z4hdo2",
      "source_id": "r-1ntqzs0",
      "source_n": 30
    },
    {
      "model_key": "gpt-6.1-sol",
      "model_as_reported": "GPT-6.1 Sol",
      "effort": "xhigh",
      "harness": "lab-internal",
      "score_pct": 71.9,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 0.7886,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6.1 Sol launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-1-sol/",
      "published": "2026-09-29",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-6.1 Sol\",\"modelLabel\":\"GPT-6.1 Sol\",\"cost\":0.7886,\"score\":0.7190,\"effortLabel\":\"Xhigh\"}",
      "notes": "Corrects t5-community.json (71.0% via GitHub issue #102). Exact score is 71.90% with cost $0.7886.",
      "display_name": "GPT-6.1 Sol",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1xg0j7k",
      "source_id": "r-1ntqzs0",
      "source_n": 30
    },
    {
      "model_key": "mimo-v2.6-pro",
      "model_as_reported": "MiMo-V2.6-Pro",
      "effort": "max",
      "harness": "lab-internal",
      "score_pct": 71.9,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Xiaomi MiMo-V2.6 Launch Announcement and Appendix",
      "source_url": "https://mimo.xiaomi.com/mimo-v2-6",
      "published": "2026-09-21",
      "accessed": "2026-10-09",
      "quote": "In Xiaomi’s launch appendix, MiMo-V2.6-Pro scored 71.9 on DeepSWE v1.1.",
      "notes": "1.02T total / 42B active parameter MoE model. Launch appendix reports 71.9% on DeepSWE v1.1. Also documented on Kingy.ai and FoneArena.",
      "display_name": "MiMo-V2.6-Pro",
      "lab": "Xiaomi",
      "open_weights": true,
      "row_id": "row-b1kgds",
      "source_id": "r-1mttwp9",
      "source_n": 24
    },
    {
      "model_key": "deepseek-v4.1-flash",
      "model_as_reported": "DeepSeek V4.1 Flash",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 71.7,
      "ci_pct": 6.3,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "DeepSeek V4.1 Flash | effort: max | score: 71.7% ± 6.3% | n_samples: 339 | provider: DeepSeek",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "DeepSeek V4.1 Flash",
      "lab": "DeepSeek",
      "open_weights": true,
      "row_id": "row-1n55rtx",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "claude-opus-5",
      "model_as_reported": "Opus 5",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 71.4,
      "ci_pct": 6.8,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Opus 5 | effort: max | score: 71.4% ± 6.8% | n_samples: 339 | provider: Anthropic",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Claude Opus 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-nhlnrl",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gemini-3.8-flash",
      "model_as_reported": "Gemini 3.8 Flash",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 71.4,
      "ci_pct": 6.8,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Gemini 3.8 Flash | effort: high | score: 71.4% ± 6.8% | n_samples: 339 | provider: Google",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Gemini 3.8 Flash",
      "lab": "Google",
      "open_weights": false,
      "row_id": "row-1kjx08d",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gemini-3.8-flash",
      "model_as_reported": "gemini-3-8-flash",
      "effort": "medium",
      "harness": "mini-swe-agent",
      "score_pct": 71.02,
      "ci_pct": 2.28,
      "n_samples": 4,
      "cost_per_task_usd": 1.9671,
      "tokens_per_task": 16988624,
      "input_tokens_per_task": 16863940,
      "output_tokens_per_task": 124684,
      "steps_per_task": 147.3,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-09-01",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gemini_3_8_flash_medium: pass_at_1 0.7101769911504425, ci_half 0.02280804572877925, n_runs 4, mean_cost_usd 1.9671024789823006",
      "notes": "pass@4 83.2%; mean 14 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Gemini 3.8 Flash",
      "lab": "Google",
      "open_weights": false,
      "row_id": "row-itcfkk",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "claude-sonnet-5.5",
      "model_as_reported": "Claude Sonnet 5.5",
      "effort": "max",
      "harness": "lab-internal",
      "score_pct": 71,
      "ci_pct": null,
      "n_samples": 5,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Claude Sonnet 5.5 System Card",
      "source_url": "https://www-cdn.anthropic.com/870c8f525702625d2c62fc6dd04c857e3250bec1/Claude%20Sonnet%205.5%20System%20Card.pdf",
      "published": "2026-09-28",
      "accessed": "2026-10-09",
      "quote": "Sonnet 5.5 scored an average of 71.0% over five trials.",
      "notes": "Section 8.3 reports 71.0% over 5 trials with adaptive thinking/max reasoning budget in internal evaluation. No cost or token figures given.",
      "display_name": "Claude Sonnet 5.5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1ns4aaz",
      "source_id": "r-1tiaoyx",
      "source_n": 29
    },
    {
      "model_key": "grok-4.7",
      "model_as_reported": "Grok 4.7",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 71,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "xAI Grok 4.7 Model Card",
      "source_url": "https://media.x.ai/v1/website/card4p7-3a96f40b.pdf",
      "published": "2026-09-21",
      "accessed": "2026-10-09",
      "quote": "Grok 4.7 (xHigh): 71.0% on DeepSWE v1.1",
      "notes": "xAI card p.7: \"Grok 4.7 scores 71.0% at high\"; footnote: results taken from evaluations conducted by Datacurve, which never posted them on its public board. Published by xAI in official model card under mini-swe-agent.",
      "display_name": "Grok 4.7",
      "lab": "xAI",
      "open_weights": false,
      "row_id": "row-1kwajgk",
      "source_id": "r-1g5kt5u",
      "source_n": 23
    },
    {
      "model_key": "gpt-5.6-sol",
      "model_as_reported": "gpt-5-6-sol",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 70.73,
      "ci_pct": 0.82,
      "n_samples": 4,
      "cost_per_task_usd": 3.5986,
      "tokens_per_task": 4300063,
      "input_tokens_per_task": 4259318,
      "output_tokens_per_task": 40745,
      "steps_per_task": 44,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-10",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_6_sol_xhigh: pass_at_1 0.7073170731707317, ci_half 0.008191117217343103, n_runs 4, mean_cost_usd 3.5986416572286237",
      "notes": "pass@4 85.8%; mean 13 min per attempt; leaderboard-live.json gives $4.70 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.6 Sol",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-tdaaio",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "claude-sonnet-5.5",
      "model_as_reported": "Sonnet 5.5",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 70.5,
      "ci_pct": 7.7,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Sonnet 5.5 | effort: max | score: 70.5% ± 7.7% | n_samples: 339 | provider: Anthropic",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Claude Sonnet 5.5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-qoxc9u",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "deepseek-v4.1-flash",
      "model_as_reported": "DS-V4.1-Flash (DSH Standard)",
      "effort": "max",
      "harness": "lab-internal",
      "score_pct": 70.5,
      "ci_pct": null,
      "n_samples": 8,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "DeepSeek-V4.1-Flash HuggingFace Model Card",
      "source_url": "https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash",
      "published": "2026-09-10",
      "accessed": "2026-10-09",
      "quote": "| DeepSWE v1.1 (Resolved) | 69.8 | 65.6 | 65.5 | 66.2 | 74.2 | 72.6 | 70.5 | 67.6 |",
      "notes": "Evaluated with DeepSeek Harness Standard (DSH Standard), reasoning_effort=100, N=8, temp=1.0, top_p=0.95, 1M context, max_steps=500.",
      "display_name": "DeepSeek V4.1 Flash",
      "lab": "DeepSeek",
      "open_weights": true,
      "row_id": "row-106utim",
      "source_id": "r-1ftq9xa",
      "source_n": 20
    },
    {
      "model_key": "glm-5.3",
      "model_as_reported": "GLM 5.3",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 70.5,
      "ci_pct": 6.3,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "GLM 5.3 | effort: max | score: 70.5% ± 6.3% | n_samples: 339 | provider: Zhipu",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "GLM-5.3",
      "lab": "Z.ai",
      "open_weights": true,
      "row_id": "row-18ttaw",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gpt-5.6-sol",
      "model_as_reported": "GPT 5.6 Sol",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 70.5,
      "ci_pct": 6.9,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "GPT 5.6 Sol | effort: max | score: 70.5% ± 6.9% | n_samples: 339 | provider: OpenAI",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "GPT-5.6 Sol",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1c4imbl",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gpt-6-sol",
      "model_as_reported": "GPT 6 Sol",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 70.5,
      "ci_pct": 7.2,
      "n_samples": 338,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "GPT 6 Sol | effort: max | score: 70.5% ± 7.2% | n_samples: 338 | provider: OpenAI",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "GPT-6 Sol",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-tdvxi6",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "claude-opus-5",
      "model_as_reported": "Opus 5",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 70.2,
      "ci_pct": 6.9,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Opus 5 | effort: xhigh | score: 70.2% ± 6.9% | n_samples: 339 | provider: Anthropic",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Claude Opus 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1dtxecg",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gpt-5.6-terra",
      "model_as_reported": "GPT 5.6 Terra",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 70.2,
      "ci_pct": 7.2,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "GPT 5.6 Terra | effort: max | score: 70.2% ± 7.2% | n_samples: 339 | provider: OpenAI",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "GPT-5.6 Terra",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-n5l7y",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "claude-fable-5",
      "model_as_reported": "claude-fable-5",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 69.91,
      "ci_pct": 3.24,
      "n_samples": 4,
      "cost_per_task_usd": 13.4145,
      "tokens_per_task": 7409513,
      "input_tokens_per_task": 7329161,
      "output_tokens_per_task": 80352,
      "steps_per_task": 68.4,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_fable_5_xhigh: pass_at_1 0.6991150442477876, ci_half 0.03244917575471319, n_runs 4, mean_cost_usd 13.414521495535714",
      "notes": "pass@4 88.5%; mean 24 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Fable 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1v14szb",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "deepseek-v4.1-flash",
      "model_as_reported": "DS-V4.1-Flash (Claude Code)",
      "effort": "max",
      "harness": "Claude Code",
      "score_pct": 69.8,
      "ci_pct": null,
      "n_samples": 8,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "DeepSeek-V4.1-Flash HuggingFace Model Card",
      "source_url": "https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash",
      "published": "2026-09-10",
      "accessed": "2026-10-09",
      "quote": "| DeepSWE v1.1 (Resolved) | 69.8 | 65.6 | 65.5 | 66.2 | 74.2 | 72.6 | 70.5 | 67.6 |",
      "notes": "Performance across agent scaffolds table. Evaluated with Claude Code scaffold, reasoning_effort=100, N=8, temp=1.0, top_p=0.95, 1M context limit, max_steps=500.",
      "display_name": "DeepSeek V4.1 Flash",
      "lab": "DeepSeek",
      "open_weights": true,
      "row_id": "row-vejsjt",
      "source_id": "r-1ftq9xa",
      "source_n": 20
    },
    {
      "model_key": "claude-fable-5",
      "model_as_reported": "claude-fable-5",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 69.72,
      "ci_pct": 4.03,
      "n_samples": 4,
      "cost_per_task_usd": 21.6347,
      "tokens_per_task": 12757425,
      "input_tokens_per_task": 12638832,
      "output_tokens_per_task": 118593,
      "steps_per_task": 88.4,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_fable_5_max: pass_at_1 0.6972477064220184, ci_half 0.04033474250807936, n_runs 4, mean_cost_usd 21.634702092592594",
      "notes": "pass@4 84.1%; mean 35 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Fable 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-vw58k6",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-5.6-terra",
      "model_as_reported": "gpt-5-6-terra",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 69.62,
      "ci_pct": 2.56,
      "n_samples": 4,
      "cost_per_task_usd": 3.9567,
      "tokens_per_task": 9302499,
      "input_tokens_per_task": 9230561,
      "output_tokens_per_task": 71939,
      "steps_per_task": 75.9,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-10",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_6_terra_max: pass_at_1 0.6962305986696231, ci_half 0.025569389504458404, n_runs 4, mean_cost_usd 3.956677811086475",
      "notes": "pass@4 88.5%; mean 17 min per attempt; leaderboard-live.json gives $4.95 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.6 Terra",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1ly571m",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-5.6-sol",
      "model_as_reported": "gpt-5-6-sol",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 69.4,
      "ci_pct": 1.43,
      "n_samples": 4,
      "cost_per_task_usd": 2.6618,
      "tokens_per_task": 2742023,
      "input_tokens_per_task": 2713572,
      "output_tokens_per_task": 28450,
      "steps_per_task": 36.9,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-10",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_6_sol_high: pass_at_1 0.6940133037694013, ci_half 0.014321452203772355, n_runs 4, mean_cost_usd 2.661832088842267",
      "notes": "pass@4 86.7%; mean 10 min per attempt; leaderboard-live.json gives $3.47 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.6 Sol",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-ajdh5w",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "glm-5.3",
      "model_as_reported": "glm-5-3",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 68.96,
      "ci_pct": 3.02,
      "n_samples": 4,
      "cost_per_task_usd": 3.9934,
      "tokens_per_task": 13352780,
      "input_tokens_per_task": 13272344,
      "output_tokens_per_task": 80436,
      "steps_per_task": 124.5,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-08-20",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_glm_5_3_max: pass_at_1 0.6895787139689579, ci_half 0.030186512715587, n_runs 4, mean_cost_usd 3.9933584893126386",
      "notes": "pass@4 87.6%; mean 35 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GLM-5.3",
      "lab": "Z.ai",
      "open_weights": true,
      "row_id": "row-fm38un",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "claude-opus-5",
      "model_as_reported": "claude-opus-5",
      "effort": "medium",
      "harness": "mini-swe-agent",
      "score_pct": 68.9,
      "ci_pct": 1.17,
      "n_samples": 4,
      "cost_per_task_usd": 3.2898,
      "tokens_per_task": 3616797,
      "input_tokens_per_task": 3579815,
      "output_tokens_per_task": 36982,
      "steps_per_task": 52.3,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-25",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_opus_5_medium: pass_at_1 0.6890380313199105, ci_half 0.011725586658930944, n_runs 4, mean_cost_usd 3.28978764541387",
      "notes": "pass@4 89.4%; mean 13 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Opus 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1g3833n",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-6-sol",
      "model_as_reported": "GPT-6 Sol",
      "effort": "max",
      "harness": "lab-internal",
      "score_pct": 68.81,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 2.7439,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6 Sol and Luna launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-sol-and-luna/",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-6 Sol\",\"cost_label\":\"$$2.74\",\"score\":0.6881,\"cost\":2.7439,\"effortLabel\":\"max\"}",
      "notes": "Reported in embedded Vega-Lite chart and text: 'On DeepSWE v1.1... GPT-6 Sol at max effort scores 68.8%, within 1.1 percentage points of Claude Fable 5’s highest score in the evaluation—69.9% at xhigh effort—at approximately 80% lower cost per task.'",
      "display_name": "GPT-6 Sol",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-15hrore",
      "source_id": "r-1m7ewdf",
      "source_n": 26
    },
    {
      "model_key": "claude-opus-5",
      "model_as_reported": "Claude Opus 5",
      "effort": "max",
      "harness": "lab-internal",
      "score_pct": 68.8,
      "ci_pct": null,
      "n_samples": 5,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Claude Opus 5 System Card",
      "source_url": "https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf",
      "published": "2026-07-24",
      "accessed": "2026-10-09",
      "quote": "Claude Opus 5 scored an average of 68.8% over five trials.",
      "notes": "Section 8.3 DeepSWE v1.1 (p. 149) and Table 8.1 (p. 147).",
      "display_name": "Claude Opus 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-kte6r1",
      "source_id": "r-l1zffg",
      "source_n": 9
    },
    {
      "model_key": "claude-fable-5",
      "model_as_reported": "claude-fable-5",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 68.6,
      "ci_pct": 1.12,
      "n_samples": 4,
      "cost_per_task_usd": 9.1776,
      "tokens_per_task": 4890465,
      "input_tokens_per_task": 4833178,
      "output_tokens_per_task": 57287,
      "steps_per_task": 58.7,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_fable_5_high: pass_at_1 0.686046511627907, ci_half 0.011208427174303877, n_runs 4, mean_cost_usd 9.177638214788733",
      "notes": "pass@4 86.7%; mean 18 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Fable 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1kfzh6g",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "kimi-k3",
      "model_as_reported": "kimi-k3",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 68.51,
      "ci_pct": 4.54,
      "n_samples": 4,
      "cost_per_task_usd": 4.6547,
      "tokens_per_task": 9776936,
      "input_tokens_per_task": 9695436,
      "output_tokens_per_task": 81500,
      "steps_per_task": 97.6,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-18",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_kimi_k3_max: pass_at_1 0.6851441241685144, ci_half 0.045370140792823206, n_runs 4, mean_cost_usd 4.654682129933482",
      "notes": "pass@4 89.4%; mean 76 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Kimi K3",
      "lab": "Moonshot AI",
      "open_weights": null,
      "row_id": "row-123w736",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "mimo-v2.6-flash",
      "model_as_reported": "MiMo-V2.6-Flash",
      "effort": "max",
      "harness": "lab-internal",
      "score_pct": 67.9,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Xiaomi MiMo-V2.6 Launch Announcement and Appendix",
      "source_url": "https://mimo.xiaomi.com/mimo-v2-6",
      "published": "2026-09-21",
      "accessed": "2026-10-09",
      "quote": "In Xiaomi’s launch appendix, MiMo-V2.6-Flash registered 67.9 on DeepSWE v1.1.",
      "notes": "Efficiency-optimized MoE model. Launch appendix reports 67.9% on DeepSWE v1.1.",
      "display_name": "MiMo-V2.6-Flash",
      "lab": "Xiaomi",
      "open_weights": true,
      "row_id": "row-q48lwe",
      "source_id": "r-1mttwp9",
      "source_n": 24
    },
    {
      "model_key": "step-5-preview",
      "model_as_reported": "Step 5 Preview",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 67.7,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "StepFun Step 5 Preview Announcement",
      "source_url": "https://www.stepfun.com/",
      "published": "2026-09-20",
      "accessed": "2026-10-09",
      "quote": "Step 5 Preview scored 67.7% on DeepSWE v1.1 under its High reasoning effort configuration.",
      "notes": "Source is StepFun’s homepage rather than a dated release page; the figure has not been re-checked on a page that is still reachable. 600B total / 27B active MoE foundation model with 1M context window. Evaluated at high reasoning effort with SWE-agent / mini-swe-agent harness.",
      "display_name": "Step 5 Preview",
      "lab": "StepFun",
      "open_weights": true,
      "row_id": "row-nj1v5q",
      "source_id": "r-1o5c47s",
      "source_n": 22
    },
    {
      "model_key": "deepseek-v4.1-flash",
      "model_as_reported": "DS-V4.1-Flash (DSH PTC)",
      "effort": "max",
      "harness": "lab-internal",
      "score_pct": 67.6,
      "ci_pct": null,
      "n_samples": 8,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "DeepSeek-V4.1-Flash HuggingFace Model Card",
      "source_url": "https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash",
      "published": "2026-09-10",
      "accessed": "2026-10-09",
      "quote": "| DeepSWE v1.1 (Resolved) | 69.8 | 65.6 | 65.5 | 66.2 | 74.2 | 72.6 | 70.5 | 67.6 |",
      "notes": "Evaluated with DeepSeek Harness Program-Test-Cycle (DSH PTC), reasoning_effort=100, N=8, temp=1.0, top_p=0.95, 1M context, max_steps=500.",
      "display_name": "DeepSeek V4.1 Flash",
      "lab": "DeepSeek",
      "open_weights": true,
      "row_id": "row-1r66635",
      "source_id": "r-1ftq9xa",
      "source_n": 20
    },
    {
      "model_key": "kimi-k3",
      "model_as_reported": "Kimi K3",
      "effort": "max",
      "harness": "lab-internal",
      "score_pct": 67.5,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Kimi-K3 HuggingFace Model Card",
      "source_url": "https://huggingface.co/moonshotai/Kimi-K3",
      "published": "2026-07-23",
      "accessed": "2026-10-09",
      "quote": "DeepSWE | 67.5 | 70.0 | 73.0 | 59.0 | 67.0 | 46.2",
      "notes": "Footnote 2: DeepSWE. Kimi K3 is evaluated with the Kimi Code harness. We report the DeepSWE v1.1 tasks. All Kimi K3 results are obtained with reasoning effort set to 'max', temperature = 1.0, top-p = 1.0 for agentic tasks.",
      "display_name": "Kimi K3",
      "lab": "Moonshot AI",
      "open_weights": null,
      "row_id": "row-ub0egg",
      "source_id": "r-kyj4bu",
      "source_n": 8
    },
    {
      "model_key": "grok-4.6",
      "model_as_reported": "grok-4-6",
      "effort": "medium",
      "harness": "mini-swe-agent",
      "score_pct": 67.48,
      "ci_pct": 2.28,
      "n_samples": 4,
      "cost_per_task_usd": 3.449,
      "tokens_per_task": 5774959,
      "input_tokens_per_task": 5725195,
      "output_tokens_per_task": 49764,
      "steps_per_task": 70.3,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-08-12",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_grok_4_6_medium: pass_at_1 0.6747787610619469, ci_half 0.0228080457287792, n_runs 4, mean_cost_usd 3.4489837522123894",
      "notes": "pass@4 84.1%; mean 15 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Grok 4.6",
      "lab": "xAI",
      "open_weights": false,
      "row_id": "row-cax4uv",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "claude-fable-5.1",
      "model_as_reported": "Claude Fable 5.1",
      "effort": "max",
      "harness": "lab-internal",
      "score_pct": 67.4,
      "ci_pct": null,
      "n_samples": 5,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Claude Fable 5.1 & Claude Mythos 5.1 System Card",
      "source_url": "https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32c215d14235/Claude%20Fable%205.1%20&%20Claude%20Mythos%205.1%20System%20Card.pdf",
      "published": "2026-09-01",
      "accessed": "2026-10-09",
      "quote": "Fable 5.1 scored an average of 67.4% over five trials. A note on scoring: DeepSWE grades each task with hidden tests that are often written for a single reference solution. When we reviewed failing transcripts, we found that Fable 5.1 implemented some ambiguous tasks more thoroughly than the task required, for example by validating inputs and raising exceptions, or by keeping new code consistent with codebase conventions.",
      "notes": "Section 8.3 DeepSWE v1.1 (p. 168). Also noted over-implementation penalty where more rigorous implementations failed single-solution reference tests.",
      "display_name": "Claude Fable 5.1",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-2f9pfs",
      "source_id": "r-1b2qf5e",
      "source_n": 16
    },
    {
      "model_key": "claude-fable-5",
      "model_as_reported": "Fable 5",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 67.3,
      "ci_pct": 7.2,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Fable 5 | effort: max | score: 67.3% ± 7.2% | n_samples: 339 | provider: Anthropic",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Claude Fable 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1e5knmg",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "claude-fable-5.1",
      "model_as_reported": "Fable 5.1",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 67.3,
      "ci_pct": 6.9,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Fable 5.1 | effort: high | score: 67.3% ± 6.9% | n_samples: 339 | provider: Anthropic",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Claude Fable 5.1",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-t03qkb",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gpt-5.6-luna",
      "model_as_reported": "gpt-5-6-luna",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 67.19,
      "ci_pct": 3.99,
      "n_samples": 4,
      "cost_per_task_usd": 0.6056,
      "tokens_per_task": 15517118,
      "input_tokens_per_task": 15443718,
      "output_tokens_per_task": 73400,
      "steps_per_task": 101.7,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-10",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_6_luna_max: pass_at_1 0.671875, ci_half 0.03992001235671055, n_runs 4, mean_cost_usd 0.6056233620535714",
      "notes": "pass@4 90.3%; mean 19 min per attempt; leaderboard-live.json gives $3.03 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.6 Luna",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1lhlm4s",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-5.5",
      "model_as_reported": "gpt-5-5",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 67.04,
      "ci_pct": 6.47,
      "n_samples": 4,
      "cost_per_task_usd": 7.2262,
      "tokens_per_task": 8419222,
      "input_tokens_per_task": 8372928,
      "output_tokens_per_task": 46295,
      "steps_per_task": 82,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_5_xhigh: pass_at_1 0.6703539823008849, ci_half 0.06465646340908864, n_runs 4, mean_cost_usd 7.226236674778761",
      "notes": "pass@4 88.5%; mean 30 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.5",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1st9iuc",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-6-astra",
      "model_as_reported": "gpt-6-astra",
      "effort": "low",
      "harness": "mini-swe-agent",
      "score_pct": 67.04,
      "ci_pct": 1.3,
      "n_samples": 4,
      "cost_per_task_usd": 1.5952,
      "tokens_per_task": 505086,
      "input_tokens_per_task": 494507,
      "output_tokens_per_task": 10580,
      "steps_per_task": 19.5,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-09-03",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_6_astra_low: pass_at_1 0.6703539823008849, ci_half 0.013008610516858731, n_runs 4, mean_cost_usd 1.595223439159292",
      "notes": "pass@4 79.6%; mean 10 min per attempt; cost basis: Current pricing at all context lengths: $10/M uncached input, $12.50/M cache writes, $1/M cache reads, $50/M output; no separate compute-unit fee.; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-6 Astra",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-tb7k4q",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-6-luna",
      "model_as_reported": "GPT 6 Luna",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 67,
      "ci_pct": 6.8,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "GPT 6 Luna | effort: max | score: 67% ± 6.8% | n_samples: 339 | provider: OpenAI",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "GPT-6 Luna",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-12m5fyd",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "glm-5.3",
      "model_as_reported": "GLM-5.3",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 66.9,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "GLM-5.3 HuggingFace Model Card",
      "source_url": "https://huggingface.co/zai-org/GLM-5.3",
      "published": "2026-08-25",
      "accessed": "2026-10-09",
      "quote": "| DeepSWE (v1.1) | 66.9 | 46.2 | 67.5 | 62.7 | 56.6 | 58.0 | 69.7 | **72.7** |",
      "notes": "Footnote: We run DeepSWE using the mini-swe-agent harness with temperature=0.95, top_p=1.0, timeout=6h and 400K context. (Official board re-run reached 69.0% ±3% on mini-swe-agent as reported in CodingFleet).",
      "display_name": "GLM-5.3",
      "lab": "Z.ai",
      "open_weights": true,
      "row_id": "row-1so5ru7",
      "source_id": "r-1171e08",
      "source_n": 14
    },
    {
      "model_key": "grok-4.6",
      "model_as_reported": "grok-4-6",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 66.74,
      "ci_pct": 2.18,
      "n_samples": 4,
      "cost_per_task_usd": 5.4977,
      "tokens_per_task": 9415426,
      "input_tokens_per_task": 9344022,
      "output_tokens_per_task": 71404,
      "steps_per_task": 87.2,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-08-12",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_grok_4_6_xhigh: pass_at_1 0.6674057649667405, ci_half 0.02175684224405049, n_runs 4, mean_cost_usd 5.497668745011087",
      "notes": "pass@4 85%; mean 21 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Grok 4.6",
      "lab": "xAI",
      "open_weights": false,
      "row_id": "row-huu76o",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-6-luna",
      "model_as_reported": "GPT-6 Luna",
      "effort": "max",
      "harness": "lab-internal",
      "score_pct": 66.59,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 0.2169,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6 Sol and Luna launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-sol-and-luna/",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-6 Luna\",\"cost_label\":\"$$0.2169\",\"score\":0.6659,\"cost\":0.2169,\"effortLabel\":\"max\"}",
      "notes": "Reported in embedded Vega-Lite chart in GPT-6 Sol and Luna launch post.",
      "display_name": "GPT-6 Luna",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-m482jr",
      "source_id": "r-1m7ewdf",
      "source_n": 26
    },
    {
      "model_key": "gpt-6-sol",
      "model_as_reported": "GPT-6 Sol",
      "effort": "xhigh",
      "harness": "lab-internal",
      "score_pct": 66.59,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 1.0033,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6 Sol and Luna launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-sol-and-luna/",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-6 Sol\",\"cost_label\":\"$$1.00\",\"score\":0.6659,\"cost\":1.0033,\"effortLabel\":\"xhigh\"}",
      "notes": "Reported in embedded Vega-Lite chart and text: 'On DeepSWE v1.1... GPT-6 Sol at max effort scores 68.8%, within 1.1 percentage points of Claude Fable 5’s highest score in the evaluation—69.9% at xhigh effort—at approximately 80% lower cost per task.'",
      "display_name": "GPT-6 Sol",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1d2hsmx",
      "source_id": "r-1m7ewdf",
      "source_n": 26
    },
    {
      "model_key": "kimi-k3",
      "model_as_reported": "K3 max",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 66.4,
      "ci_pct": null,
      "n_samples": 113,
      "cost_per_task_usd": 4.74,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Fireworks AI Ember-1 Announcement",
      "source_url": "https://fireworks.ai/blog/ember-1",
      "published": "2026-09-23",
      "accessed": "2026-10-09",
      "quote": "DeepSWE 1.1 | N=113 | K3 Low: 55.8% | K3 High: 62.8% | K3 max: 66.4% | Ember-1: 75.2% | -23.7% / -126.9 USD",
      "notes": "Evaluated by Fireworks AI across 113 tasks. Fireworks reports total run cost was $535.44 ($4.74/task), based on Kimi rates $3.00/M uncached, $0.30/M cached, $15.00/M output.",
      "display_name": "Kimi K3",
      "lab": "Moonshot AI",
      "open_weights": null,
      "row_id": "row-11vuxf9",
      "source_id": "r-1vlr43v",
      "source_n": 27
    },
    {
      "model_key": "deepseek-v4.1-flash",
      "model_as_reported": "DS-V4.1-Flash (Pi)",
      "effort": "max",
      "harness": "Pi",
      "score_pct": 66.2,
      "ci_pct": null,
      "n_samples": 8,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "DeepSeek-V4.1-Flash HuggingFace Model Card",
      "source_url": "https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash",
      "published": "2026-09-10",
      "accessed": "2026-10-09",
      "quote": "| DeepSWE v1.1 (Resolved) | 69.8 | 65.6 | 65.5 | 66.2 | 74.2 | 72.6 | 70.5 | 67.6 |",
      "notes": "Evaluated with Pi scaffold, reasoning_effort=100, N=8, temp=1.0, top_p=0.95, 1M context, max_steps=500.",
      "display_name": "DeepSeek V4.1 Flash",
      "lab": "DeepSeek",
      "open_weights": true,
      "row_id": "row-1wpexpo",
      "source_id": "r-1ftq9xa",
      "source_n": 20
    },
    {
      "model_key": "gpt-5.5",
      "model_as_reported": "GPT 5.5",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 66.1,
      "ci_pct": 6.9,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "GPT 5.5 | effort: xhigh | score: 66.1% ± 6.9% | n_samples: 339 | provider: OpenAI",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "GPT-5.5",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-e5sg10",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "grok-4.6",
      "model_as_reported": "Grok 4.6",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 65.9,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "xAI Grok 4.6 Model Card",
      "source_url": "https://media.x.ai/v1/website/card-4p6-4cd2dc57.pdf",
      "published": "2026-08-18",
      "accessed": "2026-10-09",
      "quote": "Grok 4.6 is evaluated with the mini-swe-agent harness. Grok 4.6 scores 65.9% at high thinking effort and 67.0% at xhigh thinking effort.",
      "notes": "Section 2.4 DeepSWE v1.1 (p. 11). Score at high thinking effort (reported as 65.2% in Grok 4.7 comparison table).",
      "display_name": "Grok 4.6",
      "lab": "xAI",
      "open_weights": false,
      "row_id": "row-18oxn14",
      "source_id": "r-tt3io3",
      "source_n": 12
    },
    {
      "model_key": "deepseek-v4.1-flash",
      "model_as_reported": "DS-V4.1-Flash (Codex)",
      "effort": "max",
      "harness": "Codex",
      "score_pct": 65.6,
      "ci_pct": null,
      "n_samples": 8,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "DeepSeek-V4.1-Flash HuggingFace Model Card",
      "source_url": "https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash",
      "published": "2026-09-10",
      "accessed": "2026-10-09",
      "quote": "| DeepSWE v1.1 (Resolved) | 69.8 | 65.6 | 65.5 | 66.2 | 74.2 | 72.6 | 70.5 | 67.6 |",
      "notes": "Evaluated with Codex scaffold, reasoning_effort=100, N=8, temp=1.0, top_p=0.95, 1M context, max_steps=500.",
      "display_name": "DeepSeek V4.1 Flash",
      "lab": "DeepSeek",
      "open_weights": true,
      "row_id": "row-ils0lb",
      "source_id": "r-1ftq9xa",
      "source_n": 20
    },
    {
      "model_key": "nex-n2.5-max",
      "model_as_reported": "Nex-N2.5-Max",
      "effort": "",
      "harness": "NexAU",
      "score_pct": 65.6,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Nex-AGI Nex-N2.5 GitHub README / Model Card",
      "source_url": "https://github.com/nex-agi/Nex-N2.5",
      "published": "2026-09-08",
      "accessed": "2026-10-09",
      "quote": "DeepSWE v1.1 | 36.1 | 55.8 | 65.6 | 73.7 | 72.7 | 67.5 | 66.9 | 62.8 | 69.3",
      "notes": "Coding tasks evaluated using the NexAU harness. Sampling parameters: temperature = 0.7, top_p = 0.95, top_k = 40. 1.6T parameter text-only MoE foundation model.",
      "display_name": "Nex-N2.5-Max",
      "lab": "Nex-AGI",
      "open_weights": true,
      "row_id": "row-1ogzpoy",
      "source_id": "r-pixn6e",
      "source_n": 18
    },
    {
      "model_key": "deepseek-v4.1-flash",
      "model_as_reported": "DS-V4.1-Flash (OpenCode)",
      "effort": "max",
      "harness": "OpenCode",
      "score_pct": 65.5,
      "ci_pct": null,
      "n_samples": 8,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "DeepSeek-V4.1-Flash HuggingFace Model Card",
      "source_url": "https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash",
      "published": "2026-09-10",
      "accessed": "2026-10-09",
      "quote": "| DeepSWE v1.1 (Resolved) | 69.8 | 65.6 | 65.5 | 66.2 | 74.2 | 72.6 | 70.5 | 67.6 |",
      "notes": "Evaluated with OpenCode scaffold, reasoning_effort=100, N=8, temp=1.0, top_p=0.95, 1M context, max_steps=500.",
      "display_name": "DeepSeek V4.1 Flash",
      "lab": "DeepSeek",
      "open_weights": true,
      "row_id": "row-1n2qy16",
      "source_id": "r-1ftq9xa",
      "source_n": 20
    },
    {
      "model_key": "gemini-3.7-flash",
      "model_as_reported": "Gemini 3.7 Flash",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 65.5,
      "ci_pct": 7.1,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Gemini 3.7 Flash | effort: high | score: 65.5% ± 7.1% | n_samples: 339 | provider: Google",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Gemini 3.7 Flash",
      "lab": "Google",
      "open_weights": false,
      "row_id": "row-1rhhb5e",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gemini-3.7-flash",
      "model_as_reported": "gemini-3-7-flash",
      "effort": "medium",
      "harness": "mini-swe-agent",
      "score_pct": 65.49,
      "ci_pct": 3.09,
      "n_samples": 4,
      "cost_per_task_usd": 2.0251,
      "tokens_per_task": 15426057,
      "input_tokens_per_task": 15332066,
      "output_tokens_per_task": 93991,
      "steps_per_task": 117.4,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-08-13",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gemini_3_7_flash_medium: pass_at_1 0.6548672566371682, ci_half 0.030865322764155222, n_runs 4, mean_cost_usd 2.0251351826880533",
      "notes": "pass@4 83.2%; mean 21 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Gemini 3.7 Flash",
      "lab": "Google",
      "open_weights": false,
      "row_id": "row-1q9if39",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "claude-fable-5",
      "model_as_reported": "claude-fable-5",
      "effort": "medium",
      "harness": "mini-swe-agent",
      "score_pct": 65.37,
      "ci_pct": 4.42,
      "n_samples": 4,
      "cost_per_task_usd": 6.0882,
      "tokens_per_task": 2982453,
      "input_tokens_per_task": 2942252,
      "output_tokens_per_task": 40201,
      "steps_per_task": 48.4,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_fable_5_medium: pass_at_1 0.6536697247706422, ci_half 0.044219288792381656, n_runs 4, mean_cost_usd 6.088186528935186",
      "notes": "pass@4 83.2%; mean 14 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Fable 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-f4ic5k",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gemini-3.7-flash",
      "model_as_reported": "gemini-3-7-flash",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 65.27,
      "ci_pct": 1.79,
      "n_samples": 4,
      "cost_per_task_usd": 2.1763,
      "tokens_per_task": 17107101,
      "input_tokens_per_task": 16999853,
      "output_tokens_per_task": 107248,
      "steps_per_task": 124.5,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-08-13",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gemini_3_7_flash_high: pass_at_1 0.6526548672566371, ci_half 0.017878625067843174, n_runs 4, mean_cost_usd 2.1762860351216813",
      "notes": "pass@4 82.3%; mean 21 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Gemini 3.7 Flash",
      "lab": "Google",
      "open_weights": false,
      "row_id": "row-z10ec",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-6-sol",
      "model_as_reported": "GPT-6 Sol",
      "effort": "high",
      "harness": "lab-internal",
      "score_pct": 65.27,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 0.6404,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6 Sol and Luna launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-sol-and-luna/",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-6 Sol\",\"cost_label\":\"$$0.64\",\"score\":0.6527,\"cost\":0.6404,\"effortLabel\":\"high\"}",
      "notes": "Reported in embedded Vega-Lite chart and text: 'On DeepSWE v1.1... GPT-6 Sol at max effort scores 68.8%, within 1.1 percentage points of Claude Fable 5’s highest score in the evaluation—69.9% at xhigh effort—at approximately 80% lower cost per task.'",
      "display_name": "GPT-6 Sol",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-wxw9bn",
      "source_id": "r-1m7ewdf",
      "source_n": 26
    },
    {
      "model_key": "grok-4.6",
      "model_as_reported": "grok-4-6",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 65.19,
      "ci_pct": 1.53,
      "n_samples": 4,
      "cost_per_task_usd": 4.3849,
      "tokens_per_task": 7409649,
      "input_tokens_per_task": 7348488,
      "output_tokens_per_task": 61161,
      "steps_per_task": 79,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-08-12",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_grok_4_6_high: pass_at_1 0.6518847006651884, ci_half 0.015336008186572951, n_runs 4, mean_cost_usd 4.3848732239467845",
      "notes": "pass@4 85%; mean 18 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Grok 4.6",
      "lab": "xAI",
      "open_weights": false,
      "row_id": "row-12wnw3m",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "grok-4.6",
      "model_as_reported": "Grok 4.6 (Grok Build)",
      "effort": "high",
      "harness": "Grok Build",
      "score_pct": 65,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Artificial Analysis (articles/benchmarking-grok-4-7)",
      "source_url": "https://artificialanalysis.ai/articles/benchmarking-grok-4-7",
      "published": "2026-09-21",
      "accessed": "2026-10-09",
      "quote": "Grok 4.6 | DeepSWE v1.1: 65.0% | Evaluated with Grok Build | Artificial Analysis Coding Agent Index",
      "notes": "Artificial Analysis benchmark baseline for Grok 4.6 on DeepSWE v1.1 component of Coding Agent Index.",
      "display_name": "Grok 4.6",
      "lab": "xAI",
      "open_weights": false,
      "row_id": "row-69o7y9",
      "source_id": "r-17v6sj",
      "source_n": 35
    },
    {
      "model_key": "claude-fable-5.1",
      "model_as_reported": "Fable 5.1",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 64.9,
      "ci_pct": 7.2,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Fable 5.1 | effort: max | score: 64.9% ± 7.2% | n_samples: 339 | provider: Anthropic",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Claude Fable 5.1",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-q5njty",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gpt-5.5",
      "model_as_reported": "gpt-5-5",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 64.38,
      "ci_pct": 3.12,
      "n_samples": 4,
      "cost_per_task_usd": 5.1004,
      "tokens_per_task": 4648941,
      "input_tokens_per_task": 4617782,
      "output_tokens_per_task": 31159,
      "steps_per_task": 61.9,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_5_high: pass_at_1 0.6438053097345132, ci_half 0.031168426495054677, n_runs 4, mean_cost_usd 5.100437004424779",
      "notes": "pass@4 90.3%; mean 34 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.5",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-10u6f80",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-6.1-sol",
      "model_as_reported": "GPT-6.1 Sol",
      "effort": "low",
      "harness": "lab-internal",
      "score_pct": 64.38,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 0.1714,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6.1 Sol launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-1-sol/",
      "published": "2026-09-29",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-6.1 Sol\",\"modelLabel\":\"GPT-6.1 Sol\",\"cost\":0.1714,\"score\":0.6438,\"effortLabel\":\"Low\"}",
      "notes": "Reported in embedded Vega-Lite spec on launch page.",
      "display_name": "GPT-6.1 Sol",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1u8z7hs",
      "source_id": "r-1ntqzs0",
      "source_n": 30
    },
    {
      "model_key": "gpt-5.6-luna",
      "model_as_reported": "GPT 5.6 Luna",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 64.3,
      "ci_pct": 7.4,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "GPT 5.6 Luna | effort: max | score: 64.3% ± 7.4% | n_samples: 339 | provider: OpenAI",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "GPT-5.6 Luna",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-y14yn6",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "hy4-preview",
      "model_as_reported": "Hy4 preview",
      "effort": "",
      "harness": "lab-internal",
      "score_pct": 64.3,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Tencent Hy4-preview Technical Report & Model Card",
      "source_url": "https://huggingface.co/tencent/Hy4-preview",
      "published": "2026-08-27",
      "accessed": "2026-10-09",
      "quote": "Hy4 preview achieved 64.3% on DeepSWE in the official Benchmark Appendix (up from 28.0% on Hy3).",
      "notes": "770B total / 49B active MoE flagship model. Evaluated on DeepSWE in official Benchmark Appendix.",
      "display_name": "Hy4 Preview",
      "lab": "Tencent",
      "open_weights": true,
      "row_id": "row-lcrv5b",
      "source_id": "r-p3aj15",
      "source_n": 15
    },
    {
      "model_key": "grok-4.6",
      "model_as_reported": "Grok 4.6",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 63.7,
      "ci_pct": 7.4,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Grok 4.6 | effort: high | score: 63.7% ± 7.4% | n_samples: 339 | provider: xAI",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Grok 4.6",
      "lab": "xAI",
      "open_weights": false,
      "row_id": "row-2clyyf",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "glm-5.3-flash",
      "model_as_reported": "glm-5-3-flash",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 63.39,
      "ci_pct": 4.38,
      "n_samples": 4,
      "cost_per_task_usd": 0.241,
      "tokens_per_task": 12459855,
      "input_tokens_per_task": 12387025,
      "output_tokens_per_task": 72830,
      "steps_per_task": 122.9,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-08-26",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_glm_5_3_flash_max: pass_at_1 0.6339285714285714, ci_half 0.04377970751693644, n_runs 4, mean_cost_usd 0.24098185622767856",
      "notes": "pass@4 85%; mean 26 min per attempt; leaderboard-live.json gives $0.48 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GLM-5.3 Flash",
      "lab": "Z.ai",
      "open_weights": true,
      "row_id": "row-pa1eq4",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "deepseek-v4-pro",
      "model_as_reported": "deepseek-v4-pro",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 62.83,
      "ci_pct": 6.33,
      "n_samples": 4,
      "cost_per_task_usd": 1.666,
      "tokens_per_task": 24297605,
      "input_tokens_per_task": 24191606,
      "output_tokens_per_task": 105999,
      "steps_per_task": 154.7,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-08-12",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_deepseek_v4_pro_max: pass_at_1 0.6283185840707964, ci_half 0.06333430597228876, n_runs 4, mean_cost_usd 1.6660232187256647",
      "notes": "pass@4 88.5%; mean 37 min per attempt; leaderboard-live.json gives $0.24 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "DeepSeek V4 Pro",
      "lab": "DeepSeek",
      "open_weights": true,
      "row_id": "row-eakg9h",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "kimi-k3",
      "model_as_reported": "K3 High",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 62.8,
      "ci_pct": null,
      "n_samples": 113,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Fireworks AI Ember-1 Announcement",
      "source_url": "https://fireworks.ai/blog/ember-1",
      "published": "2026-09-23",
      "accessed": "2026-10-09",
      "quote": "DeepSWE 1.1 | N=113 | K3 Low: 55.8% | K3 High: 62.8% | K3 max: 66.4% | Ember-1: 75.2%",
      "notes": "Evaluated by Fireworks AI across all 113 DeepSWE 1.1 tasks under high reasoning effort configuration.",
      "display_name": "Kimi K3",
      "lab": "Moonshot AI",
      "open_weights": null,
      "row_id": "row-2nx1rj",
      "source_id": "r-1vlr43v",
      "source_n": 27
    },
    {
      "model_key": "deepseek-v4-pro",
      "model_as_reported": "DS-V4-Pro",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 62.7,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "DeepSeek-V4.1-Flash HuggingFace Model Card",
      "source_url": "https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash",
      "published": "2026-09-10",
      "accessed": "2026-10-09",
      "quote": "| DeepSWE v1.1 (Resolved) | 74.0 | 73.0 | 67.5 | 66.9 | 62.7 | 54.4 | **74.2** |",
      "notes": "Reported under frontier comparisons table in DS-V4.1-Flash README for DS-V4-Pro at max reasoning effort using mini-SWE harness.",
      "display_name": "DeepSeek V4 Pro",
      "lab": "DeepSeek",
      "open_weights": true,
      "row_id": "row-1fioxzl",
      "source_id": "r-1ftq9xa",
      "source_n": 20
    },
    {
      "model_key": "gpt-5.6-luna",
      "model_as_reported": "GPT-5.6 Luna",
      "effort": "max",
      "harness": "lab-internal",
      "score_pct": 62.17,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 0.532,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6 Sol and Luna launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-sol-and-luna/",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-5.6 Luna\",\"cost_label\":\"$$0.532\",\"score\":0.6217,\"cost\":0.532,\"effortLabel\":\"max\"}",
      "notes": "OpenAI re-ran GPT-5.6 Luna internally for its GPT-6 Sol/Luna post; the official board measured 67.2% at max. Explains discrepancy: 62.17% was OpenAI internal re-evaluation chart data alongside GPT-6 Sol/Luna, whereas 67.2% (67.19%) was Datacurve official board 4-pass verified rollout.",
      "display_name": "GPT-5.6 Luna",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1csb5tq",
      "source_id": "r-1m7ewdf",
      "source_n": 26
    },
    {
      "model_key": "mistral-large-4",
      "model_as_reported": "Mistral Large 4",
      "effort": "thinking",
      "harness": "partner eval (Artificial Analysis, Surge AI)",
      "score_pct": 61.7,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Mistral Large 4 Announcement",
      "source_url": "https://mistral.ai/news/mistral-large-4/",
      "published": "2026-10-06",
      "accessed": "2026-10-09",
      "quote": "scoring 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4.",
      "notes": "Mistral reports the run was done with Artificial Analysis and Surge AI. Evaluated in partnership with Artificial Analysis and Surge AI using Artificial Analysis coding agent harness. Sole Mistral model with DeepSWE 1.1 score.; harness as reported: other name",
      "display_name": "Mistral Large 4",
      "lab": "Mistral",
      "open_weights": true,
      "row_id": "row-2arw03",
      "source_id": "r-d2tk0p",
      "source_n": 32
    },
    {
      "model_key": "gpt-6-luna",
      "model_as_reported": "GPT-6 Luna",
      "effort": "xhigh",
      "harness": "lab-internal",
      "score_pct": 61.28,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 0.1096,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6 Sol and Luna launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-sol-and-luna/",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-6 Luna\",\"cost_label\":\"$$0.1096\",\"score\":0.6128,\"cost\":0.1096,\"effortLabel\":\"xhigh\"}",
      "notes": "Reported in embedded Vega-Lite chart in GPT-6 Sol and Luna launch post.",
      "display_name": "GPT-6 Luna",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-ddelns",
      "source_id": "r-1m7ewdf",
      "source_n": 26
    },
    {
      "model_key": "gpt-5.6-sol",
      "model_as_reported": "gpt-5-6-sol",
      "effort": "medium",
      "harness": "mini-swe-agent",
      "score_pct": 61.06,
      "ci_pct": 1.58,
      "n_samples": 4,
      "cost_per_task_usd": 1.4158,
      "tokens_per_task": 1524219,
      "input_tokens_per_task": 1505794,
      "output_tokens_per_task": 18425,
      "steps_per_task": 30.9,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-10",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_6_sol_medium: pass_at_1 0.6106194690265486, ci_half 0.01583357649307215, n_runs 4, mean_cost_usd 1.4157802656590275",
      "notes": "pass@4 80.5%; mean 7 min per attempt; leaderboard-live.json gives $1.86 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.6 Sol",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-5a37tj",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-5.6-terra",
      "model_as_reported": "gpt-5-6-terra",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 60.18,
      "ci_pct": 2.12,
      "n_samples": 4,
      "cost_per_task_usd": 1.7018,
      "tokens_per_task": 3285972,
      "input_tokens_per_task": 3246355,
      "output_tokens_per_task": 39617,
      "steps_per_task": 43.1,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-10",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_6_terra_xhigh: pass_at_1 0.6017699115044248, ci_half 0.021242972019271257, n_runs 4, mean_cost_usd 1.7017533805309735",
      "notes": "pass@4 80.5%; mean 10 min per attempt; leaderboard-live.json gives $2.13 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.6 Terra",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-ouwfw",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "claude-haiku-5.5",
      "model_as_reported": "Haiku 5.5",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 59.9,
      "ci_pct": 7.7,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX OSS DeepSWE Leaderboard",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Haiku 5.5 | effort: max | score: 59.9% ± 7.7% | n_samples: 339 | provider: Anthropic",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Independent run on Mercor APEX. Confirms no Anthropic official DeepSWE figure exists.",
      "display_name": "Claude Haiku 5.5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-bm6mdq",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "claude-opus-4.8",
      "model_as_reported": "Opus 4.8",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 59.6,
      "ci_pct": 7.1,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Opus 4.8 | effort: max | score: 59.6% ± 7.1% | n_samples: 339 | provider: Anthropic",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Claude Opus 4.8",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1gtq4fm",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "claude-fable-5",
      "model_as_reported": "claude-fable-5",
      "effort": "low",
      "harness": "mini-swe-agent",
      "score_pct": 59.58,
      "ci_pct": 2.79,
      "n_samples": 4,
      "cost_per_task_usd": 3.7579,
      "tokens_per_task": 1716493,
      "input_tokens_per_task": 1691250,
      "output_tokens_per_task": 25243,
      "steps_per_task": 37.8,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_fable_5_low: pass_at_1 0.5958429561200924, ci_half 0.027940054451297294, n_runs 4, mean_cost_usd 3.757873125874126",
      "notes": "pass@4 81.4%; mean 11 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Fable 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-18j1o47",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-6-luna",
      "model_as_reported": "GPT-6 Luna",
      "effort": "high",
      "harness": "lab-internal",
      "score_pct": 59.29,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 0.0838,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6 Sol and Luna launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-sol-and-luna/",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-6 Luna\",\"cost_label\":\"$$0.0838\",\"score\":0.5929,\"cost\":0.0838,\"effortLabel\":\"high\"}",
      "notes": "Reported in embedded Vega-Lite chart in GPT-6 Sol and Luna launch post.",
      "display_name": "GPT-6 Luna",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-3vfxsy",
      "source_id": "r-1m7ewdf",
      "source_n": 26
    },
    {
      "model_key": "claude-opus-4.8",
      "model_as_reported": "claude-opus-4-8",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 58.97,
      "ci_pct": 1.76,
      "n_samples": 4,
      "cost_per_task_usd": 13.2226,
      "tokens_per_task": 17182970,
      "input_tokens_per_task": 17047939,
      "output_tokens_per_task": 135032,
      "steps_per_task": 120,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_opus_4_8_max: pass_at_1 0.5897435897435898, ci_half 0.01764815347429254, n_runs 4, mean_cost_usd 13.222583593240094",
      "notes": "pass@4 79.3%; mean 58 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Opus 4.8",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1ojpt5b",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "qwen3.8-flash-next",
      "model_as_reported": "Qwen3.8-Flash-Next",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 58.7,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Qwen3.8-Flash-Next HuggingFace Model Card",
      "source_url": "https://huggingface.co/Qwen/Qwen3.8-Flash-Next",
      "published": "2026-08-24",
      "accessed": "2026-10-09",
      "quote": "<td class=\"benchmark-cell\" ...><div class=\"benchmark-capability\" ...>Agentic coding</div><div class=\"benchmark-name\" ...>DeepSWE 1.1</div></td>\n<td ...><strong>58.7</strong></td>",
      "notes": "Footnote 1: DeepSWE 1.1: evaluated with the Claude Code and mini-SWE-agent harnesses, temp=1.0, top_p=0.95, 256K context window. We report the highest score across the two harnesses; notably, Qwen3.8-Flash-Next performs best on mini-SWE-agent.",
      "display_name": "Qwen3.8-Flash-Next",
      "lab": "Alibaba",
      "open_weights": true,
      "row_id": "row-kipywi",
      "source_id": "r-qxom7y",
      "source_n": 13
    },
    {
      "model_key": "claude-opus-5",
      "model_as_reported": "claude-opus-5",
      "effort": "low",
      "harness": "mini-swe-agent",
      "score_pct": 58.13,
      "ci_pct": 2.33,
      "n_samples": 4,
      "cost_per_task_usd": 1.6626,
      "tokens_per_task": 1584172,
      "input_tokens_per_task": 1564288,
      "output_tokens_per_task": 19884,
      "steps_per_task": 35.6,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-25",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_opus_5_low: pass_at_1 0.5812917594654788, ci_half 0.023306879466842723, n_runs 4, mean_cost_usd 1.6626223919821828",
      "notes": "pass@4 85%; mean 8 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Opus 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-blybci",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "qwen3.8-max",
      "model_as_reported": "qwen3-8-max",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 57.46,
      "ci_pct": 2.66,
      "n_samples": 4,
      "cost_per_task_usd": 3.7291,
      "tokens_per_task": 13091142,
      "input_tokens_per_task": 12996066,
      "output_tokens_per_task": 95075,
      "steps_per_task": 111.3,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-08-04",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_qwen3_8_max_xhigh: pass_at_1 0.5746102449888641, ci_half 0.026627926725075454, n_runs 4, mean_cost_usd 3.7290646500445437",
      "notes": "pass@4 83.2%; mean 43 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Qwen3.8 Max",
      "lab": "Alibaba",
      "open_weights": true,
      "row_id": "row-1nr3gb4",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-5.6-luna",
      "model_as_reported": "gpt-5-6-luna",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 56.86,
      "ci_pct": 2.17,
      "n_samples": 4,
      "cost_per_task_usd": 0.3071,
      "tokens_per_task": 7655255,
      "input_tokens_per_task": 7610578,
      "output_tokens_per_task": 44678,
      "steps_per_task": 71.1,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-10",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_6_luna_xhigh: pass_at_1 0.5685840707964602, ci_half 0.021681017528097975, n_runs 4, mean_cost_usd 0.30712385371681417",
      "notes": "pass@4 78.8%; mean 12 min per attempt; leaderboard-live.json gives $1.54 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.6 Luna",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-19e0252",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-6-sol",
      "model_as_reported": "GPT-6 Sol",
      "effort": "medium",
      "harness": "lab-internal",
      "score_pct": 56.64,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 0.3798,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6 Sol and Luna launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-sol-and-luna/",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-6 Sol\",\"cost_label\":\"$$0.38\",\"score\":0.5664,\"cost\":0.3798,\"effortLabel\":\"medium\"}",
      "notes": "Reported in embedded Vega-Lite chart and text: 'On DeepSWE v1.1... GPT-6 Sol at max effort scores 68.8%, within 1.1 percentage points of Claude Fable 5’s highest score in the evaluation—69.9% at xhigh effort—at approximately 80% lower cost per task.'",
      "display_name": "GPT-6 Sol",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-vphi25",
      "source_id": "r-1m7ewdf",
      "source_n": 26
    },
    {
      "model_key": "deepseek-v4-flash",
      "model_as_reported": "DeepSeek V4 Flash",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 56.6,
      "ci_pct": 8.8,
      "n_samples": 113,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "DeepSeek V4 Flash | effort: max | score: 56.6% ± 8.8% | n_samples: 113 | provider: DeepSeek",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "DeepSeek V4 Flash",
      "lab": "DeepSeek",
      "open_weights": true,
      "row_id": "row-1vnfpxa",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "qwen3.8-max",
      "model_as_reported": "Qwen3.8-Max",
      "effort": "",
      "harness": "Claude Code",
      "score_pct": 56.6,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Qwen3.8-2.4T-A95B (Qwen3.8-Max) HuggingFace Model Card",
      "source_url": "https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B",
      "published": "2026-08-08",
      "accessed": "2026-10-09",
      "quote": "DeepSWE 1.1 | 59.0 | 70.0 | 73.0 | 21.6 | 56.6",
      "notes": "Footnote 4: DeepSWE 1.1: Evaluated with the Claude Code and mini-SWE-agent harnesses, temp=1.0, top_p=0.95, and a 256K context window. We report the highest score among both harnesses; notably, Qwen3.8-Max performs best on Claude Code.",
      "display_name": "Qwen3.8 Max",
      "lab": "Alibaba",
      "open_weights": true,
      "row_id": "row-17gjj3j",
      "source_id": "r-ouk7wk",
      "source_n": 11
    },
    {
      "model_key": "deepseek-v4-pro-0813",
      "model_as_reported": "DeepSeek V4 Pro 0813",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 56.3,
      "ci_pct": 7.2,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "DeepSeek V4 Pro 0813 | effort: max | score: 56.3% ± 7.2% | n_samples: 339 | provider: DeepSeek",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "DeepSeek V4 Pro 0813",
      "lab": "DeepSeek",
      "open_weights": true,
      "row_id": "row-10qeqvl",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gpt-5.6-luna",
      "model_as_reported": "GPT-5.6 Luna",
      "effort": "xhigh",
      "harness": "lab-internal",
      "score_pct": 56.19,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 0.2737,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6 Sol and Luna launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-sol-and-luna/",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-5.6 Luna\",\"cost_label\":\"$$0.2737\",\"score\":0.5619,\"cost\":0.2737,\"effortLabel\":\"xhigh\"}",
      "notes": "OpenAI re-ran GPT-5.6 Luna internally for its GPT-6 Sol/Luna post; the official board measured 67.2% at max. Full effort breakdown from low to max reported in GPT-6 Sol and Luna launch post chart data.",
      "display_name": "GPT-5.6 Luna",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-15b4q5d",
      "source_id": "r-1m7ewdf",
      "source_n": 26
    },
    {
      "model_key": "kimi-k3",
      "model_as_reported": "K3 Low",
      "effort": "low",
      "harness": "mini-swe-agent",
      "score_pct": 55.8,
      "ci_pct": null,
      "n_samples": 113,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Fireworks AI Ember-1 Announcement",
      "source_url": "https://fireworks.ai/blog/ember-1",
      "published": "2026-09-23",
      "accessed": "2026-10-09",
      "quote": "DeepSWE 1.1 | N=113 | K3 Low: 55.8% | K3 High: 62.8% | K3 max: 66.4% | Ember-1: 75.2%",
      "notes": "Evaluated by Fireworks AI across all 113 DeepSWE 1.1 tasks under low reasoning effort configuration.",
      "display_name": "Kimi K3",
      "lab": "Moonshot AI",
      "open_weights": null,
      "row_id": "row-uvwqhb",
      "source_id": "r-1vlr43v",
      "source_n": 27
    },
    {
      "model_key": "nex-n2.5-pro",
      "model_as_reported": "Nex-N2.5-Pro",
      "effort": "",
      "harness": "NexAU",
      "score_pct": 55.8,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Nex-AGI Nex-N2.5 GitHub README / Model Card",
      "source_url": "https://github.com/nex-agi/Nex-N2.5",
      "published": "2026-09-08",
      "accessed": "2026-10-09",
      "quote": "DeepSWE v1.1 | 36.1 | 55.8 | 65.6 | 73.7 | 72.7 | 67.5 | 66.9 | 62.8 | 69.3",
      "notes": "Coding tasks evaluated using the NexAU harness. Sampling parameters: temperature = 0.7, top_p = 0.95, top_k = 40. Mid-sized multimodal agentic model.",
      "display_name": "Nex-N2.5-Pro",
      "lab": "Nex-AGI",
      "open_weights": true,
      "row_id": "row-1bzp4ea",
      "source_id": "r-pixn6e",
      "source_n": 18
    },
    {
      "model_key": "muse-spark-1.2",
      "model_as_reported": "muse-spark-1-2",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 54.87,
      "ci_pct": 2.12,
      "n_samples": 4,
      "cost_per_task_usd": 3.6955,
      "tokens_per_task": 13793051,
      "input_tokens_per_task": 13693824,
      "output_tokens_per_task": 99226,
      "steps_per_task": 100.8,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-08-07",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_muse_spark_1_2_xhigh: pass_at_1 0.5486725663716814, ci_half 0.02124297201927127, n_runs 4, mean_cost_usd 3.6955048180309733",
      "notes": "pass@4 81.4%; mean 15 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Muse Spark 1.2",
      "lab": "Meta",
      "open_weights": false,
      "row_id": "row-qte2qf",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "deepseek-v4-flash",
      "model_as_reported": "DS-V4-Flash",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 54.4,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "DeepSeek-V4.1-Flash HuggingFace Model Card",
      "source_url": "https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash",
      "published": "2026-09-10",
      "accessed": "2026-10-09",
      "quote": "| DeepSWE v1.1 (Resolved) | 74.0 | 73.0 | 67.5 | 66.9 | 62.7 | 54.4 | **74.2** |",
      "notes": "Reported under frontier comparisons table in DS-V4.1-Flash README for DS-V4-Flash (also referenced as DeepSeek-V4-Flash-0731 in Qwen3.8-Flash-Next README).",
      "display_name": "DeepSeek V4 Flash",
      "lab": "DeepSeek",
      "open_weights": true,
      "row_id": "row-1yrhm2q",
      "source_id": "r-1ftq9xa",
      "source_n": 20
    },
    {
      "model_key": "claude-opus-4.8",
      "model_as_reported": "claude-opus-4-8",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 54.36,
      "ci_pct": 3.71,
      "n_samples": 4,
      "cost_per_task_usd": 8.0064,
      "tokens_per_task": 9940741,
      "input_tokens_per_task": 9854652,
      "output_tokens_per_task": 86089,
      "steps_per_task": 94.6,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_opus_4_8_xhigh: pass_at_1 0.5436241610738255, ci_half 0.03713430628094476, n_runs 4, mean_cost_usd 8.006355704138702",
      "notes": "pass@4 80.5%; mean 37 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Opus 4.8",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-f3dla0",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "claude-opus-4.7",
      "model_as_reported": "claude-opus-4-7",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 54.2,
      "ci_pct": 4.72,
      "n_samples": null,
      "cost_per_task_usd": 18.19,
      "tokens_per_task": 28617242,
      "input_tokens_per_task": 28513814,
      "output_tokens_per_task": 103428,
      "steps_per_task": 202.8,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "claude-opus-4-7 [max]: 54.2% ±4.72%, cost $18.19, output tokens 103428, steps 202.8",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "Claude Opus 4.7",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-z88wg6",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "gpt-5.5",
      "model_as_reported": "gpt-5-5",
      "effort": "medium",
      "harness": "mini-swe-agent",
      "score_pct": 53.98,
      "ci_pct": 2.55,
      "n_samples": 4,
      "cost_per_task_usd": 2.7493,
      "tokens_per_task": 2458164,
      "input_tokens_per_task": 2438539,
      "output_tokens_per_task": 19625,
      "steps_per_task": 46,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_5_medium: pass_at_1 0.5398230088495575, ci_half 0.02553087495290979, n_runs 4, mean_cost_usd 2.749342734513274",
      "notes": "pass@4 77.9%; mean 21 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.5",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1meuekt",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "claude-sonnet-5",
      "model_as_reported": "claude-sonnet-5",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 53.85,
      "ci_pct": 4.24,
      "n_samples": 4,
      "cost_per_task_usd": 26.3999,
      "tokens_per_task": 72629874,
      "input_tokens_per_task": 72415756,
      "output_tokens_per_task": 214118,
      "steps_per_task": 268.5,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-02",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_sonnet_5_max: pass_at_1 0.5384615384615384, ci_half 0.04236916470174386, n_runs 4, mean_cost_usd 26.399858950791852",
      "notes": "pass@4 78.8%; mean 80 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Sonnet 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1q1xs9y",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gemini-3.7-flash",
      "model_as_reported": "gemini-3-7-flash",
      "effort": "low",
      "harness": "mini-swe-agent",
      "score_pct": 53.76,
      "ci_pct": 2.59,
      "n_samples": 4,
      "cost_per_task_usd": 1.8323,
      "tokens_per_task": 15129930,
      "input_tokens_per_task": 15056565,
      "output_tokens_per_task": 73365,
      "steps_per_task": 130.4,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-08-13",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gemini_3_7_flash_low: pass_at_1 0.5376106194690266, ci_half 0.02589649081831869, n_runs 4, mean_cost_usd 1.8322645629424779",
      "notes": "pass@4 77.9%; mean 21 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Gemini 3.7 Flash",
      "lab": "Google",
      "open_weights": false,
      "row_id": "row-7fe1ld",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-5.6-terra",
      "model_as_reported": "gpt-5-6-terra",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 53.76,
      "ci_pct": 4.33,
      "n_samples": 4,
      "cost_per_task_usd": 0.9075,
      "tokens_per_task": 1578570,
      "input_tokens_per_task": 1557053,
      "output_tokens_per_task": 21517,
      "steps_per_task": 33.5,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-10",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_6_terra_high: pass_at_1 0.5376106194690266, ci_half 0.04328970467213552, n_runs 4, mean_cost_usd 0.9075010212389382",
      "notes": "pass@4 80.5%; mean 6 min per attempt; leaderboard-live.json gives $1.13 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.6 Terra",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1imalva",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "grok-4.5",
      "model_as_reported": "grok-4-5",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 53.76,
      "ci_pct": 2.28,
      "n_samples": 4,
      "cost_per_task_usd": 2.4157,
      "tokens_per_task": 4094708,
      "input_tokens_per_task": 4059183,
      "output_tokens_per_task": 35525,
      "steps_per_task": 61.3,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-16",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_grok_4_5_high: pass_at_1 0.5376106194690266, ci_half 0.0228080457287792, n_runs 4, mean_cost_usd 2.4157481725663716",
      "notes": "pass@4 77.9%; mean 8 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Grok 4.5",
      "lab": "xAI",
      "open_weights": false,
      "row_id": "row-1jii3v9",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "deepseek-v4-flash",
      "model_as_reported": "deepseek-v4-flash",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 53.32,
      "ci_pct": 3.57,
      "n_samples": 4,
      "cost_per_task_usd": 0.4641,
      "tokens_per_task": 19937618,
      "input_tokens_per_task": 19829931,
      "output_tokens_per_task": 107687,
      "steps_per_task": 152.9,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-08-06",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_deepseek_v4_flash_max: pass_at_1 0.5331858407079646, ci_half 0.035669502150324287, n_runs 4, mean_cost_usd 0.46407388012739975",
      "notes": "pass@4 80.5%; mean 24 min per attempt; leaderboard-live.json gives $0.10 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "DeepSeek V4 Flash",
      "lab": "DeepSeek",
      "open_weights": true,
      "row_id": "row-2xh2qw",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "muse-spark-1.1",
      "model_as_reported": "muse-spark-1-1",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 53.32,
      "ci_pct": 3.04,
      "n_samples": 4,
      "cost_per_task_usd": 2.3611,
      "tokens_per_task": 12079408,
      "input_tokens_per_task": 12005399,
      "output_tokens_per_task": 74008,
      "steps_per_task": 95.8,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-14",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_muse_spark_1_1_xhigh: pass_at_1 0.5331858407079646, ci_half 0.030353424539337114, n_runs 4, mean_cost_usd 2.361142644469026",
      "notes": "pass@4 79.6%; mean 15 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Muse Spark 1.1",
      "lab": "Meta",
      "open_weights": false,
      "row_id": "row-ylw86l",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "qwen3.8-max",
      "model_as_reported": "Qwen 3.8 Max",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 53.1,
      "ci_pct": 8.8,
      "n_samples": 113,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Qwen 3.8 Max | effort: xhigh | score: 53.1% ± 8.8% | n_samples: 113 | provider: Alibaba",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Qwen3.8 Max",
      "lab": "Alibaba",
      "open_weights": true,
      "row_id": "row-e8s2o9",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "claude-opus-4.8",
      "model_as_reported": "claude-opus-4-8",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 51.77,
      "ci_pct": 4.56,
      "n_samples": 4,
      "cost_per_task_usd": 4.2819,
      "tokens_per_task": 4958751,
      "input_tokens_per_task": 4908687,
      "output_tokens_per_task": 50064,
      "steps_per_task": 72.5,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_opus_4_8_high: pass_at_1 0.5176991150442478, ci_half 0.04561609145755843, n_runs 4, mean_cost_usd 4.281940017699115",
      "notes": "pass@4 77.9%; mean 18 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Opus 4.8",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1o3rvd8",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-5.4",
      "model_as_reported": "gpt-5-4",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 51.77,
      "ci_pct": 1.5,
      "n_samples": 4,
      "cost_per_task_usd": 5.6525,
      "tokens_per_task": 9490330,
      "input_tokens_per_task": 9418921,
      "output_tokens_per_task": 71409,
      "steps_per_task": 70.5,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_4_xhigh: pass_at_1 0.5176991150442478, ci_half 0.015021049567382833, n_runs 4, mean_cost_usd 5.6524577245575225",
      "notes": "pass@4 77.9%; mean 23 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.4",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-2ndzjq",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "claude-sonnet-5",
      "model_as_reported": "claude-sonnet-5",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 49.67,
      "ci_pct": 3.45,
      "n_samples": 4,
      "cost_per_task_usd": 11.8906,
      "tokens_per_task": 30816326,
      "input_tokens_per_task": 30695628,
      "output_tokens_per_task": 120699,
      "steps_per_task": 185.5,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-02",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_sonnet_5_xhigh: pass_at_1 0.49667405764966743, ci_half 0.034545489916778645, n_runs 4, mean_cost_usd 11.890634540022173",
      "notes": "pass@4 75.2%; mean 40 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Sonnet 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-k22gcp",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "claude-sonnet-5",
      "model_as_reported": "Sonnet 5",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 49.6,
      "ci_pct": 7.2,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Sonnet 5 | effort: max | score: 49.6% ± 7.2% | n_samples: 339 | provider: Anthropic",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Claude Sonnet 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-495gbu",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "claude-opus-4.7",
      "model_as_reported": "Opus 4.7",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 49.3,
      "ci_pct": 6.9,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Opus 4.7 | effort: max | score: 49.3% ± 6.9% | n_samples: 339 | provider: Anthropic",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Claude Opus 4.7",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1sa9dgv",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gemini-3.6-flash",
      "model_as_reported": "Gemini 3.6 Flash",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 49.3,
      "ci_pct": 7.1,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Gemini 3.6 Flash | effort: high | score: 49.3% ± 7.1% | n_samples: 339 | provider: Google",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Gemini 3.6 Flash",
      "lab": "Google",
      "open_weights": false,
      "row_id": "row-v4yd6l",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "grok-4.5",
      "model_as_reported": "Grok 4.5",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 49.3,
      "ci_pct": 7.4,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Grok 4.5 | effort: high | score: 49.3% ± 7.4% | n_samples: 339 | provider: xAI",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Grok 4.5",
      "lab": "xAI",
      "open_weights": false,
      "row_id": "row-dvkuwq",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gpt-5.4",
      "model_as_reported": "GPT 5.4",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 49,
      "ci_pct": 6.8,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "GPT 5.4 | effort: xhigh | score: 49% ± 6.8% | n_samples: 339 | provider: OpenAI",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "GPT-5.4",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-oo7mzf",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "claude-opus-4.8",
      "model_as_reported": "claude-opus-4-8",
      "effort": "medium",
      "harness": "mini-swe-agent",
      "score_pct": 48.67,
      "ci_pct": 2.24,
      "n_samples": 4,
      "cost_per_task_usd": 3.4439,
      "tokens_per_task": 3884916,
      "input_tokens_per_task": 3843603,
      "output_tokens_per_task": 41313,
      "steps_per_task": 65.6,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_opus_4_8_medium: pass_at_1 0.48672566371681414, ci_half 0.02239205861737452, n_runs 4, mean_cost_usd 3.4438571028761062",
      "notes": "pass@4 76.1%; mean 15 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Opus 4.8",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1gl1fwy",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "claude-sonnet-5",
      "model_as_reported": "claude-sonnet-5",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 48.23,
      "ci_pct": 4.51,
      "n_samples": 4,
      "cost_per_task_usd": 7.4256,
      "tokens_per_task": 18341786,
      "input_tokens_per_task": 18254491,
      "output_tokens_per_task": 87295,
      "steps_per_task": 146.6,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-02",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_sonnet_5_high: pass_at_1 0.4823008849557522, ci_half 0.045063148702148406, n_runs 4, mean_cost_usd 7.425560161197339",
      "notes": "pass@4 79.6%; mean 29 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Sonnet 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1ca4w9c",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gemini-3.6-flash",
      "model_as_reported": "gemini-3-6-flash",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 46.68,
      "ci_pct": 3.7,
      "n_samples": 4,
      "cost_per_task_usd": 2.2095,
      "tokens_per_task": 12691802,
      "input_tokens_per_task": 12595957,
      "output_tokens_per_task": 95845,
      "steps_per_task": 116.7,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-22",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gemini_3_6_flash_high: pass_at_1 0.4668141592920354, ci_half 0.037048538992472776, n_runs 4, mean_cost_usd 2.209506638666667",
      "notes": "pass@4 75.2%; mean 25 min per attempt; leaderboard-live.json gives $4.42 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Gemini 3.6 Flash",
      "lab": "Google",
      "open_weights": false,
      "row_id": "row-qiuq8l",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "glm-5.2",
      "model_as_reported": "GLM-5.2",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 46.2,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "GLM-5.2 HuggingFace Model Card",
      "source_url": "https://huggingface.co/zai-org/GLM-5.2",
      "published": "2026-06-16",
      "accessed": "2026-10-09",
      "quote": "| DeepSWE | 46.2 | 18 | 18 | 20 | 8 | 58 | 70 | 10 |",
      "notes": "Footnote: We run DeepSWE with the official pier evaluation framework and the mini-swe-agent harness (temperature=1.0, top_p=1.0, timeout=2h, 400K context). Each task is solved in an isolated container with 2 CPUs, 8 GB RAM, and no internet access.",
      "display_name": "GLM-5.2",
      "lab": "Z.ai",
      "open_weights": true,
      "row_id": "row-1138rgg",
      "source_id": "r-11h0zp7",
      "source_n": 6
    },
    {
      "model_key": "gpt-5.6-sol",
      "model_as_reported": "gpt-5-6-sol",
      "effort": "low",
      "harness": "mini-swe-agent",
      "score_pct": 45.35,
      "ci_pct": 2.39,
      "n_samples": 4,
      "cost_per_task_usd": 0.817,
      "tokens_per_task": 701413,
      "input_tokens_per_task": 690834,
      "output_tokens_per_task": 10579,
      "steps_per_task": 23.4,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-10",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_6_sol_low: pass_at_1 0.45353982300884954, ci_half 0.0238819467145892, n_runs 4, mean_cost_usd 0.8170052734939612",
      "notes": "pass@4 71.7%; mean 4 min per attempt; leaderboard-live.json gives $1.07 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.6 Sol",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-255enc",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "claude-opus-4.7",
      "model_as_reported": "claude-opus-4-7",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 44.69,
      "ci_pct": 2.88,
      "n_samples": null,
      "cost_per_task_usd": 8.58,
      "tokens_per_task": 12224722,
      "input_tokens_per_task": 12159101,
      "output_tokens_per_task": 65621,
      "steps_per_task": 125.9,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "claude-opus-4-7 [xhigh]: 44.69% ±2.88%, cost $8.58, output tokens 65621, steps 125.9",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "Claude Opus 4.7",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-13997ak",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "gpt-6-luna",
      "model_as_reported": "GPT-6 Luna",
      "effort": "medium",
      "harness": "lab-internal",
      "score_pct": 44.47,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 0.0518,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6 Sol and Luna launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-sol-and-luna/",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-6 Luna\",\"cost_label\":\"$$0.0518\",\"score\":0.4447,\"cost\":0.0518,\"effortLabel\":\"medium\"}",
      "notes": "Reported in embedded Vega-Lite chart in GPT-6 Sol and Luna launch post.",
      "display_name": "GPT-6 Luna",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1w8ic1j",
      "source_id": "r-1m7ewdf",
      "source_n": 26
    },
    {
      "model_key": "beam",
      "model_as_reported": "Beam",
      "effort": "",
      "harness": "",
      "score_pct": 44.4,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Reflection AI Beam launch post",
      "source_url": "https://reflection.ai/blog/introducing-beam",
      "published": "2026-10-05",
      "accessed": "2026-10-09",
      "quote": "DeepSWE v1.1 | 44.4",
      "notes": "Open-weight 501B/23B MoE under Apache 2.0 license. Reported in the agentic coding evaluation table under DeepSWE v1.1. Evaluated across reasoning effort levels during RL training.",
      "display_name": "Beam",
      "lab": "Reflection AI",
      "open_weights": true,
      "row_id": "row-16j4ryu",
      "source_id": "r-1x8ip3p",
      "source_n": 31
    },
    {
      "model_key": "gpt-5.6-luna",
      "model_as_reported": "gpt-5-6-luna",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 44.25,
      "ci_pct": 2.92,
      "n_samples": 4,
      "cost_per_task_usd": 0.1556,
      "tokens_per_task": 3398256,
      "input_tokens_per_task": 3372478,
      "output_tokens_per_task": 25778,
      "steps_per_task": 49,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-10",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_6_luna_high: pass_at_1 0.4424778761061947, ci_half 0.029195672479165317, n_runs 4, mean_cost_usd 0.15558014999999997",
      "notes": "pass@4 75.2%; mean 8 min per attempt; leaderboard-live.json gives $0.78 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.6 Luna",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1kerb9g",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "glm-5.2",
      "model_as_reported": "glm-5-2",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 43.78,
      "ci_pct": 1.73,
      "n_samples": 4,
      "cost_per_task_usd": 3.9199,
      "tokens_per_task": 12684736,
      "input_tokens_per_task": 12606560,
      "output_tokens_per_task": 78175,
      "steps_per_task": 129.1,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-21",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_glm_5_2_max: pass_at_1 0.43777777777777777, ci_half 0.01725629620580811, n_runs 4, mean_cost_usd 3.9199075743111114",
      "notes": "pass@4 77%; mean 44 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GLM-5.2",
      "lab": "Z.ai",
      "open_weights": true,
      "row_id": "row-grzg6t",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "glm-5.3-flash",
      "model_as_reported": "GLM 5.3 Flash",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 42.5,
      "ci_pct": 6.5,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "GLM 5.3 Flash | effort: max | score: 42.5% ± 6.5% | n_samples: 339 | provider: Zhipu",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "GLM-5.3 Flash",
      "lab": "Z.ai",
      "open_weights": true,
      "row_id": "row-j8z20a",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gpt-5.6-luna",
      "model_as_reported": "GPT-5.6 Luna",
      "effort": "high",
      "harness": "lab-internal",
      "score_pct": 42.37,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 0.1272,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6 Sol and Luna launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-sol-and-luna/",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-5.6 Luna\",\"cost_label\":\"$$0.1272\",\"score\":0.4237,\"cost\":0.1272,\"effortLabel\":\"high\"}",
      "notes": "OpenAI re-ran GPT-5.6 Luna internally for its GPT-6 Sol/Luna post; the official board measured 67.2% at max. Full effort breakdown from low to max reported in GPT-6 Sol and Luna launch post chart data.",
      "display_name": "GPT-5.6 Luna",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1fvmcu6",
      "source_id": "r-1m7ewdf",
      "source_n": 26
    },
    {
      "model_key": "qwen3.8-27b",
      "model_as_reported": "Qwen3.8-27B",
      "effort": "",
      "harness": "Claude Code",
      "score_pct": 42.2,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Qwen3.8-27B HuggingFace Model Card",
      "source_url": "https://huggingface.co/Qwen/Qwen3.8-27B",
      "published": "2026-08-05",
      "accessed": "2026-10-09",
      "quote": "<td class=\"benchmark-cell\" ...><div class=\"benchmark-capability\" ...>Agentic coding</div><div class=\"benchmark-name\" ...>DeepSWE 1.1</div></td>\n<td ...><strong>42.2</strong></td>",
      "notes": "Footnote 3: DeepSWE 1.1: Evaluated with the Claude Code harness at temp=1.0, top_p=0.95, and a 256K context window.",
      "display_name": "Qwen3.8-27B",
      "lab": "Alibaba",
      "open_weights": true,
      "row_id": "row-gz1a0a",
      "source_id": "r-190dm7p",
      "source_n": 10
    },
    {
      "model_key": "grok-4.6",
      "model_as_reported": "grok-4-6",
      "effort": "low",
      "harness": "mini-swe-agent",
      "score_pct": 41.65,
      "ci_pct": 2.32,
      "n_samples": 4,
      "cost_per_task_usd": 1.0424,
      "tokens_per_task": 1663378,
      "input_tokens_per_task": 1646920,
      "output_tokens_per_task": 16458,
      "steps_per_task": 44.2,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-08-12",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_grok_4_6_low: pass_at_1 0.41648106904231624, ci_half 0.023186517644299014, n_runs 4, mean_cost_usd 1.0423604231625836",
      "notes": "pass@4 69%; mean 5 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Grok 4.6",
      "lab": "xAI",
      "open_weights": false,
      "row_id": "row-1ucz10t",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "claude-opus-4.8",
      "model_as_reported": "claude-opus-4-8",
      "effort": "low",
      "harness": "mini-swe-agent",
      "score_pct": 40.8,
      "ci_pct": 1.46,
      "n_samples": 4,
      "cost_per_task_usd": 2.2934,
      "tokens_per_task": 2446911,
      "input_tokens_per_task": 2417988,
      "output_tokens_per_task": 28923,
      "steps_per_task": 54,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_opus_4_8_low: pass_at_1 0.4079822616407982, ci_half 0.014635847585691517, n_runs 4, mean_cost_usd 2.293369216740577",
      "notes": "pass@4 68.1%; mean 11 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Opus 4.8",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1pcx9r4",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "laguna-s-2.1",
      "model_as_reported": "Laguna S 2.1",
      "effort": "max",
      "harness": "pool",
      "score_pct": 40.4,
      "ci_pct": null,
      "n_samples": 3,
      "cost_per_task_usd": null,
      "tokens_per_task": 249000,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Poolside Laguna S 2.1 launch post",
      "source_url": "https://poolside.ai/blog/introducing-laguna-s-2-1",
      "published": "2026-07-21",
      "accessed": "2026-10-09",
      "quote": "On DeepSWE v1.1, Laguna S 2.1 scores 40.4 in thinking mode in pool harness.",
      "notes": "Evaluated using Poolside's proprietary 'pool' agent harness (not mini-swe-agent as aggregators cited) with 500-step limit and 6-hour timeout. pass@1 averaged over 3 attempts per task. Thinking mode lifts score from 16.5% (99k tokens, thinking off) to 40.4% (249k tokens, thinking max).",
      "display_name": "Laguna S 2.1",
      "lab": "Poolside",
      "open_weights": true,
      "row_id": "row-befdi5",
      "cost_estimate": {
        "low": 0.0224,
        "high": 0.0448,
        "basis": "249000 total tokens with no input/output split, priced all as input (low) to all as output (high) at poolside/laguna-s-2.1 list prices (AI Gateway, 2026-10-09); cached input would cost less"
      },
      "source_id": "r-12e8ru1",
      "source_n": 7
    },
    {
      "model_key": "muse-spark-1.1",
      "model_as_reported": "Muse Spark 1.1",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 40.4,
      "ci_pct": 6.5,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Muse Spark 1.1 | effort: xhigh | score: 40.4% ± 6.5% | n_samples: 339 | provider: Meta",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Muse Spark 1.1",
      "lab": "Meta",
      "open_weights": false,
      "row_id": "row-1bkgt8e",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "claude-opus-4.7",
      "model_as_reported": "claude-opus-4-7",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 40.27,
      "ci_pct": 3.5,
      "n_samples": null,
      "cost_per_task_usd": 4.83,
      "tokens_per_task": 6214791,
      "input_tokens_per_task": 6169055,
      "output_tokens_per_task": 45736,
      "steps_per_task": 87.2,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "claude-opus-4-7 [high]: 40.27% ±3.5%, cost $4.83, output tokens 45736, steps 87.2",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "Claude Opus 4.7",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-r2zw6y",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "claude-sonnet-5",
      "model_as_reported": "claude-sonnet-5",
      "effort": "medium",
      "harness": "mini-swe-agent",
      "score_pct": 39.78,
      "ci_pct": 3.13,
      "n_samples": 4,
      "cost_per_task_usd": 4.079,
      "tokens_per_task": 9343779,
      "input_tokens_per_task": 9286962,
      "output_tokens_per_task": 56817,
      "steps_per_task": 107.6,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-02",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_sonnet_5_medium: pass_at_1 0.3977777777777778, ci_half 0.03127576206329727, n_runs 4, mean_cost_usd 4.079036229",
      "notes": "pass@4 64.6%; mean 19 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Sonnet 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1iia63p",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "glm-5.2",
      "model_as_reported": "GLM 5.2",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 38.9,
      "ci_pct": 6.5,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "GLM 5.2 | effort: max | score: 38.9% ± 6.5% | n_samples: 339 | provider: Zhipu",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "GLM-5.2",
      "lab": "Z.ai",
      "open_weights": true,
      "row_id": "row-1k7amnx",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "swe-1.7",
      "model_as_reported": "SWE-1.7",
      "effort": "",
      "harness": "Devin CLI",
      "score_pct": 37.7,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Cognition SWE-2 launch post",
      "source_url": "https://cognition.com/blog/swe-2",
      "published": "2026-09-10",
      "accessed": "2026-10-09",
      "quote": "SWE-1.7 | 37.7%",
      "notes": "Reported retrospectively in the SWE-2 launch post DeepSWE 1.1 comparison table. The original SWE-1.7 launch post on 2026-07-08 (https://cognition.com/blog/swe-1-7) predated and did not report DeepSWE 1.1.",
      "display_name": "SWE-1.7",
      "lab": "Cognition",
      "open_weights": false,
      "row_id": "row-1bfjpb6",
      "source_id": "r-prwi23",
      "source_n": 19
    },
    {
      "model_key": "gpt-6-sol",
      "model_as_reported": "GPT-6 Sol",
      "effort": "low",
      "harness": "lab-internal",
      "score_pct": 37.17,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 0.1623,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6 Sol and Luna launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-sol-and-luna/",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-6 Sol\",\"cost_label\":\"$$0.16\",\"score\":0.3717,\"cost\":0.1623,\"effortLabel\":\"low\"}",
      "notes": "Reported in embedded Vega-Lite chart and text: 'On DeepSWE v1.1... GPT-6 Sol at max effort scores 68.8%, within 1.1 percentage points of Claude Fable 5’s highest score in the evaluation—69.9% at xhigh effort—at approximately 80% lower cost per task.'",
      "display_name": "GPT-6 Sol",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1bw434n",
      "source_id": "r-1m7ewdf",
      "source_n": 26
    },
    {
      "model_key": "gemini-3.5-flash",
      "model_as_reported": "Gemini 3.5 Flash",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 36.6,
      "ci_pct": 6.9,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Gemini 3.5 Flash | effort: high | score: 36.6% ± 6.9% | n_samples: 339 | provider: Google",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Gemini 3.5 Flash",
      "lab": "Google",
      "open_weights": false,
      "row_id": "row-1twb52n",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "glm-5.2",
      "model_as_reported": "glm-5-2",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 36.28,
      "ci_pct": 4.75,
      "n_samples": 4,
      "cost_per_task_usd": 2.8355,
      "tokens_per_task": 9124851,
      "input_tokens_per_task": 9070606,
      "output_tokens_per_task": 54246,
      "steps_per_task": 121.9,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-21",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_glm_5_2_high: pass_at_1 0.36283185840707965, ci_half 0.047500729479216554, n_runs 4, mean_cost_usd 2.8355164725663715",
      "notes": "pass@4 68.1%; mean 30 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GLM-5.2",
      "lab": "Z.ai",
      "open_weights": true,
      "row_id": "row-a6zyiq",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "nex-n2.5-mini",
      "model_as_reported": "Nex-N2.5-mini",
      "effort": "",
      "harness": "NexAU",
      "score_pct": 36.1,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Nex-AGI Nex-N2.5 GitHub README / Model Card",
      "source_url": "https://github.com/nex-agi/Nex-N2.5",
      "published": "2026-09-08",
      "accessed": "2026-10-09",
      "quote": "DeepSWE v1.1 | 36.1 | 55.8 | 65.6 | 73.7 | 72.7 | 67.5 | 66.9 | 62.8 | 69.3",
      "notes": "Coding tasks evaluated using the NexAU harness. Sampling parameters: temperature = 0.7, top_p = 0.95, top_k = 40. Lightweight multimodal agentic model.",
      "display_name": "Nex-N2.5-Mini",
      "lab": "Nex-AGI",
      "open_weights": true,
      "row_id": "row-92bvuc",
      "source_id": "r-pixn6e",
      "source_n": 18
    },
    {
      "model_key": "gemini-3.5-flash",
      "model_as_reported": "gemini-3-5-flash",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 36.06,
      "ci_pct": 3.97,
      "n_samples": 4,
      "cost_per_task_usd": 3.4467,
      "tokens_per_task": 7704802,
      "input_tokens_per_task": 7629072,
      "output_tokens_per_task": 75730,
      "steps_per_task": 105.3,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gemini_3_5_flash_high: pass_at_1 0.3606194690265487, ci_half 0.03966303010520439, n_runs 4, mean_cost_usd 3.4467114588691796",
      "notes": "pass@4 63.7%; mean 19 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Gemini 3.5 Flash",
      "lab": "Google",
      "open_weights": false,
      "row_id": "row-vnphv3",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-5.6-terra",
      "model_as_reported": "gpt-5-6-terra",
      "effort": "medium",
      "harness": "mini-swe-agent",
      "score_pct": 35.11,
      "ci_pct": 3.38,
      "n_samples": 4,
      "cost_per_task_usd": 0.4666,
      "tokens_per_task": 736847,
      "input_tokens_per_task": 725101,
      "output_tokens_per_task": 11747,
      "steps_per_task": 25.1,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-10",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_6_terra_medium: pass_at_1 0.3511111111111111, ci_half 0.03381964303406632, n_runs 4, mean_cost_usd 0.46659949511111115",
      "notes": "pass@4 60.2%; mean 4 min per attempt; leaderboard-live.json gives $0.58 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.6 Terra",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-poihd0",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "claude-sonnet-4.6",
      "model_as_reported": "Sonnet 4.6",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 34.2,
      "ci_pct": 6.3,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Sonnet 4.6 | effort: high | score: 34.2% ± 6.3% | n_samples: 339 | provider: Anthropic",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Claude Sonnet 4.6",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-kcyzlp",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "grok-4.7",
      "model_as_reported": "Grok 4.7",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 33.3,
      "ci_pct": 6,
      "n_samples": 337,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Grok 4.7 | effort: xhigh | score: 33.3% ± 6% | n_samples: 337 | provider: xAI",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Grok 4.7",
      "lab": "xAI",
      "open_weights": false,
      "row_id": "row-bfbz9v",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "claude-opus-4.6",
      "model_as_reported": "Opus 4.6",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 31.9,
      "ci_pct": 6.3,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Opus 4.6 | effort: max | score: 31.9% ± 6.3% | n_samples: 339 | provider: Anthropic",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Claude Opus 4.6",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-vnot5p",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "claude-opus-4.7",
      "model_as_reported": "claude-opus-4-7",
      "effort": "medium",
      "harness": "mini-swe-agent",
      "score_pct": 31.64,
      "ci_pct": 4.03,
      "n_samples": null,
      "cost_per_task_usd": 2.34,
      "tokens_per_task": 2572887,
      "input_tokens_per_task": 2545442,
      "output_tokens_per_task": 27445,
      "steps_per_task": 55.4,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "claude-opus-4-7 [medium]: 31.64% ±4.03%, cost $2.34, output tokens 27445, steps 55.4",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "Claude Opus 4.7",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-3ph296",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "kimi-k2.7-code",
      "model_as_reported": "kimi-k2-7-code",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 30.53,
      "ci_pct": 0.5,
      "n_samples": 4,
      "cost_per_task_usd": 2.8155,
      "tokens_per_task": 13067339,
      "input_tokens_per_task": 13008041,
      "output_tokens_per_task": 59297,
      "steps_per_task": 149.1,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_kimi_k2_7_code_default: pass_at_1 0.3053097345132743, ci_half 0.005007016522460913, n_runs 4, mean_cost_usd 2.815536243274336",
      "notes": "pass@4 61.1%; mean 36 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Kimi K2.7 Code",
      "lab": "Moonshot AI",
      "open_weights": null,
      "row_id": "row-1yoj145",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "claude-sonnet-5",
      "model_as_reported": "claude-sonnet-5",
      "effort": "low",
      "harness": "mini-swe-agent",
      "score_pct": 30.51,
      "ci_pct": 1.13,
      "n_samples": 4,
      "cost_per_task_usd": 2.1866,
      "tokens_per_task": 4570071,
      "input_tokens_per_task": 4534476,
      "output_tokens_per_task": 35595,
      "steps_per_task": 76.9,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-02",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_sonnet_5_low: pass_at_1 0.3051224944320713, ci_half 0.011260175414427497, n_runs 4, mean_cost_usd 2.1865550832962137",
      "notes": "pass@4 57.5%; mean 12 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Sonnet 5",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1em597q",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "kimi-k2.7-code",
      "model_as_reported": "Kimi K2.7 Code",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 30.1,
      "ci_pct": 8.4,
      "n_samples": 113,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Kimi K2.7 Code | effort: high | score: 30.1% ± 8.4% | n_samples: 113 | provider: Kimi",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Kimi K2.7 Code",
      "lab": "Moonshot AI",
      "open_weights": null,
      "row_id": "row-12ms1es",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "claude-sonnet-4.6",
      "model_as_reported": "claude-sonnet-4-6",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 29.93,
      "ci_pct": 4.09,
      "n_samples": 4,
      "cost_per_task_usd": 5.5224,
      "tokens_per_task": 12946788,
      "input_tokens_per_task": 12870628,
      "output_tokens_per_task": 76160,
      "steps_per_task": 133.7,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_claude_sonnet_4_6_high: pass_at_1 0.29933481152993346, ci_half 0.04092141269189136, n_runs 4, mean_cost_usd 5.5223517987804875",
      "notes": "pass@4 56.6%; mean 47 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Claude Sonnet 4.6",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-uei8w1",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gemini-3.5-flash",
      "model_as_reported": "gemini-3-5-flash",
      "effort": "medium",
      "harness": "mini-swe-agent",
      "score_pct": 28.32,
      "ci_pct": 3.54,
      "n_samples": null,
      "cost_per_task_usd": 7.42,
      "tokens_per_task": 12736169,
      "input_tokens_per_task": 12547145,
      "output_tokens_per_task": 189024,
      "steps_per_task": 75.1,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "gemini-3-5-flash [medium]: 28.32% ±3.54%, cost $7.42, output tokens 189024, steps 75.1",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "Gemini 3.5 Flash",
      "lab": "Google",
      "open_weights": false,
      "row_id": "row-1bibl2y",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "hy3",
      "model_as_reported": "Hy3",
      "effort": "",
      "harness": "lab-internal",
      "score_pct": 28,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Tencent Hy4-preview Technical Report & Model Card",
      "source_url": "https://huggingface.co/tencent/Hy4-preview",
      "published": "2026-08-27",
      "accessed": "2026-10-09",
      "quote": "Hy3 recorded 28.0% on DeepSWE, reported as previous generation baseline in the Hy4 Benchmark Appendix.",
      "notes": "295B MoE baseline model released mid-2026.",
      "display_name": "Hy3",
      "lab": "Tencent",
      "open_weights": true,
      "row_id": "row-1ilk7um",
      "source_id": "r-p3aj15",
      "source_n": 15
    },
    {
      "model_key": "claude-opus-4.6",
      "model_as_reported": "claude-opus-4-6",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 27.6,
      "ci_pct": 3.68,
      "n_samples": null,
      "cost_per_task_usd": 5.39,
      "tokens_per_task": 7176852,
      "input_tokens_per_task": 7132371,
      "output_tokens_per_task": 44481,
      "steps_per_task": 103,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "claude-opus-4-6 [max]: 27.6% ±3.68%, cost $5.39, output tokens 44481, steps 103.0",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "Claude Opus 4.6",
      "lab": "Anthropic",
      "open_weights": false,
      "row_id": "row-1thnwnf",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "gpt-5.5",
      "model_as_reported": "gpt-5-5",
      "effort": "low",
      "harness": "mini-swe-agent",
      "score_pct": 26.99,
      "ci_pct": 2.29,
      "n_samples": 4,
      "cost_per_task_usd": 1.2002,
      "tokens_per_task": 786194,
      "input_tokens_per_task": 776752,
      "output_tokens_per_task": 9443,
      "steps_per_task": 28.1,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_5_low: pass_at_1 0.26991150442477874, ci_half 0.02294503222007181, n_runs 4, mean_cost_usd 1.2002025442477877",
      "notes": "pass@4 47.8%; mean 9 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.5",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-165hhz9",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "grok-4.6",
      "model_as_reported": "Grok 4.6",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 24.8,
      "ci_pct": 6.8,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Grok 4.6 | effort: xhigh | score: 24.8% ± 6.8% | n_samples: 339 | provider: xAI",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Grok 4.6",
      "lab": "xAI",
      "open_weights": false,
      "row_id": "row-f6bkeh",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gpt-5.4-mini",
      "model_as_reported": "gpt-5-4-mini",
      "effort": "xhigh",
      "harness": "mini-swe-agent",
      "score_pct": 24.34,
      "ci_pct": 2.96,
      "n_samples": null,
      "cost_per_task_usd": 2.08,
      "tokens_per_task": 12313472,
      "input_tokens_per_task": 12178956,
      "output_tokens_per_task": 134516,
      "steps_per_task": 85.9,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "gpt-5-4-mini [xhigh]: 24.34% ±2.96%, cost $2.08, output tokens 134516, steps 85.9",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "GPT-5.4 mini",
      "lab": "OpenAI",
      "open_weights": null,
      "row_id": "row-3oes9n",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "gpt-5.6-terra",
      "model_as_reported": "gpt-5-6-terra",
      "effort": "low",
      "harness": "mini-swe-agent",
      "score_pct": 24.05,
      "ci_pct": 0.78,
      "n_samples": 4,
      "cost_per_task_usd": 0.3422,
      "tokens_per_task": 489138,
      "input_tokens_per_task": 480565,
      "output_tokens_per_task": 8572,
      "steps_per_task": 21.5,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-10",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_6_terra_low: pass_at_1 0.24053452115812918, ci_half 0.007767613669447833, n_runs 4, mean_cost_usd 0.3421962084632517",
      "notes": "pass@4 44.2%; mean 3 min per attempt; leaderboard-live.json gives $0.43 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.6 Terra",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-1wrol78",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "kimi-k2.6",
      "model_as_reported": "kimi-k2-6",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 23.89,
      "ci_pct": 2.45,
      "n_samples": null,
      "cost_per_task_usd": 3.16,
      "tokens_per_task": 12102875,
      "input_tokens_per_task": 12018459,
      "output_tokens_per_task": 84416,
      "steps_per_task": 146.8,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "kimi-k2-6 [None]: 23.89% ±2.45%, cost $3.16, output tokens 84416, steps 146.8",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "Kimi K2.6",
      "lab": "Moonshot AI",
      "open_weights": null,
      "row_id": "row-1r0o7od",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "qwen3.7-max",
      "model_as_reported": "Qwen3.7-Max",
      "effort": "",
      "harness": "Claude Code",
      "score_pct": 21.6,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Qwen3.8-2.4T-A95B HuggingFace Model Card",
      "source_url": "https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B",
      "published": "2026-08-08",
      "accessed": "2026-10-09",
      "quote": "DeepSWE 1.1 | 59.0 | 70.0 | 73.0 | 21.6 | 56.6",
      "notes": "Reported as previous generation baseline in Qwen3.8-Max benchmark table. Evaluated with temp=1.0, top_p=0.95, 256K context window.",
      "display_name": "Qwen3.7 Max",
      "lab": "Alibaba",
      "open_weights": false,
      "row_id": "row-cpxpc0",
      "source_id": "r-ouk7wk",
      "source_n": 11
    },
    {
      "model_key": "minimax-m3",
      "model_as_reported": "minimax-m3",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 20.44,
      "ci_pct": 3.73,
      "n_samples": null,
      "cost_per_task_usd": 5.57,
      "tokens_per_task": 41744490,
      "input_tokens_per_task": 41646948,
      "output_tokens_per_task": 97542,
      "steps_per_task": 314.2,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "minimax-m3 [None]: 20.44% ±3.73%, cost $5.57, output tokens 97542, steps 314.2",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "MiniMax M3",
      "lab": "MiniMax",
      "open_weights": null,
      "row_id": "row-9zllbx",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "mimo-v2.5-pro",
      "model_as_reported": "mimo-v2-5-pro",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 19.47,
      "ci_pct": 1.87,
      "n_samples": null,
      "cost_per_task_usd": 1.99,
      "tokens_per_task": 8685488,
      "input_tokens_per_task": 8636167,
      "output_tokens_per_task": 49321,
      "steps_per_task": 121.5,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "mimo-v2-5-pro [None]: 19.47% ±1.87%, cost $1.99, output tokens 49321, steps 121.5",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "MiMo-V2.5-Pro",
      "lab": "Xiaomi",
      "open_weights": true,
      "row_id": "row-1q4sncw",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "glm-5.1",
      "model_as_reported": "GLM 5.1",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 18.9,
      "ci_pct": 5.5,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "GLM 5.1 | effort: None | score: 18.9% ± 5.5% | n_samples: 339 | provider: Zhipu",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "GLM-5.1",
      "lab": "Z.ai",
      "open_weights": true,
      "row_id": "row-65jvtm",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "qwen3.7-max",
      "model_as_reported": "qwen3-7-max",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 17.7,
      "ci_pct": 1.42,
      "n_samples": null,
      "cost_per_task_usd": 2.12,
      "tokens_per_task": 6486282,
      "input_tokens_per_task": 6443835,
      "output_tokens_per_task": 42447,
      "steps_per_task": 110.3,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "qwen3-7-max [None]: 17.7% ±1.42%, cost $2.12, output tokens 42447, steps 110.3",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "Qwen3.7 Max",
      "lab": "Alibaba",
      "open_weights": false,
      "row_id": "row-1rl4w7q",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "glm-5.1",
      "model_as_reported": "glm-5-1",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 17.52,
      "ci_pct": 0.79,
      "n_samples": null,
      "cost_per_task_usd": 7.46,
      "tokens_per_task": 13589155,
      "input_tokens_per_task": 13539998,
      "output_tokens_per_task": 49157,
      "steps_per_task": 176.6,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "glm-5-1 [None]: 17.52% ±0.79%, cost $7.46, output tokens 49157, steps 176.6",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "GLM-5.1",
      "lab": "Z.ai",
      "open_weights": true,
      "row_id": "row-q9lc3f",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "minimax-m3",
      "model_as_reported": "MiniMax-M3 (Extended Window)",
      "effort": "thinking",
      "harness": "mini-swe-agent",
      "score_pct": 16.8,
      "ci_pct": null,
      "n_samples": 113,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Independent DeepSWE Audit by entrpi",
      "source_url": "https://entrpi.github.io/misc/deep-swe-minimax-m3/",
      "published": "2026-06-02",
      "accessed": "2026-10-09",
      "quote": "Extended pass rate: 16.8% (19/113, counting 4 solutions finished past the standard time window).",
      "notes": "19/113 tasks solved when counting 4 solutions completed past the 90-minute deadline but within the 135-minute ceiling.",
      "display_name": "MiniMax M3",
      "lab": "MiniMax",
      "open_weights": null,
      "row_id": "row-1hc37at",
      "source_id": "r-hzsdm7",
      "source_n": 34
    },
    {
      "model_key": "qwen3.7-plus",
      "model_as_reported": "Qwen3.7-Plus",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 16.5,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Qwen3.8-Flash-Next HuggingFace Model Card",
      "source_url": "https://huggingface.co/Qwen/Qwen3.8-Flash-Next",
      "published": "2026-08-24",
      "accessed": "2026-10-09",
      "quote": "<td class=\"benchmark-name\" ...>DeepSWE 1.1</div></td>\n<td ...><strong>58.7</strong></td>\n<td ...>42.2</td>\n<td ...>16.5</td>",
      "notes": "Reported as baseline in Qwen3.8-Flash-Next README (16.5%) and Qwen3.8-27B README (14.2% on Claude Code). Evaluated with temp=1.0, top_p=0.95, 256K context window.",
      "display_name": "Qwen3.7 Plus",
      "lab": "Alibaba",
      "open_weights": null,
      "row_id": "row-1exbwot",
      "source_id": "r-qxom7y",
      "source_n": 13
    },
    {
      "model_key": "minimax-m3",
      "model_as_reported": "MiniMax-M3",
      "effort": "thinking",
      "harness": "mini-swe-agent",
      "score_pct": 13.3,
      "ci_pct": null,
      "n_samples": 113,
      "cost_per_task_usd": 7.48,
      "tokens_per_task": 80000,
      "input_tokens_per_task": null,
      "output_tokens_per_task": 80000,
      "steps_per_task": 325,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Independent DeepSWE Audit by entrpi",
      "source_url": "https://entrpi.github.io/misc/deep-swe-minimax-m3/",
      "published": "2026-06-02",
      "accessed": "2026-10-09",
      "quote": "Strict pass rate: 13.3% (15/113 evaluated problems resolved within budget); Extended pass rate: 16.8% (19/113); Median steps: 325; Median output volume: 80k tokens; Median price: $7.48 per scenario.",
      "notes": "Third-party audit using unmodified mini-swe-agent in Pier with Harbor task format. MiniMax Anthropic-compatible API, pass@1, 90 min working time limit per task, air-gapped container.",
      "display_name": "MiniMax M3",
      "lab": "MiniMax",
      "open_weights": null,
      "row_id": "row-x0507n",
      "source_id": "r-hzsdm7",
      "source_n": 34
    },
    {
      "model_key": "qwen3.6-27b",
      "model_as_reported": "Qwen3.6-27B",
      "effort": "",
      "harness": "Claude Code",
      "score_pct": 13.3,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "Qwen3.8-27B HuggingFace Model Card",
      "source_url": "https://huggingface.co/Qwen/Qwen3.8-27B",
      "published": "2026-08-05",
      "accessed": "2026-10-09",
      "quote": "| DeepSWE 1.1 | 42.2 | 13.3 | 14.2 | -- | -- |",
      "notes": "Reported as baseline in Qwen3.8-27B README benchmark table. Claude Code harness, temp=1.0, top_p=0.95, 256K context window.",
      "display_name": "Qwen3.6-27B",
      "lab": "Alibaba",
      "open_weights": true,
      "row_id": "row-1fk1m2l",
      "source_id": "r-190dm7p",
      "source_n": 10
    },
    {
      "model_key": "grok-build-0.1",
      "model_as_reported": "grok-build-0-1",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 13.08,
      "ci_pct": 2.44,
      "n_samples": null,
      "cost_per_task_usd": 6.6,
      "tokens_per_task": 10309658,
      "input_tokens_per_task": 10257762,
      "output_tokens_per_task": 51896,
      "steps_per_task": 176.1,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "grok-build-0-1 [None]: 13.08% ±2.44%, cost $6.6, output tokens 51896, steps 176.1",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "Grok Build 0.1",
      "lab": "xAI",
      "open_weights": null,
      "row_id": "row-mbkfkb",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "gemini-3.1-pro-preview",
      "model_as_reported": "gemini-3-1-pro-preview",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 11.73,
      "ci_pct": 1.48,
      "n_samples": 4,
      "cost_per_task_usd": 2.1434,
      "tokens_per_task": 2432046,
      "input_tokens_per_task": 2403677,
      "output_tokens_per_task": 28369,
      "steps_per_task": 75.6,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gemini_3_1_pro_preview_high: pass_at_1 0.1172566371681416, ci_half 0.01481095461108845, n_runs 4, mean_cost_usd 2.1434045568888886",
      "notes": "pass@4 28.3%; mean 14 min per attempt; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "Gemini 3.1 Pro (preview)",
      "lab": "Google",
      "open_weights": false,
      "row_id": "row-18zp7di",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-5.6-luna",
      "model_as_reported": "gpt-5-6-luna",
      "effort": "medium",
      "harness": "mini-swe-agent",
      "score_pct": 11.28,
      "ci_pct": 0.83,
      "n_samples": 4,
      "cost_per_task_usd": 0.0433,
      "tokens_per_task": 626007,
      "input_tokens_per_task": 617827,
      "output_tokens_per_task": 8180,
      "steps_per_task": 23.7,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-10",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_6_luna_medium: pass_at_1 0.11283185840707964, ci_half 0.008303197562056514, n_runs 4, mean_cost_usd 0.04326195796460178",
      "notes": "pass@4 27.4%; mean 3 min per attempt; leaderboard-live.json gives $0.22 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.6 Luna",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-7ygxig",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gemini-3.1-pro-preview",
      "model_as_reported": "gemini-3-1-pro-preview",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 9.73,
      "ci_pct": 2.83,
      "n_samples": null,
      "cost_per_task_usd": 1.84,
      "tokens_per_task": 2100631,
      "input_tokens_per_task": 2047640,
      "output_tokens_per_task": 52991,
      "steps_per_task": 74.1,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "gemini-3-1-pro-preview [None]: 9.73% ±2.83%, cost $1.84, output tokens 52991, steps 74.1",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "Gemini 3.1 Pro (preview)",
      "lab": "Google",
      "open_weights": false,
      "row_id": "row-xdw8fl",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "gpt-5.6-luna",
      "model_as_reported": "GPT-5.6 Luna",
      "effort": "medium",
      "harness": "lab-internal",
      "score_pct": 9.29,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 0.0313,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6 Sol and Luna launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-sol-and-luna/",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-5.6 Luna\",\"cost_label\":\"$$0.0313\",\"score\":0.0929,\"cost\":0.0313,\"effortLabel\":\"medium\"}",
      "notes": "OpenAI re-ran GPT-5.6 Luna internally for its GPT-6 Sol/Luna post; the official board measured 67.2% at max. Full effort breakdown from low to max reported in GPT-6 Sol and Luna launch post chart data.",
      "display_name": "GPT-5.6 Luna",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-yvja1r",
      "source_id": "r-1m7ewdf",
      "source_n": 26
    },
    {
      "model_key": "deepseek-v4-pro",
      "model_as_reported": "DeepSeek V4 Pro",
      "effort": "max",
      "harness": "mini-swe-agent",
      "score_pct": 9.1,
      "ci_pct": 3.8,
      "n_samples": 337,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "DeepSeek V4 Pro | effort: max | score: 9.1% ± 3.8% | n_samples: 337 | provider: DeepSeek",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "DeepSeek V4 Pro",
      "lab": "DeepSeek",
      "open_weights": true,
      "row_id": "row-1gab77k",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "minimax-m3",
      "model_as_reported": "MiniMax M3",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 9.1,
      "ci_pct": 3.1,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "MiniMax M3 | effort: high | score: 9.1% ± 3.1% | n_samples: 339 | provider: MiniMax",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "MiniMax M3",
      "lab": "MiniMax",
      "open_weights": null,
      "row_id": "row-1ggd9h",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gemini-3.1-pro",
      "model_as_reported": "Gemini 3.1 Pro",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 7.7,
      "ci_pct": 3.1,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Gemini 3.1 Pro | effort: high | score: 7.7% ± 3.1% | n_samples: 339 | provider: Google",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Gemini 3.1 Pro",
      "lab": "Google",
      "open_weights": false,
      "row_id": "row-se9nz3",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "deepseek-v4-pro",
      "model_as_reported": "deepseek-v4-pro",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 7.52,
      "ci_pct": 2.7,
      "n_samples": null,
      "cost_per_task_usd": 4.22,
      "tokens_per_task": 8403574,
      "input_tokens_per_task": 8353625,
      "output_tokens_per_task": 49949,
      "steps_per_task": 111.3,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "deepseek-v4-pro [None]: 7.52% ±2.7%, cost $4.22, output tokens 49949, steps 111.3",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "DeepSeek V4 Pro",
      "lab": "DeepSeek",
      "open_weights": true,
      "row_id": "row-183svb6",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "gemini-3-flash-preview",
      "model_as_reported": "gemini-3-flash-preview",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 5.14,
      "ci_pct": 2.48,
      "n_samples": null,
      "cost_per_task_usd": 1.53,
      "tokens_per_task": 5989122,
      "input_tokens_per_task": 5756060,
      "output_tokens_per_task": 233062,
      "steps_per_task": 70.8,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "gemini-3-flash-preview [None]: 5.14% ±2.48%, cost $1.53, output tokens 233062, steps 70.8",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "Gemini 3 Flash (preview)",
      "lab": "Google",
      "open_weights": null,
      "row_id": "row-14jryq4",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "inkling",
      "model_as_reported": "Inkling",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 5,
      "ci_pct": 2.8,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Inkling | effort: high | score: 5% ± 2.8% | n_samples: 339 | provider: Thinking Machines",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Inkling",
      "lab": "Thinking Machines",
      "open_weights": false,
      "row_id": "row-1rprkfy",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "qwen3.6-plus",
      "model_as_reported": "qwen3-6-plus",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 2.65,
      "ci_pct": 1.23,
      "n_samples": null,
      "cost_per_task_usd": 4.25,
      "tokens_per_task": 12729719,
      "input_tokens_per_task": 12662826,
      "output_tokens_per_task": 66893,
      "steps_per_task": 164.4,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "qwen3-6-plus [None]: 2.65% ±1.23%, cost $4.25, output tokens 66893, steps 164.4",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "Qwen3.6 Plus",
      "lab": "Alibaba",
      "open_weights": null,
      "row_id": "row-1yu8727",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "gpt-6-luna",
      "model_as_reported": "GPT-6 Luna",
      "effort": "low",
      "harness": "lab-internal",
      "score_pct": 2.43,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 0.0057,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6 Sol and Luna launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-sol-and-luna/",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-6 Luna\",\"cost_label\":\"$$0.0057\",\"score\":0.0243,\"cost\":0.0057,\"effortLabel\":\"low\"}",
      "notes": "Reported in embedded Vega-Lite chart in GPT-6 Sol and Luna launch post.",
      "display_name": "GPT-6 Luna",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-6tewm6",
      "source_id": "r-1m7ewdf",
      "source_n": 26
    },
    {
      "model_key": "glm-5",
      "model_as_reported": "GLM 5",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 2.1,
      "ci_pct": 1.9,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "GLM 5 | effort: None | score: 2.1% ± 1.9% | n_samples: 339 | provider: Zhipu",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "GLM-5",
      "lab": "Z.ai",
      "open_weights": true,
      "row_id": "row-w6r0fq",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gpt-5.6-luna",
      "model_as_reported": "gpt-5-6-luna",
      "effort": "low",
      "harness": "mini-swe-agent",
      "score_pct": 1.55,
      "ci_pct": 0.83,
      "n_samples": 4,
      "cost_per_task_usd": 0.0145,
      "tokens_per_task": 153159,
      "input_tokens_per_task": 150032,
      "output_tokens_per_task": 3128,
      "steps_per_task": 12.5,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard",
      "source_url": "https://deepswe.datacurve.ai/",
      "published": "2026-07-10",
      "accessed": "2026-10-09",
      "quote": "config mini_swe_agent_gpt_5_6_luna_low: pass_at_1 0.015486725663716814, ci_half 0.008303197562056514, n_runs 4, mean_cost_usd 0.014481247876106192",
      "notes": "pass@4 4.4%; mean 1 min per attempt; leaderboard-live.json gives $0.07 per task under its own cost basis; source: rows embedded in https://deepswe.datacurve.ai/",
      "display_name": "GPT-5.6 Luna",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-7db7b6",
      "source_id": "r-16wwn5l",
      "source_n": 1
    },
    {
      "model_key": "gpt-5.6-luna",
      "model_as_reported": "GPT-5.6 Luna",
      "effort": "low",
      "harness": "lab-internal",
      "score_pct": 1.22,
      "ci_pct": null,
      "n_samples": null,
      "cost_per_task_usd": 0.0106,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "lab_self_reported",
      "verified_primary": true,
      "source_name": "OpenAI GPT-6 Sol and Luna launch post",
      "source_url": "https://openai.com/index/introducing-gpt-6-sol-and-luna/",
      "published": "2026-09-22",
      "accessed": "2026-10-09",
      "quote": "{\"model\":\"GPT-5.6 Luna\",\"cost_label\":\"$$0.0106\",\"score\":0.0122,\"cost\":0.0106,\"effortLabel\":\"low\"}",
      "notes": "OpenAI re-ran GPT-5.6 Luna internally for its GPT-6 Sol/Luna post; the official board measured 67.2% at max. Full effort breakdown from low to max reported in GPT-6 Sol and Luna launch post chart data.",
      "display_name": "GPT-5.6 Luna",
      "lab": "OpenAI",
      "open_weights": false,
      "row_id": "row-10y4p19",
      "source_id": "r-1m7ewdf",
      "source_n": 26
    },
    {
      "model_key": "deepseek-v3.2",
      "model_as_reported": "DeepSeek V3.2",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 0.6,
      "ci_pct": 0.9,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "DeepSeek V3.2 | effort: None | score: 0.6% ± 0.9% | n_samples: 339 | provider: DeepSeek",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "DeepSeek V3.2",
      "lab": "DeepSeek",
      "open_weights": true,
      "row_id": "row-1irgvx5",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "nemotron-3-ultra",
      "model_as_reported": "Nemotron 3 Ultra",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 0.6,
      "ci_pct": 0.7,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Nemotron 3 Ultra | effort: high | score: 0.6% ± 0.7% | n_samples: 339 | provider: NVIDIA",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Nemotron 3 Ultra",
      "lab": "NVIDIA",
      "open_weights": true,
      "row_id": "row-p7xtiy",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "claude-haiku-4.5",
      "model_as_reported": "claude-haiku-4-5",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 0.22,
      "ci_pct": 0.43,
      "n_samples": null,
      "cost_per_task_usd": 0.84,
      "tokens_per_task": 5332574,
      "input_tokens_per_task": 5293386,
      "output_tokens_per_task": 39188,
      "steps_per_task": 109.3,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "claude-haiku-4-5 [None]: 0.22% ±0.43%, cost $0.84, output tokens 39188, steps 109.3",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "Claude Haiku 4.5",
      "lab": "Anthropic",
      "open_weights": null,
      "row_id": "row-1lndd2c",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "minimax-m2.7",
      "model_as_reported": "minimax-m2-7",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 0.22,
      "ci_pct": 0.43,
      "n_samples": null,
      "cost_per_task_usd": 0.7,
      "tokens_per_task": 9519377,
      "input_tokens_per_task": 9459354,
      "output_tokens_per_task": 60023,
      "steps_per_task": 135.6,
      "source_type": "official_leaderboard",
      "verified_primary": true,
      "source_name": "Datacurve DeepSWE 1.1 leaderboard, launch snapshot (Wayback)",
      "source_url": "https://web.archive.org/web/20260615111157/https://deepswe.datacurve.ai/",
      "published": "2026-06-15",
      "accessed": "2026-10-09",
      "quote": "minimax-m2-7 [None]: 0.22% ±0.43%, cost $0.7, output tokens 60023, steps 135.6",
      "notes": "On the v1.1 launch board on 2026-06-15; no longer on the live board.",
      "display_name": "MiniMax M2.7",
      "lab": "MiniMax",
      "open_weights": false,
      "row_id": "row-9s98xd",
      "source_id": "r-1fnts5d",
      "source_n": 2
    },
    {
      "model_key": "gemma-4-31b",
      "model_as_reported": "Gemma 4 31B",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 0,
      "ci_pct": 0,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Gemma 4 31B | effort: None | score: 0% ± 0% | n_samples: 339 | provider: Google",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Gemma 4 31B",
      "lab": "Google",
      "open_weights": true,
      "row_id": "row-4yiqqu",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "gpt-oss-120b",
      "model_as_reported": "GPT OSS 120B",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 0,
      "ci_pct": 0,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "GPT OSS 120B | effort: high | score: 0% ± 0% | n_samples: 339 | provider: OpenAI",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "gpt-oss-120b",
      "lab": "OpenAI",
      "open_weights": true,
      "row_id": "row-1tvbm70",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "kimi-k2",
      "model_as_reported": "Kimi K2",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 0,
      "ci_pct": 0,
      "n_samples": 339,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Kimi K2 | effort: high | score: 0% ± 0% | n_samples: 339 | provider: Kimi",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Kimi K2",
      "lab": "Moonshot AI",
      "open_weights": false,
      "row_id": "row-1dj6krb",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "minimax-m2.7",
      "model_as_reported": "MiniMax M2.7",
      "effort": "high",
      "harness": "mini-swe-agent",
      "score_pct": 0,
      "ci_pct": 0,
      "n_samples": 337,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "MiniMax M2.7 | effort: high | score: 0% ± 0% | n_samples: 337 | provider: MiniMax",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "MiniMax M2.7",
      "lab": "MiniMax",
      "open_weights": false,
      "row_id": "row-1lhbuv",
      "source_id": "r-14zikky",
      "source_n": 36
    },
    {
      "model_key": "qwen3.5",
      "model_as_reported": "Qwen 3.5",
      "effort": "",
      "harness": "mini-swe-agent",
      "score_pct": 0,
      "ci_pct": 0,
      "n_samples": 113,
      "cost_per_task_usd": null,
      "tokens_per_task": null,
      "input_tokens_per_task": null,
      "output_tokens_per_task": null,
      "steps_per_task": null,
      "source_type": "third_party_run",
      "verified_primary": true,
      "source_name": "Mercor APEX (mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/)",
      "source_url": "https://www.mercor.com/apex/oss-benchmarks/oss-deep-swe-leaderboard/",
      "published": "",
      "accessed": "2026-10-09",
      "quote": "Qwen 3.5 | effort: None | score: 0% ± 0% | n_samples: 113 | provider: Alibaba",
      "notes": "Mercor shows no per-model dates; on its board as of the 2026-10-07 capture. Mercor independent evaluation using mini-swe-agent (500 max steps, 2 hour time limit), unit tests all must pass, judge of None, k=3 (113 tasks, 339 samples; 113 samples for k=1).; harness as reported: mini-swe-agent (500 max steps, 2 hour time limit)",
      "display_name": "Qwen3.5",
      "lab": "Alibaba",
      "open_weights": true,
      "row_id": "row-hdfwj0",
      "source_id": "r-14zikky",
      "source_n": 36
    }
  ]
}
