Skip to content

    AI Progress Tracker, SWE-bench Verified (2024–2026)

    Last updated 2026-06-24 · Updated weekly

    SWE-BENCH VERIFIED · 500 REAL GITHUB ISSUES · Y = % ISSUES RESOLVED
    ★ = SOTA OPEN-SOURCE AT LAUNCH · TAP/HOVER DOTS FOR DETAIL · TRACKING HOW FAST AI IS MOVING, AND WHAT IT MEANS FOR YOU

    Date:
    SWE-bench score:
    Type:

    ⏱ Time for OSS to catch each frontier model

    Frontier
    Claude Mythos 5 (95.5%)
    2026-06-09
    First OSS Match
    Not yet caught
    No OSS model has matched this score
    open gap
    Frontier
    Claude Fable 5 (95%)
    2026-06-09
    First OSS Match
    Not yet caught
    No OSS model has matched this score
    open gap
    Frontier
    GPT-5.5 (88.7%)
    2026-04-23
    First OSS Match
    Not yet caught
    No OSS model has matched this score
    open gap
    Frontier
    Claude Opus 4.7 (87.6%)
    2026-04-16
    First OSS Match
    Not yet caught
    No OSS model has matched this score
    open gap
    Frontier
    GPT-5.3 Codex (85%)
    2026-04-10
    First OSS Match
    Not yet caught
    No OSS model has matched this score
    open gap
    Frontier
    Claude Mythos Preview (93.9%)
    2026-04-07
    First OSS Match
    Not yet caught
    No OSS model has matched this score
    open gap
    Frontier
    GPT-5.2 (80%)
    2025-12-11
    First OSS Match
    MiniMax M2.5
    2026-02-12 · 80.2%
    63d
    2.1 mo
    Frontier
    GPT-5.1 (76.3%)
    2025-11-12
    First OSS Match
    Kimi K2.5
    2026-01-26 · 76.8%
    75d
    2.5 mo
    Frontier
    Claude Opus 4.5 (80.9%)
    2025-11-01
    First OSS Match
    Not yet caught
    No OSS model has matched this score
    open gap
    Frontier
    Claude Sonnet 4.5 (77.2%)
    2025-09-29
    First OSS Match
    GLM-5
    2026-02-11 · 77.8%
    135d
    4.4 mo
    Frontier
    GPT-5 (74.9%)
    2025-08-07
    First OSS Match
    Kimi K2.5
    2026-01-26 · 76.8%
    172d
    5.7 mo
    Frontier
    Claude Opus 4.1 (74.5%)
    2025-08-05
    First OSS Match
    Kimi K2.5
    2026-01-26 · 76.8%
    174d
    5.7 mo
    Frontier
    Claude 4 Sonnet (72.7%)
    2025-05-22
    First OSS Match
    DeepSeek-V3.2
    2025-12-01 · 73.1%
    193d
    6.3 mo
    Frontier
    o3 (71.7%)
    2025-04-16
    First OSS Match
    DeepSeek-V3.2
    2025-12-01 · 73.1%
    229d
    7.5 mo
    Frontier
    Claude 3.7 Sonnet (70.3%)
    2025-02-24
    First OSS Match
    DeepSeek-V3.2
    2025-12-01 · 73.1%
    280d
    9.2 mo
    Frontier
    Claude 3.5 Sonnet v2 (49%)
    2024-10-22
    First OSS Match
    DeepSeek-R1
    2025-01-20 · 49.2%
    90d
    3.0 mo
    Frontier
    o1 (48.9%)
    2024-09-12
    First OSS Match
    DeepSeek-R1
    2025-01-20 · 49.2%
    130d
    4.3 mo
    Frontier
    GPT-4o (33.2%)
    2024-05-13
    First OSS Match
    Qwen2.5-Coder-32B
    2024-11-11 · 39%
    182d
    6.0 mo
    ▾ FULL DATA TABLE, all models, dates, scores, sources
    #ModelReleasedScoreBarNotesSource
    1GLM-5.2OSS2026-06-16Pending
    ⚠ SWE-bench Verified pending — Z.AI published no benchmarks at launch (third-party SWE-bench Pro: 62.1%). Flagship 744B-total / 40B-active MoE with usable 1M-token context and dual reasoning modes.www.marktechpost.com
    2Kimi K2.7 CodeOSS2026-06-1278.2%
    ⚠ Vendor-reported: 78.2% SWE-bench Verified (Moonshot AI); no independent third-party verification at launch. Coding-focused 1T-parameter open-weight; uses 30% fewer reasoning tokens than K2.6.huggingface.co
    3DiffusionGemma 26B-A4BOSS2026-06-10Pending
    ⚠ SWE-bench Verified pending — speed-focused research model with no published score. Google DeepMind experimental open-weight diffusion-based text LLM (26B MoE, Apache 2.0); generates 4× faster than autoregressive Gemma 4.blog.google
    4Claude Fable 52026-06-0995.0%
    95.0% SWE-bench Verified — new public SOTA, #2 on BenchLM leaderboard. First publicly-available Mythos-class model (tier above Opus 4.8); 1M context.www.anthropic.com
    5Claude Mythos 52026-06-0995.5%
    95.5% SWE-bench Verified — #1 on BenchLM leaderboard. Restricted-access variant of Fable 5 (Project Glasswing tier); same weights, lifted safeguards for authorized cybersecurity orgs.www.anthropic.com
    6North Mini Code 1.0OSS2026-06-0980.2%
    ⚠ Caveat: 80.2% pass@10 SWE-bench Verified (Cohere blog) — not pass@1, so direct comparison is misleading. 30B-A3B MoE, Apache 2.0; single H100 deployable.cohere.com
    7Unisound U22026-06-0775.0%
    ⚠ Vendor-reported 75% SWE-bench Verified (Unisound press release). Chinese agentic foundation model capable of 100+ step autonomous workflows.www.prnewswire.com
    8Nemotron 3 UltraOSS2026-06-0471.9%
    71.9% SWE-bench Verified (NVIDIA technical report, BenchLM); 550B-total / 55B-active hybrid Mamba-Transformer MoE optimized for long-running agents.research.nvidia.com
    9MAI-Code-1-Flash2026-06-0271.6%
    71.6% SWE-bench Verified (Microsoft AI announcement); first in-house Microsoft coding model — 137B sparse MoE, 5B active. Launched at Build 2026; beats Claude Haiku 4.5 (66.6%).microsoft.ai
    10MAI-Thinking-12026-06-0273.5%
    73.5% SWE-bench Verified (BenchLM); first in-house Microsoft reasoning model. 35B active / 1T total MoE, trained from scratch without OpenAI distillation.microsoft.ai
    11MiniMax M3OSS2026-06-0180.5%
    ⚠ Contested: 80.5% SWE-bench Verified (vendor-reported, llm-stats listing); only 59% SWE-bench Pro suggests harness gap. Open-weight 1M-context MoE with MSA architecture.venturebeat.com
    12Grok Build 0.12026-05-2970.8%
    ⚠ Vendor-reported: 70.8% SWE-bench Verified (xAI announcement, own harness); no independent score yet. Dedicated coding model behind Grok Build CLI, 256K context.x.ai
    13Step 3.7 FlashOSS2026-05-29Pending
    ⚠ SWE-bench Verified pending — no standalone score published. StepFun 198B sparse MoE vision-language coding model, Apache 2.0; vendor claims ~97% of Claude Opus 4.6's Verified performance with Advisor Mode.www.marktechpost.com
    14Claude Opus 4.82026-05-2888.6%
    88.6% SWE-bench Verified (llm-stats.com, BenchLM); adaptive-thinking Opus refresh with 1M context window.www.anthropic.com
    15Qwen3.7-Max2026-05-2080.4%
    80.4% SWE-bench Verified (felloai.com, DataCamp, startupfortune.com); 60.6% SWE-bench Pro (highest at launch). Closed API flagship. 1M context. Launched at Alibaba Cloud Summit Hangzhou.qwen.ai
    16Gemini 3.5 Flash2026-05-1978.0%
    78% SWE-bench Verified (Google I/O 2026); beats Gemini 3.1 Pro on 11/15 agentic benchmarks; 4× faster than rival frontier models; 55.1% SWE-bench Pro. Closed model.blog.google
    17GPT-5.5 Instant2026-05-05Pending
    ⚠ SWE-bench Verified pending — no independent score published. Lighter/faster default-tier variant of GPT-5.5; became the new ChatGPT default May 5. $1.50/$6/M tokens.techcrunch.com
    18Grok 4.32026-04-30Pending
    ⚠ SWE-bench Verified pending (no official score published). Full API release April 30; improved agentic performance over Grok 4.20; 40% price cut, $0.80/$3.20/M tokens.artificialanalysis.ai
    19Mistral Medium 3.5OSS2026-04-2977.6%
    77.6% SWE-bench Verified (MarkTechPost / winbuzzer); unified 128B model merging chat+reasoning+code, replaces Devstral 2 and Magistral. Modified MIT license. 256K context, $1.50/M input.huggingface.co
    20DeepSeek-V4-ProOSS2026-04-2480.6%
    80.6% SWE-bench Verified — new OSS SOTA, 0.4pp above MiniMax M2.5/Kimi K2.6. 1.6T param MoE, 49B active, 1M context. MIT license. $0.30/$1.20/M tokens.huggingface.co
    21DeepSeek-V4-FlashOSS2026-04-2479.0%
    79.0% SWE-bench Verified (BenchLM); 284B-A13B MoE — 12× cheaper than V4-Pro at 1.6pp lower SWE-bench Verified. MIT license, 1M context.huggingface.co
    22GPT-5.52026-04-2388.7%
    88.7% SWE-bench Verified — new public SOTA, beats Claude Opus 4.7 (87.6%). First fully retrained base since GPT-4.5. 1M context, $5/$30/M tokens. Codename 'Spud'.openai.com
    23Hy3-previewOSS2026-04-2374.4%
    74.4% SWE-bench Verified (Artificial Analysis, multiple sources); 40% jump from Hy2 (53.0%). Tencent. 295B MoE / 21B active, 256K context. Tencent Hy Community License.huggingface.co
    24Qwen3.6-27BOSS2026-04-2277.2%
    77.2% SWE-bench Verified (llm-stats.com / the-decoder.com); 27B dense model, Apache 2.0. Surpasses Qwen3.5-397B-A17B on all major coding benchmarks despite 14× fewer params.huggingface.co
    25MiMo-V2.5-ProOSS2026-04-2278.9%
    78.9% SWE-bench Verified (artificialanalysis.ai, the-decoder.com); 57.2% SWE-bench Pro. Xiaomi. 1.02T MoE, 42B active, MIT license, 1M context.huggingface.co
    26Qwen3.6-Max-Preview2026-04-20Pending
    ⚠ SWE-bench Verified pending (no confirmed score). Leads SWE-bench Pro, Terminal-Bench 2.0, SkillsBench at launch (#1 on 6 benchmarks). Closed API flagship from Alibaba. 260K context.qwen.ai
    27Kimi K2.6OSS2026-04-2080.2%
    80.2% SWE-bench Verified; 58.6% SWE-bench Pro. 1T params / 32B active MoE. 300-agent swarm, 4000-step horizon, 262K context. Modified MIT.huggingface.co
    28Claude Opus 4.72026-04-1687.6%
    87.6% SWE-bench Verified; 64.3% SWE-bench Pro (#1 public model at launch). xhigh effort, 3x vision, /ultrareview, 1M context. $5/$25 per M tokens.www.anthropic.com
    29Qwen3.6-35B-A3BOSS2026-04-1673.4%
    73.4% SWE-bench Verified. Apache 2.0. 35B total / 3B active MoE — runs on a laptop. Highest OSS SWE score on consumer hardware.huggingface.co
    30GPT-5.3 Codex2026-04-1085.0%
    85.0% SWE-bench Verified — OpenAI coding-specialist Codex variant; 56.8% SWE-bench Pro (benchlm.ai)benchlm.ai
    31Muse Spark2026-04-0877.4%
    77.4% SWE-bench Verified (The Next Web, llm-stats.com, vals.ai); 55.0% SWE-bench Pro. Meta Superintelligence Labs — first flagship closed-source Meta model. Natively multimodal.ai.meta.com
    32Claude Mythos Preview2026-04-0793.9%
    93.9% SWE-bench Verified — most capable model ever built; NOT publicly released. Restricted to 50 orgs under Project Glasswing (cybersecurity). Triggered ASL-4 safety protocol. $25/$125 per M tokens.www.buildfastwithai.com
    33GLM-5.1OSS2026-04-0777.8%
    77.8% SWE-bench Verified (llm-stats.com); 58.4% SWE-bench Pro — SOTA at launch. Post-training upgrade to GLM-5 for agentic coding. 744B MoE, MIT license.huggingface.co
    34Qwen3.6 Plus2026-04-0278.8%
    78.8% SWE-bench Verified (OpenRouter benchmarks; 56.6% SWE-bench Pro); closed API flagship from Alibaba; 1M context. Tops Terminal-Bench 2.0 agentic coding at launch.qwen.ai
    35Gemma 4 26B A4B (MoE)OSS2026-04-0264.0%
    ⚠ SWE-bench pending (no official Google score). Multiple third-party sources ~64% on SWE-bench Verified; 77.1% LiveCodeBench v6. Runs on 24GB GPU — 3.8B active params; 256K context; Apache 2.0.huggingface.co
    36Gemma 4 31B (Dense)OSS2026-04-0264.0%
    ⚠ SWE-bench pending (no official Google score). Multiple third-party sources confirm ~64% on SWE-bench Verified. LiveCodeBench v6: 80.0%. #3 open model on Arena AI (ELO 1452); single 80GB GPU; 256K context; Apache 2.0.huggingface.co
    37MiniMax M2.7OSS2026-03-1778.0%
    78.0% SWE-bench Verified (regression from M2.5's 80.2%; self-improvement focus over coding perf); 56.22% SWE-Pro (SOTA at launch); SWE Multilingual 76.5%. MIT license, 204K context.www.minimax.io
    38GPT-5.42026-03-0578.2%
    GPT-5.4: 78.2% bash-only (vals.ai); 57.7% SWE-bench Prowww.vals.ai
    39Claude Sonnet 4.62026-02-1679.6%
    79.6% SWE-bench Verified (llm-stats.com); $3/$15 per M tokenswww.anthropic.com
    40Qwen3.5-27BOSS2026-02-1672.4%
    Dense 27B model; remarkable performance for its sizeawesomeagents.ai
    41Qwen3.5-397B-A17BOSS2026-02-1676.4%
    Large MoE Qwen3.5; strong all-round but behind K2.5llm-stats.com
    42Claude Opus 4.62026-02-1380.8%
    80.8% (llm-stats.com); 78.2% bash-only (vals.ai)llm-stats.com
    43Gemini 3.1 Pro2026-02-1380.6%
    80.6% (llm-stats.com); 78.8% bash-only (vals.ai)llm-stats.com
    44MiniMax M2.5OSS2026-02-1280.2%
    80.2% SWE-bench Verified — highest-scoring open-weight model at launch; 230B MoE, MIT license, $0.30/$1.20 per M tokenshuggingface.co
    45GLM-5OSS2026-02-1177.8%
    77.8% SWE-bench Verified (llm-stats.com); OSS SOTA for 1 day before MiniMax M2.5. 744B MoE, 40B active, MIT license. Z.AI (Zhipu).huggingface.co
    46Qwen3-Coder-NextOSS2026-02-0370.6%
    70.6% SWE-Agent / 71.3% OpenHands; 80B/3B-active MoE, 262K ctx — purpose-built agentic coding, Apache 2.0qwen.ai
    47Step-3.5-FlashOSS2026-02-0274.4%
    74.4% SWE-bench Verified; 196B-A11B MoE, 256K context. StepFun. Apache 2.0.huggingface.co
    48Kimi K2.5OSS2026-01-2676.8%
    K2.5 — best open model at launch; visual agentic capabilitieswww.kimi.com
    49GLM-4.7OSS2025-12-2273.8%
    73.8% SWE-bench Verified — OSS SOTA at launch (+5.8pp over GLM-4.6 / DeepSeek-V3.2); 355B MoE, MIT, 200K context, preserved thinking across turns. Z.AI (Zhipu).huggingface.co
    50Gemini 3 Flash2025-12-1678.0%
    78.0% SWE-bench Verified (llm-stats.com); Flash-level latency at Pro-grade reasoningllm-stats.com
    51GPT-5.22025-12-1180.0%
    GPT-5.2 at 80% (llm-stats.com)llm-stats.com
    52Kimi K2 ThinkingOSS2025-12-1072.8%
    K2 + thinking mode; best OSS pass@1 on swe-rebench Dec 2025swe-rebench.com
    53Devstral 2 (Dec 2512)OSS2025-12-0972.2%
    72.2% SWE-bench Verified (Mistral official); 123B dense model, modified MIT license, 256K context. Ships with Mistral Vibe CLI for end-to-end code automation.mistral.ai
    54Devstral Small 2OSS2025-12-0968.0%
    68.0% SWE-bench Verified (Mistral official); 24B compact variant, Apache 2.0 — laptop-deployable companion to Devstral 2. 28× fewer params, only 4.2pp below the 123B sibling.mistral.ai
    55DeepSeek-V3.2OSS2025-12-0173.1%
    73.1% — top open model through late 2025 (llm-stats.com)llm-stats.com
    56Gemini 3 Pro2025-11-1876.2%
    Gemini 3 Pro (vals.ai)www.vals.ai
    57GPT-5.12025-11-1276.3%
    76.3% SWE-bench Verified (llm-stats.com); configurable reasoning effortllm-stats.com
    58Claude Opus 4.52025-11-0180.9%
    80.9% — #1 overall at launch (codesota.com / llm-stats.com)www.codesota.com
    59Claude Sonnet 4.52025-09-2977.2%
    77.2% SWE-bench Verified (82.0% high-compute); 30-hour autonomous coding capability; #1 coding + computer-use at launch (TechCrunch, Fortune, Anthropic blog)www.anthropic.com
    60Grok Code Fast 12025-08-2670.8%
    ⚠ Contested score: xAI self-reported 70.8% (own harness); independent third-party evaluation ~57.6% — significant harness gap. Purpose-built agentic coder; 314B MoE, 256K context. $0.20/$1.50/M tokens. Available via GitHub Copilot, Cursor, Cline at launch.x.ai
    61GPT-52025-08-0774.9%
    GPT-5 SWE-bench Verified (vals.ai)www.vals.ai
    62Claude Opus 4.12025-08-0574.5%
    74.5% SWE-bench Verified; improved multi-file refactoring and long-context reasoning over Opus 4 (Anthropic blog, InfoQ)www.anthropic.com
    63DeepSeek-V3.1OSS2025-08-0167.0%
    Hybrid thinking mode; strong agentic/tool-use workflowswww.bentoml.com
    64Qwen3-Coder-480BOSS2025-07-2269.6%
    69.6% (500-turn); 67.0% standard — rivalled Claude Sonnet 4qwenlm.github.io
    65Kimi K2OSS2025-07-1665.8%
    Best open model at launch; 1T-param MoE, Apache 2.0arxiv.org
    66DeepSeek-R1-0528OSS2025-05-2857.0%
    R1 update — improved reasoning, MIT licensegithub.com
    67Claude 4 Sonnet2025-05-2272.7%
    72.7% (llm-stats.com)llm-stats.com
    68Claude Opus 42025-05-2272.5%
    72.5% (llm-stats.com)llm-stats.com
    69Devstral (May 2505)OSS2025-05-0146.8%
    Mistral purpose-built SWE coding agent; first of its kindmistral.ai
    70Qwen3-235B-A22BOSS2025-04-2959.0%
    First open model above 55% on SWE-bench Verifiedqwenlm.github.io
    71o32025-04-1671.7%
    First OpenAI model above 70% SWE-bench Verifiedopenai.com
    72Llama 4 MaverickOSS2025-04-0532.0%
    400B MoE; strong multimodal, weaker SWE vs Qwen/DeepSeekai.meta.com
    73Gemini 2.5 Pro2025-03-2563.8%
    63.8% with custom agent setup (Google blog)blog.google
    74DeepSeek-V3-0324OSS2025-03-2446.0%
    V3 update with improved reasoning; MIT licenseen.wikipedia.org
    75Gemma 3 27BOSS2025-03-1217.0%
    Strong general model, not coding-focused; low SWE scoreblog.google
    76GPT-4.52025-02-2738.0%
    Less reasoning-focused than o1; lower SWE scoreopenai.com
    77Claude 3.7 Sonnet2025-02-2470.3%
    70.3% custom harness; 63.7% bare agent (Anthropic blog)www.anthropic.com
    78Gemini 2.0 Flash2025-01-2135.0%
    SWE-bench Verified approximateblog.google
    79DeepSeek-R1OSS2025-01-2049.2%
    First open reasoning model; matched OpenAI o1 on math/codegithub.com
    80DeepSeek-V3OSS2024-12-2642.0%
    GPT-4o parity at ~$6M training cost — shocked the industryen.wikipedia.org
    81Qwen2.5-Coder-32BOSS2024-11-1139.0%
    Best OSS coder at launch; matched GPT-4o on codingqwenlm.github.io
    82Claude 3.5 Sonnet v22024-10-2249.0%
    Oct 2024 update; 49% SWE-bench Verifiedwww.codesota.com
    83Qwen2.5-72BOSS2024-09-1918.0%
    General model; moderate SWE-bench via Agentless scaffoldqwenlm.github.io
    84o12024-09-1248.9%
    First OpenAI chain-of-thought reasoning modelopenai.com
    85Llama 3.1 70BOSS2024-07-2311.0%
    Best open-source non-specialist at the timeai.meta.com
    86Claude 3.5 Sonnet2024-06-2015.0%
    Initial release; ~15% via early agent scaffoldswww.anthropic.com
    87GPT-4o2024-05-1333.2%
    OpenAI self-reported best result on SWE-bench Verifiedopenai.com
    88Claude 3 Opus2024-03-048.0%
    SWE-bench Verified via RAG scaffoldwww.anthropic.com
    89Gemini 1.5 Pro2024-02-156.0%
    Approximate via early agent evalsblog.google
    90Gemini 1.0 Pro2023-12-061.5%
    Estimated — very limited coding capability at launchblog.google
    91SWE-Llama 13BOSS2023-10-103.5%
    Princeton baseline — first open model on SWE-benchwww.swebench.com
    92Claude 22023-07-114.8%
    SWE-bench original via RAG; pre-Verified periodwww.anthropic.com
    93GPT-42023-03-1416.0%
    Best scaffold (CodeR) on SWE-bench Verified; raw model ~2–4%openai.com
    94GPT-3.5 (ChatGPT)2022-11-303.8%
    SWE-bench original (not Verified); earliest public LLM entryopenai.com