AI Progress Tracker, SWE-bench Verified (2024–2026)
Last updated 2026-06-24 · Updated weeklySWE-BENCH VERIFIED · 500 REAL GITHUB ISSUES · Y = % ISSUES RESOLVED
★ = SOTA OPEN-SOURCE AT LAUNCH · TAP/HOVER DOTS FOR DETAIL · TRACKING HOW FAST AI IS MOVING, AND WHAT IT MEANS FOR YOU
⏱ Time for OSS to catch each frontier model
Frontier
Claude Mythos 5 (95.5%)
2026-06-09
First OSS Match
Not yet caught
No OSS model has matched this score
–
open gap
Frontier
Claude Fable 5 (95%)
2026-06-09
First OSS Match
Not yet caught
No OSS model has matched this score
–
open gap
Frontier
GPT-5.5 (88.7%)
2026-04-23
First OSS Match
Not yet caught
No OSS model has matched this score
–
open gap
Frontier
Claude Opus 4.7 (87.6%)
2026-04-16
First OSS Match
Not yet caught
No OSS model has matched this score
–
open gap
Frontier
GPT-5.3 Codex (85%)
2026-04-10
First OSS Match
Not yet caught
No OSS model has matched this score
–
open gap
Frontier
Claude Mythos Preview (93.9%)
2026-04-07
First OSS Match
Not yet caught
No OSS model has matched this score
–
open gap
Frontier
GPT-5.2 (80%)
2025-12-11
First OSS Match
MiniMax M2.5
2026-02-12 · 80.2%
63d
2.1 mo
Frontier
GPT-5.1 (76.3%)
2025-11-12
First OSS Match
Kimi K2.5
2026-01-26 · 76.8%
75d
2.5 mo
Frontier
Claude Opus 4.5 (80.9%)
2025-11-01
First OSS Match
Not yet caught
No OSS model has matched this score
–
open gap
Frontier
Claude Sonnet 4.5 (77.2%)
2025-09-29
First OSS Match
GLM-5
2026-02-11 · 77.8%
135d
4.4 mo
Frontier
GPT-5 (74.9%)
2025-08-07
First OSS Match
Kimi K2.5
2026-01-26 · 76.8%
172d
5.7 mo
Frontier
Claude Opus 4.1 (74.5%)
2025-08-05
First OSS Match
Kimi K2.5
2026-01-26 · 76.8%
174d
5.7 mo
Frontier
Claude 4 Sonnet (72.7%)
2025-05-22
First OSS Match
DeepSeek-V3.2
2025-12-01 · 73.1%
193d
6.3 mo
Frontier
o3 (71.7%)
2025-04-16
First OSS Match
DeepSeek-V3.2
2025-12-01 · 73.1%
229d
7.5 mo
Frontier
Claude 3.7 Sonnet (70.3%)
2025-02-24
First OSS Match
DeepSeek-V3.2
2025-12-01 · 73.1%
280d
9.2 mo
Frontier
Claude 3.5 Sonnet v2 (49%)
2024-10-22
First OSS Match
DeepSeek-R1
2025-01-20 · 49.2%
90d
3.0 mo
Frontier
o1 (48.9%)
2024-09-12
First OSS Match
DeepSeek-R1
2025-01-20 · 49.2%
130d
4.3 mo
Frontier
GPT-4o (33.2%)
2024-05-13
First OSS Match
Qwen2.5-Coder-32B
2024-11-11 · 39%
182d
6.0 mo
▾ FULL DATA TABLE, all models, dates, scores, sources
| # | Model | Released | Score | Bar | Notes | Source |
|---|---|---|---|---|---|---|
| 1 | GLM-5.2OSS⚠ | 2026-06-16 | Pending | ⚠ SWE-bench Verified pending — Z.AI published no benchmarks at launch (third-party SWE-bench Pro: 62.1%). Flagship 744B-total / 40B-active MoE with usable 1M-token context and dual reasoning modes. | www.marktechpost.com | |
| 2 | Kimi K2.7 CodeOSS⚠ | 2026-06-12 | 78.2% | ⚠ Vendor-reported: 78.2% SWE-bench Verified (Moonshot AI); no independent third-party verification at launch. Coding-focused 1T-parameter open-weight; uses 30% fewer reasoning tokens than K2.6. | huggingface.co | |
| 3 | DiffusionGemma 26B-A4BOSS⚠ | 2026-06-10 | Pending | ⚠ SWE-bench Verified pending — speed-focused research model with no published score. Google DeepMind experimental open-weight diffusion-based text LLM (26B MoE, Apache 2.0); generates 4× faster than autoregressive Gemma 4. | blog.google | |
| 4 | Claude Fable 5 | 2026-06-09 | 95.0% | 95.0% SWE-bench Verified — new public SOTA, #2 on BenchLM leaderboard. First publicly-available Mythos-class model (tier above Opus 4.8); 1M context. | www.anthropic.com | |
| 5 | Claude Mythos 5 | 2026-06-09 | 95.5% | 95.5% SWE-bench Verified — #1 on BenchLM leaderboard. Restricted-access variant of Fable 5 (Project Glasswing tier); same weights, lifted safeguards for authorized cybersecurity orgs. | www.anthropic.com | |
| 6 | North Mini Code 1.0OSS⚠ | 2026-06-09 | 80.2% | ⚠ Caveat: 80.2% pass@10 SWE-bench Verified (Cohere blog) — not pass@1, so direct comparison is misleading. 30B-A3B MoE, Apache 2.0; single H100 deployable. | cohere.com | |
| 7 | Unisound U2⚠ | 2026-06-07 | 75.0% | ⚠ Vendor-reported 75% SWE-bench Verified (Unisound press release). Chinese agentic foundation model capable of 100+ step autonomous workflows. | www.prnewswire.com | |
| 8 | Nemotron 3 UltraOSS | 2026-06-04 | 71.9% | 71.9% SWE-bench Verified (NVIDIA technical report, BenchLM); 550B-total / 55B-active hybrid Mamba-Transformer MoE optimized for long-running agents. | research.nvidia.com | |
| 9 | MAI-Code-1-Flash | 2026-06-02 | 71.6% | 71.6% SWE-bench Verified (Microsoft AI announcement); first in-house Microsoft coding model — 137B sparse MoE, 5B active. Launched at Build 2026; beats Claude Haiku 4.5 (66.6%). | microsoft.ai | |
| 10 | MAI-Thinking-1 | 2026-06-02 | 73.5% | 73.5% SWE-bench Verified (BenchLM); first in-house Microsoft reasoning model. 35B active / 1T total MoE, trained from scratch without OpenAI distillation. | microsoft.ai | |
| 11 | MiniMax M3OSS⚠ | 2026-06-01 | 80.5% | ⚠ Contested: 80.5% SWE-bench Verified (vendor-reported, llm-stats listing); only 59% SWE-bench Pro suggests harness gap. Open-weight 1M-context MoE with MSA architecture. | venturebeat.com | |
| 12 | Grok Build 0.1⚠ | 2026-05-29 | 70.8% | ⚠ Vendor-reported: 70.8% SWE-bench Verified (xAI announcement, own harness); no independent score yet. Dedicated coding model behind Grok Build CLI, 256K context. | x.ai | |
| 13 | Step 3.7 FlashOSS⚠ | 2026-05-29 | Pending | ⚠ SWE-bench Verified pending — no standalone score published. StepFun 198B sparse MoE vision-language coding model, Apache 2.0; vendor claims ~97% of Claude Opus 4.6's Verified performance with Advisor Mode. | www.marktechpost.com | |
| 14 | Claude Opus 4.8 | 2026-05-28 | 88.6% | 88.6% SWE-bench Verified (llm-stats.com, BenchLM); adaptive-thinking Opus refresh with 1M context window. | www.anthropic.com | |
| 15 | Qwen3.7-Max | 2026-05-20 | 80.4% | 80.4% SWE-bench Verified (felloai.com, DataCamp, startupfortune.com); 60.6% SWE-bench Pro (highest at launch). Closed API flagship. 1M context. Launched at Alibaba Cloud Summit Hangzhou. | qwen.ai | |
| 16 | Gemini 3.5 Flash | 2026-05-19 | 78.0% | 78% SWE-bench Verified (Google I/O 2026); beats Gemini 3.1 Pro on 11/15 agentic benchmarks; 4× faster than rival frontier models; 55.1% SWE-bench Pro. Closed model. | blog.google | |
| 17 | GPT-5.5 Instant⚠ | 2026-05-05 | Pending | ⚠ SWE-bench Verified pending — no independent score published. Lighter/faster default-tier variant of GPT-5.5; became the new ChatGPT default May 5. $1.50/$6/M tokens. | techcrunch.com | |
| 18 | Grok 4.3⚠ | 2026-04-30 | Pending | ⚠ SWE-bench Verified pending (no official score published). Full API release April 30; improved agentic performance over Grok 4.20; 40% price cut, $0.80/$3.20/M tokens. | artificialanalysis.ai | |
| 19 | Mistral Medium 3.5OSS | 2026-04-29 | 77.6% | 77.6% SWE-bench Verified (MarkTechPost / winbuzzer); unified 128B model merging chat+reasoning+code, replaces Devstral 2 and Magistral. Modified MIT license. 256K context, $1.50/M input. | huggingface.co | |
| 20 | DeepSeek-V4-ProOSS★ | 2026-04-24 | 80.6% | 80.6% SWE-bench Verified — new OSS SOTA, 0.4pp above MiniMax M2.5/Kimi K2.6. 1.6T param MoE, 49B active, 1M context. MIT license. $0.30/$1.20/M tokens. | huggingface.co | |
| 21 | DeepSeek-V4-FlashOSS | 2026-04-24 | 79.0% | 79.0% SWE-bench Verified (BenchLM); 284B-A13B MoE — 12× cheaper than V4-Pro at 1.6pp lower SWE-bench Verified. MIT license, 1M context. | huggingface.co | |
| 22 | GPT-5.5 | 2026-04-23 | 88.7% | 88.7% SWE-bench Verified — new public SOTA, beats Claude Opus 4.7 (87.6%). First fully retrained base since GPT-4.5. 1M context, $5/$30/M tokens. Codename 'Spud'. | openai.com | |
| 23 | Hy3-previewOSS | 2026-04-23 | 74.4% | 74.4% SWE-bench Verified (Artificial Analysis, multiple sources); 40% jump from Hy2 (53.0%). Tencent. 295B MoE / 21B active, 256K context. Tencent Hy Community License. | huggingface.co | |
| 24 | Qwen3.6-27BOSS | 2026-04-22 | 77.2% | 77.2% SWE-bench Verified (llm-stats.com / the-decoder.com); 27B dense model, Apache 2.0. Surpasses Qwen3.5-397B-A17B on all major coding benchmarks despite 14× fewer params. | huggingface.co | |
| 25 | MiMo-V2.5-ProOSS | 2026-04-22 | 78.9% | 78.9% SWE-bench Verified (artificialanalysis.ai, the-decoder.com); 57.2% SWE-bench Pro. Xiaomi. 1.02T MoE, 42B active, MIT license, 1M context. | huggingface.co | |
| 26 | Qwen3.6-Max-Preview⚠ | 2026-04-20 | Pending | ⚠ SWE-bench Verified pending (no confirmed score). Leads SWE-bench Pro, Terminal-Bench 2.0, SkillsBench at launch (#1 on 6 benchmarks). Closed API flagship from Alibaba. 260K context. | qwen.ai | |
| 27 | Kimi K2.6OSS | 2026-04-20 | 80.2% | 80.2% SWE-bench Verified; 58.6% SWE-bench Pro. 1T params / 32B active MoE. 300-agent swarm, 4000-step horizon, 262K context. Modified MIT. | huggingface.co | |
| 28 | Claude Opus 4.7 | 2026-04-16 | 87.6% | 87.6% SWE-bench Verified; 64.3% SWE-bench Pro (#1 public model at launch). xhigh effort, 3x vision, /ultrareview, 1M context. $5/$25 per M tokens. | www.anthropic.com | |
| 29 | Qwen3.6-35B-A3BOSS | 2026-04-16 | 73.4% | 73.4% SWE-bench Verified. Apache 2.0. 35B total / 3B active MoE — runs on a laptop. Highest OSS SWE score on consumer hardware. | huggingface.co | |
| 30 | GPT-5.3 Codex | 2026-04-10 | 85.0% | 85.0% SWE-bench Verified — OpenAI coding-specialist Codex variant; 56.8% SWE-bench Pro (benchlm.ai) | benchlm.ai | |
| 31 | Muse Spark | 2026-04-08 | 77.4% | 77.4% SWE-bench Verified (The Next Web, llm-stats.com, vals.ai); 55.0% SWE-bench Pro. Meta Superintelligence Labs — first flagship closed-source Meta model. Natively multimodal. | ai.meta.com | |
| 32 | Claude Mythos Preview | 2026-04-07 | 93.9% | 93.9% SWE-bench Verified — most capable model ever built; NOT publicly released. Restricted to 50 orgs under Project Glasswing (cybersecurity). Triggered ASL-4 safety protocol. $25/$125 per M tokens. | www.buildfastwithai.com | |
| 33 | GLM-5.1OSS | 2026-04-07 | 77.8% | 77.8% SWE-bench Verified (llm-stats.com); 58.4% SWE-bench Pro — SOTA at launch. Post-training upgrade to GLM-5 for agentic coding. 744B MoE, MIT license. | huggingface.co | |
| 34 | Qwen3.6 Plus | 2026-04-02 | 78.8% | 78.8% SWE-bench Verified (OpenRouter benchmarks; 56.6% SWE-bench Pro); closed API flagship from Alibaba; 1M context. Tops Terminal-Bench 2.0 agentic coding at launch. | qwen.ai | |
| 35 | Gemma 4 26B A4B (MoE)OSS⚠ | 2026-04-02 | 64.0% | ⚠ SWE-bench pending (no official Google score). Multiple third-party sources ~64% on SWE-bench Verified; 77.1% LiveCodeBench v6. Runs on 24GB GPU — 3.8B active params; 256K context; Apache 2.0. | huggingface.co | |
| 36 | Gemma 4 31B (Dense)OSS⚠ | 2026-04-02 | 64.0% | ⚠ SWE-bench pending (no official Google score). Multiple third-party sources confirm ~64% on SWE-bench Verified. LiveCodeBench v6: 80.0%. #3 open model on Arena AI (ELO 1452); single 80GB GPU; 256K context; Apache 2.0. | huggingface.co | |
| 37 | MiniMax M2.7OSS | 2026-03-17 | 78.0% | 78.0% SWE-bench Verified (regression from M2.5's 80.2%; self-improvement focus over coding perf); 56.22% SWE-Pro (SOTA at launch); SWE Multilingual 76.5%. MIT license, 204K context. | www.minimax.io | |
| 38 | GPT-5.4 | 2026-03-05 | 78.2% | GPT-5.4: 78.2% bash-only (vals.ai); 57.7% SWE-bench Pro | www.vals.ai | |
| 39 | Claude Sonnet 4.6 | 2026-02-16 | 79.6% | 79.6% SWE-bench Verified (llm-stats.com); $3/$15 per M tokens | www.anthropic.com | |
| 40 | Qwen3.5-27BOSS | 2026-02-16 | 72.4% | Dense 27B model; remarkable performance for its size | awesomeagents.ai | |
| 41 | Qwen3.5-397B-A17BOSS | 2026-02-16 | 76.4% | Large MoE Qwen3.5; strong all-round but behind K2.5 | llm-stats.com | |
| 42 | Claude Opus 4.6 | 2026-02-13 | 80.8% | 80.8% (llm-stats.com); 78.2% bash-only (vals.ai) | llm-stats.com | |
| 43 | Gemini 3.1 Pro | 2026-02-13 | 80.6% | 80.6% (llm-stats.com); 78.8% bash-only (vals.ai) | llm-stats.com | |
| 44 | MiniMax M2.5OSS★ | 2026-02-12 | 80.2% | 80.2% SWE-bench Verified — highest-scoring open-weight model at launch; 230B MoE, MIT license, $0.30/$1.20 per M tokens | huggingface.co | |
| 45 | GLM-5OSS★ | 2026-02-11 | 77.8% | 77.8% SWE-bench Verified (llm-stats.com); OSS SOTA for 1 day before MiniMax M2.5. 744B MoE, 40B active, MIT license. Z.AI (Zhipu). | huggingface.co | |
| 46 | Qwen3-Coder-NextOSS | 2026-02-03 | 70.6% | 70.6% SWE-Agent / 71.3% OpenHands; 80B/3B-active MoE, 262K ctx — purpose-built agentic coding, Apache 2.0 | qwen.ai | |
| 47 | Step-3.5-FlashOSS | 2026-02-02 | 74.4% | 74.4% SWE-bench Verified; 196B-A11B MoE, 256K context. StepFun. Apache 2.0. | huggingface.co | |
| 48 | Kimi K2.5OSS★ | 2026-01-26 | 76.8% | K2.5 — best open model at launch; visual agentic capabilities | www.kimi.com | |
| 49 | GLM-4.7OSS★ | 2025-12-22 | 73.8% | 73.8% SWE-bench Verified — OSS SOTA at launch (+5.8pp over GLM-4.6 / DeepSeek-V3.2); 355B MoE, MIT, 200K context, preserved thinking across turns. Z.AI (Zhipu). | huggingface.co | |
| 50 | Gemini 3 Flash | 2025-12-16 | 78.0% | 78.0% SWE-bench Verified (llm-stats.com); Flash-level latency at Pro-grade reasoning | llm-stats.com | |
| 51 | GPT-5.2 | 2025-12-11 | 80.0% | GPT-5.2 at 80% (llm-stats.com) | llm-stats.com | |
| 52 | Kimi K2 ThinkingOSS | 2025-12-10 | 72.8% | K2 + thinking mode; best OSS pass@1 on swe-rebench Dec 2025 | swe-rebench.com | |
| 53 | Devstral 2 (Dec 2512)OSS | 2025-12-09 | 72.2% | 72.2% SWE-bench Verified (Mistral official); 123B dense model, modified MIT license, 256K context. Ships with Mistral Vibe CLI for end-to-end code automation. | mistral.ai | |
| 54 | Devstral Small 2OSS | 2025-12-09 | 68.0% | 68.0% SWE-bench Verified (Mistral official); 24B compact variant, Apache 2.0 — laptop-deployable companion to Devstral 2. 28× fewer params, only 4.2pp below the 123B sibling. | mistral.ai | |
| 55 | DeepSeek-V3.2OSS★ | 2025-12-01 | 73.1% | 73.1% — top open model through late 2025 (llm-stats.com) | llm-stats.com | |
| 56 | Gemini 3 Pro | 2025-11-18 | 76.2% | Gemini 3 Pro (vals.ai) | www.vals.ai | |
| 57 | GPT-5.1 | 2025-11-12 | 76.3% | 76.3% SWE-bench Verified (llm-stats.com); configurable reasoning effort | llm-stats.com | |
| 58 | Claude Opus 4.5 | 2025-11-01 | 80.9% | 80.9% — #1 overall at launch (codesota.com / llm-stats.com) | www.codesota.com | |
| 59 | Claude Sonnet 4.5 | 2025-09-29 | 77.2% | 77.2% SWE-bench Verified (82.0% high-compute); 30-hour autonomous coding capability; #1 coding + computer-use at launch (TechCrunch, Fortune, Anthropic blog) | www.anthropic.com | |
| 60 | Grok Code Fast 1⚠ | 2025-08-26 | 70.8% | ⚠ Contested score: xAI self-reported 70.8% (own harness); independent third-party evaluation ~57.6% — significant harness gap. Purpose-built agentic coder; 314B MoE, 256K context. $0.20/$1.50/M tokens. Available via GitHub Copilot, Cursor, Cline at launch. | x.ai | |
| 61 | GPT-5 | 2025-08-07 | 74.9% | GPT-5 SWE-bench Verified (vals.ai) | www.vals.ai | |
| 62 | Claude Opus 4.1 | 2025-08-05 | 74.5% | 74.5% SWE-bench Verified; improved multi-file refactoring and long-context reasoning over Opus 4 (Anthropic blog, InfoQ) | www.anthropic.com | |
| 63 | DeepSeek-V3.1OSS | 2025-08-01 | 67.0% | Hybrid thinking mode; strong agentic/tool-use workflows | www.bentoml.com | |
| 64 | Qwen3-Coder-480BOSS★ | 2025-07-22 | 69.6% | 69.6% (500-turn); 67.0% standard — rivalled Claude Sonnet 4 | qwenlm.github.io | |
| 65 | Kimi K2OSS★ | 2025-07-16 | 65.8% | Best open model at launch; 1T-param MoE, Apache 2.0 | arxiv.org | |
| 66 | DeepSeek-R1-0528OSS | 2025-05-28 | 57.0% | R1 update — improved reasoning, MIT license | github.com | |
| 67 | Claude 4 Sonnet | 2025-05-22 | 72.7% | 72.7% (llm-stats.com) | llm-stats.com | |
| 68 | Claude Opus 4 | 2025-05-22 | 72.5% | 72.5% (llm-stats.com) | llm-stats.com | |
| 69 | Devstral (May 2505)OSS | 2025-05-01 | 46.8% | Mistral purpose-built SWE coding agent; first of its kind | mistral.ai | |
| 70 | Qwen3-235B-A22BOSS★ | 2025-04-29 | 59.0% | First open model above 55% on SWE-bench Verified | qwenlm.github.io | |
| 71 | o3 | 2025-04-16 | 71.7% | First OpenAI model above 70% SWE-bench Verified | openai.com | |
| 72 | Llama 4 MaverickOSS | 2025-04-05 | 32.0% | 400B MoE; strong multimodal, weaker SWE vs Qwen/DeepSeek | ai.meta.com | |
| 73 | Gemini 2.5 Pro | 2025-03-25 | 63.8% | 63.8% with custom agent setup (Google blog) | blog.google | |
| 74 | DeepSeek-V3-0324OSS | 2025-03-24 | 46.0% | V3 update with improved reasoning; MIT license | en.wikipedia.org | |
| 75 | Gemma 3 27BOSS | 2025-03-12 | 17.0% | Strong general model, not coding-focused; low SWE score | blog.google | |
| 76 | GPT-4.5 | 2025-02-27 | 38.0% | Less reasoning-focused than o1; lower SWE score | openai.com | |
| 77 | Claude 3.7 Sonnet | 2025-02-24 | 70.3% | 70.3% custom harness; 63.7% bare agent (Anthropic blog) | www.anthropic.com | |
| 78 | Gemini 2.0 Flash | 2025-01-21 | 35.0% | SWE-bench Verified approximate | blog.google | |
| 79 | DeepSeek-R1OSS★ | 2025-01-20 | 49.2% | First open reasoning model; matched OpenAI o1 on math/code | github.com | |
| 80 | DeepSeek-V3OSS★ | 2024-12-26 | 42.0% | GPT-4o parity at ~$6M training cost — shocked the industry | en.wikipedia.org | |
| 81 | Qwen2.5-Coder-32BOSS★ | 2024-11-11 | 39.0% | Best OSS coder at launch; matched GPT-4o on coding | qwenlm.github.io | |
| 82 | Claude 3.5 Sonnet v2 | 2024-10-22 | 49.0% | Oct 2024 update; 49% SWE-bench Verified | www.codesota.com | |
| 83 | Qwen2.5-72BOSS | 2024-09-19 | 18.0% | General model; moderate SWE-bench via Agentless scaffold | qwenlm.github.io | |
| 84 | o1 | 2024-09-12 | 48.9% | First OpenAI chain-of-thought reasoning model | openai.com | |
| 85 | Llama 3.1 70BOSS★ | 2024-07-23 | 11.0% | Best open-source non-specialist at the time | ai.meta.com | |
| 86 | Claude 3.5 Sonnet | 2024-06-20 | 15.0% | Initial release; ~15% via early agent scaffolds | www.anthropic.com | |
| 87 | GPT-4o | 2024-05-13 | 33.2% | OpenAI self-reported best result on SWE-bench Verified | openai.com | |
| 88 | Claude 3 Opus | 2024-03-04 | 8.0% | SWE-bench Verified via RAG scaffold | www.anthropic.com | |
| 89 | Gemini 1.5 Pro | 2024-02-15 | 6.0% | Approximate via early agent evals | blog.google | |
| 90 | Gemini 1.0 Pro | 2023-12-06 | 1.5% | Estimated — very limited coding capability at launch | blog.google | |
| 91 | SWE-Llama 13BOSS★ | 2023-10-10 | 3.5% | Princeton baseline — first open model on SWE-bench | www.swebench.com | |
| 92 | Claude 2 | 2023-07-11 | 4.8% | SWE-bench original via RAG; pre-Verified period | www.anthropic.com | |
| 93 | GPT-4 | 2023-03-14 | 16.0% | Best scaffold (CodeR) on SWE-bench Verified; raw model ~2–4% | openai.com | |
| 94 | GPT-3.5 (ChatGPT) | 2022-11-30 | 3.8% | SWE-bench original (not Verified); earliest public LLM entry | openai.com |