NLP2024intermediate10 min read
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
SWE-bench: هل تستطيع النماذج اللغوية حلّ مشكلات GitHub الحقيقية؟
Jimenez, C. E. · Yang, J. · Wettig, A. · Yao, S. · Pei, K. · Press, O. · Narasimhan, K. — ICLR
The problem
By 2023, language models had saturated most existing coding benchmarks like , which only test short, self-contained functions. But real software engineering is vastly harder: fixing a real bug means navigating thousands of files, understanding how functions across different modules interact, and producing code edits that fix the issue without breaking anything else. There was no that tested models on this kind of realistic, -scale software engineering.
The contribution
SWE-bench: an evaluation framework of 2,294 real software engineering tasks drawn from GitHub issues and pull requests across 12 popular Python repositories. Each task gives a model an issue description and a full codebase, and the model must generate a that passes real unit tests. The paper also introduces SWE-Llama (7B and 13B), fine-tuned on 19,000 issue–PR pairs. Results show that even the best model (Claude 2) resolves only 1.96% of issues with retrieval and 4.8% with oracle file retrieval.
The impact
SWE-bench became the standard benchmark for evaluating AI coding agents. It catalyzed a wave of agent-based systems (SWE-Agent, Devin, OpenHands) that dramatically improved on the initial baselines. The benchmark demonstrated that real software engineering — not isolated puzzles — is the frontier challenge for language models, shifting the field's focus from code completion to autonomous software development.
Imagine testing a medical student. Traditional coding benchmarks are like giving them multiple-choice anatomy quizzes — useful, but trivial for anyone with basic . SWE-bench is like handing them a real patient chart from a hospital with thousands of pages of records and saying: "Something is wrong. Find it, diagnose it, and write the treatment order — without introducing any new complications." Even the best AI "doctors" could only handle about 2 out of every 100 cases.
The gap: from coding puzzles to real engineering
By 2023, language models had become remarkably good at generating short code snippets. On HumanEval — the standard benchmark — models could solve over 90% of the problems. But these problems are toy puzzles: write a function that reverses a string, computes a factorial, or finds a palindrome. Each puzzle is self-contained, fits in a few lines, and requires no understanding of the surrounding codebase.
Real software engineering is fundamentally different. A developer fixing a bug in Django might need to read an issue report, navigate a codebase with 6,000+ files and 400,000+ lines of code, understand how a dozen functions across multiple modules interact, write a fix that touches several files, and ensure that hundreds of existing tests still pass. No benchmark captured this level of complexity — until SWE-bench.
Building the benchmark: from 90,000 PRs to 2,294 tasks
SWE-bench's construction pipeline starts by scraping pull requests from 12 popular open-source Python repositories — projects like Django, scikit-learn, matplotlib, sympy, and pytest. These repositories were chosen because they are well-maintained, have clear contributor guidelines, and have extensive test coverage. The initial scrape yields about 90,000 pull requests.
The pipeline then applies two rounds of filtering. First, an attribute filter keeps only PRs that (1) resolve a GitHub issue and (2) modify test files — indicating the contributor wrote tests to verify the fix. This reduces the pool to roughly 11,400 candidates. Second, an execution filter installs each codebase, applies the PR's changes, and runs the tests. Only instances where at least one test flips from "fail" to "pass" survive. After this, just 2,294 task instances remain — each one a verified, reproducible software engineering challenge.
What a task looks like: issue → codebase → patch → tests
Each SWE-bench task instance has four components. The issue text is a real GitHub issue — usually a bug report with reproduction steps, expected behavior, and actual behavior. The codebase is the full repository at the commit just before the fix was applied. The model must produce a — a .diff specifying exactly which lines to add, remove, or modify. Finally, tests evaluate the patch: "fail-to-pass" tests verify the fix works, and "pass-to-pass" tests ensure nothing else broke.
The numbers are striking: on average, a task involves a codebase with 3,010 files and 438,000 lines of code, and the gold solution edits 1.7 files, 3 functions, and 32.8 lines. Finding those 33 lines among 438,000 is like finding a specific paragraph in a stack of novels.
The retrieval challenge: finding the needle in the haystack
A typical SWE-bench codebase has hundreds of thousands of lines — far more than any model's can hold. The critical question becomes: which files do you show the model? The paper evaluates two retrieval strategies.
BM25 () uses keyword matching between the issue text and the code files to select the most relevant ones. It's practical but crude: in roughly half the cases, BM25 retrieves none of the files that actually need to be edited. The model is trying to fix code it has never seen.
provides exactly the files edited by the gold solution — a best-case scenario that tells us the ceiling of model capability. Even with perfect file retrieval, performance remains extremely low, showing that the bottleneck is not just retrieval but the models' reasoning and code editing abilities themselves.
SWE-Llama: fine-tuning an open model for the task
To benchmark open-source models alongside proprietary ones, the authors fine-tuned CodeLlama (7B and 13B parameters) on a separate training set of 19,000 issue–pull-request pairs from 37 additional Python repositories — completely disjoint from the evaluation set to prevent .
The uses LoRA (), modifying only the layer weights for memory efficiency, with sequences capped at 30,000 tokens (reducing the usable training data to about 10,000 instances). The resulting SWE-Llama models can process contexts exceeding 100,000 tokens and can run on consumer hardware.
An important finding: SWE-Llama performs reasonably well with oracle retrieval (matching the distribution it was trained on), but drops significantly when given BM25 retrieval. The model was trained to edit every file provided as context, so when BM25 includes irrelevant files, SWE-Llama tries to edit them too — a context that hurts performance.
Simplified to show the idea — not the real implementation.
# SWE-Llama fine-tuning setup (simplified)
from peft import LoraConfig
lora_config = LoraConfig(
r=16, # Low-rank dimension
lora_alpha=16, # Scaling factor
lora_dropout=0.05,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
)
# Training: lr=6e-4, batch=32, max 4 epochs # Max sequence length: 30,000 tokens # Checkpoint selection: best validation loss on 100 held-out instancesResults: all models struggle — badly
The headline result is sobering: with BM25 retrieval, Claude 2 resolves only 1.96% of SWE-bench issues — the best among all models tested. ChatGPT-3.5 manages just 0.17%, and GPT-4 only 1.31%. Even the fine-tuned SWE-Llama 13b resolves only 0.70%.
With oracle retrieval (perfect file selection), Claude 2 reaches 4.8% and SWE-Llama 13b reaches 3.97%. An even more generous "oracle-collapsed" setting — which shows only the edited lines ±15 lines of buffer — pushes Claude 2 to 5.93%. These numbers highlight that even when the model sees the right code, it struggles to understand and fix it.
Performance drops sharply as context length increases. Tasks with fewer than 20K input tokens are resolved at a higher rate, but beyond 100K tokens performance nearly vanishes. This aligns with the "lost in the middle" phenomenon where models fail to locate relevant information in long contexts.
What models get wrong: greedy, shallow patches
Qualitative analysis reveals consistent failure patterns. Models generate patches that are roughly half the length of human solutions: 19.6 lines on average versus 74.5 for gold patches. Models almost always edit a single file, while human solutions average 1.7 files. They rarely touch more than one function.
The root cause is a "greedy" approach: models fix the immediate symptom described in the issue without considering the broader codebase context. They write primitive Python instead of leveraging existing utility functions and libraries already available in the repository. They ignore code style conventions like relative imports. In contrast, human patches often include structural improvements and anticipate future issues — the kind of holistic thinking that current models lack.
Another major failure: generating well-formatted patch files. Even when models understand the fix needed, they frequently produce malformed .patch output that cannot be applied to the codebase. The patch application rate ranges from 21% (ChatGPT-3.5) to 67% (SWE-Llama 13b) in the oracle setting.
Why SWE-bench works as a benchmark
Several design choices make SWE-bench a robust and lasting benchmark. First, it is continually updatable: the same pipeline can scrape new PRs from any Python repository, generating fresh tasks from issues created after a model's training cutoff — guaranteeing the model has never seen the solution. Second, evaluation is execution-based and robust: each task has a median of 51 pass-to-pass tests alongside the fail-to-pass tests, creating a strong check against patches that fix one thing but break another. Third, the benchmark is open-ended: unlike cloze-style fill-in-the-blank tasks, models can generate novel solutions that differ entirely from the reference patch — they just need to pass the tests.
The paper also introduces SWE-bench Lite, a curated subset of 300 task instances focused on self-contained functional bug fixes. This provides a faster evaluation loop for researchers developing new approaches while preserving the benchmark's essential difficulty.
What SWE-bench unlocked
2023
SWE-bench
2,294 real GitHub issues as tasks. Best model (Claude 2) solves 1.96%. Establishes that real software engineering is far beyond current model capabilities.
2024
SWE-Agent
Agent-based approach with tool use and iterative debugging. Dramatically improves on the baseline retrieve-and-generate paradigm.
2024
SWE-bench Verified
Human-verified subset of 500 instances to ensure each task is solvable with only the issue description, addressing concerns about ambiguous or underspecified tasks.
2024
Agent Era
Systems like Devin, OpenHands, and CodeAct push SWE-bench Lite scores above 40%, showing that agents that plan, navigate, and iterate outperform single-shot generation.
SWE-bench's deepest impact is not any single score but the paradigm shift it triggered. Before SWE-bench, AI coding was about generating correct short functions. After SWE-bench, the field pivoted toward autonomous software engineering agents — systems that can navigate repositories, plan multi-step fixes, run tests iteratively, and reason about code at the scale that real engineering demands. The benchmark showed that the gap between "writes good code snippets" and "can actually fix bugs in real projects" is enormous, and closing it requires fundamentally different approaches than scaling up the same models on the same benchmarks.
CitationJimenez, Yang, Wettig, Yao, Pei, Press, Narasimhan. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. ICLR, 2024.
Terms in this paper
- Benchmarkالمعيار المرجعي
- Fine-Tuningالضبط الدقيق
- Code Generationتوليد الشيفرات
- Language Modelالنموذج اللغوي
- Information Retrievalاسترجاع المعلومات
- Sparse Retrievalالاسترجاع المتناثر
- Patchرُقعة
- Context Windowنافذة السياق
- BM25BM25
- HumanEvalHumanEval
- Low-Rank Adaptationالتكيّف مُنخَفِض الرُّتبة
- Test Setمجموعة بيانات الاختبار النهائي