NLP2024intermediate10 min read

SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

SWE-bench: هل تستطيع النماذج اللغوية حلّ مشكلات GitHub الحقيقية؟

Jimenez, C. E. · Yang, J. · Wettig, A. · Yao, S. · Pei, K. · Press, O. · Narasimhan, K. — ICLR

The problem

By 2023, language models had saturated most existing coding benchmarks like , which only test short, self-contained functions. But real software engineering is vastly harder: fixing a real bug means navigating thousands of files, understanding how functions across different modules interact, and producing code edits that fix the issue without breaking anything else. There was no that tested models on this kind of realistic, -scale software engineering.

The contribution

SWE-bench: an evaluation framework of 2,294 real software engineering tasks drawn from GitHub issues and pull requests across 12 popular Python repositories. Each task gives a model an issue description and a full codebase, and the model must generate a that passes real unit tests. The paper also introduces SWE-Llama (7B and 13B), fine-tuned on 19,000 issue–PR pairs. Results show that even the best model (Claude 2) resolves only 1.96% of issues with retrieval and 4.8% with oracle file retrieval.

The impact

SWE-bench became the standard benchmark for evaluating AI coding agents. It catalyzed a wave of agent-based systems (SWE-Agent, Devin, OpenHands) that dramatically improved on the initial baselines. The benchmark demonstrated that real software engineering — not isolated puzzles — is the frontier challenge for language models, shifting the field's focus from code completion to autonomous software development.

Imagine testing a medical student. Traditional coding benchmarks are like giving them multiple-choice anatomy quizzes — useful, but trivial for anyone with basic . SWE-bench is like handing them a real patient chart from a hospital with thousands of pages of records and saying: "Something is wrong. Find it, diagnose it, and write the treatment order — without introducing any new complications." Even the best AI "doctors" could only handle about 2 out of every 100 cases.

The gap: from coding puzzles to real engineering

By 2023, language models had become remarkably good at generating short code snippets. On HumanEval — the standard benchmark — models could solve over 90% of the problems. But these problems are toy puzzles: write a function that reverses a string, computes a factorial, or finds a palindrome. Each puzzle is self-contained, fits in a few lines, and requires no understanding of the surrounding codebase.

Real software engineering is fundamentally different. A developer fixing a bug in Django might need to read an issue report, navigate a codebase with 6,000+ files and 400,000+ lines of code, understand how a dozen functions across multiple modules interact, write a fix that touches several files, and ensure that hundreds of existing tests still pass. No benchmark captured this level of complexity — until SWE-bench.

Open in Lab
Compare the scope and complexity of traditional coding benchmarks vs. SWE-bench.
The demo wakes as you arrive…

Building the benchmark: from 90,000 PRs to 2,294 tasks

SWE-bench's construction pipeline starts by scraping pull requests from 12 popular open-source Python repositories — projects like Django, scikit-learn, matplotlib, sympy, and pytest. These repositories were chosen because they are well-maintained, have clear contributor guidelines, and have extensive test coverage. The initial scrape yields about 90,000 pull requests.

The pipeline then applies two rounds of filtering. First, an attribute filter keeps only PRs that (1) resolve a GitHub issue and (2) modify test files — indicating the contributor wrote tests to verify the fix. This reduces the pool to roughly 11,400 candidates. Second, an execution filter installs each codebase, applies the PR's changes, and runs the tests. Only instances where at least one test flips from "fail" to "pass" survive. After this, just 2,294 task instances remain — each one a verified, reproducible software engineering challenge.

Open in Lab
Step through the 3-stage pipeline that builds SWE-bench from raw GitHub data.
The demo wakes as you arrive…

What a task looks like: issue → codebase → patch → tests

Each SWE-bench task instance has four components. The issue text is a real GitHub issue — usually a bug report with reproduction steps, expected behavior, and actual behavior. The codebase is the full repository at the commit just before the fix was applied. The model must produce a — a .diff specifying exactly which lines to add, remove, or modify. Finally, tests evaluate the patch: "fail-to-pass" tests verify the fix works, and "pass-to-pass" tests ensure nothing else broke.

The numbers are striking: on average, a task involves a codebase with 3,010 files and 438,000 lines of code, and the gold solution edits 1.7 files, 3 functions, and 32.8 lines. Finding those 33 lines among 438,000 is like finding a specific paragraph in a stack of novels.

Open in Lab
Explore the four building blocks of every SWE-bench task instance.
The demo wakes as you arrive…

The retrieval challenge: finding the needle in the haystack

A typical SWE-bench codebase has hundreds of thousands of lines — far more than any model's can hold. The critical question becomes: which files do you show the model? The paper evaluates two retrieval strategies.

BM25 () uses keyword matching between the issue text and the code files to select the most relevant ones. It's practical but crude: in roughly half the cases, BM25 retrieves none of the files that actually need to be edited. The model is trying to fix code it has never seen.

provides exactly the files edited by the gold solution — a best-case scenario that tells us the ceiling of model capability. Even with perfect file retrieval, performance remains extremely low, showing that the bottleneck is not just retrieval but the models' reasoning and code editing abilities themselves.

Open in Lab
Toggle between BM25 and Oracle retrieval to see the performance ceiling.
The demo wakes as you arrive…

SWE-Llama: fine-tuning an open model for the task

To benchmark open-source models alongside proprietary ones, the authors fine-tuned CodeLlama (7B and 13B parameters) on a separate training set of 19,000 issue–pull-request pairs from 37 additional Python repositories — completely disjoint from the evaluation set to prevent .

The uses LoRA (), modifying only the layer weights for memory efficiency, with sequences capped at 30,000 tokens (reducing the usable training data to about 10,000 instances). The resulting SWE-Llama models can process contexts exceeding 100,000 tokens and can run on consumer hardware.

An important finding: SWE-Llama performs reasonably well with oracle retrieval (matching the distribution it was trained on), but drops significantly when given BM25 retrieval. The model was trained to edit every file provided as context, so when BM25 includes irrelevant files, SWE-Llama tries to edit them too — a context that hurts performance.

SWE-Llama Training Configurationpython

Simplified to show the idea — not the real implementation.

# SWE-Llama fine-tuning setup (simplified)
from peft import LoraConfig
lora_config = LoraConfig(
    r=16,              # Low-rank dimension
    lora_alpha=16,     # Scaling factor
    lora_dropout=0.05,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
)
# Training: lr=6e-4, batch=32, max 4 epochs # Max sequence length: 30,000 tokens # Checkpoint selection: best validation loss on 100 held-out instances

Results: all models struggle — badly

The headline result is sobering: with BM25 retrieval, Claude 2 resolves only 1.96% of SWE-bench issues — the best among all models tested. ChatGPT-3.5 manages just 0.17%, and GPT-4 only 1.31%. Even the fine-tuned SWE-Llama 13b resolves only 0.70%.

With oracle retrieval (perfect file selection), Claude 2 reaches 4.8% and SWE-Llama 13b reaches 3.97%. An even more generous "oracle-collapsed" setting — which shows only the edited lines ±15 lines of buffer — pushes Claude 2 to 5.93%. These numbers highlight that even when the model sees the right code, it struggles to understand and fix it.

Performance drops sharply as context length increases. Tasks with fewer than 20K input tokens are resolved at a higher rate, but beyond 100K tokens performance nearly vanishes. This aligns with the "lost in the middle" phenomenon where models fail to locate relevant information in long contexts.

Open in Lab
See how performance degrades as the model receives more code context.
The demo wakes as you arrive…

What models get wrong: greedy, shallow patches

Qualitative analysis reveals consistent failure patterns. Models generate patches that are roughly half the length of human solutions: 19.6 lines on average versus 74.5 for gold patches. Models almost always edit a single file, while human solutions average 1.7 files. They rarely touch more than one function.

The root cause is a "greedy" approach: models fix the immediate symptom described in the issue without considering the broader codebase context. They write primitive Python instead of leveraging existing utility functions and libraries already available in the repository. They ignore code style conventions like relative imports. In contrast, human patches often include structural improvements and anticipate future issues — the kind of holistic thinking that current models lack.

Another major failure: generating well-formatted patch files. Even when models understand the fix needed, they frequently produce malformed .patch output that cannot be applied to the codebase. The patch application rate ranges from 21% (ChatGPT-3.5) to 67% (SWE-Llama 13b) in the oracle setting.

Open in Lab
Compare the complexity of model-generated patches vs. human gold solutions.
The demo wakes as you arrive…

Why SWE-bench works as a benchmark

Several design choices make SWE-bench a robust and lasting benchmark. First, it is continually updatable: the same pipeline can scrape new PRs from any Python repository, generating fresh tasks from issues created after a model's training cutoff — guaranteeing the model has never seen the solution. Second, evaluation is execution-based and robust: each task has a median of 51 pass-to-pass tests alongside the fail-to-pass tests, creating a strong check against patches that fix one thing but break another. Third, the benchmark is open-ended: unlike cloze-style fill-in-the-blank tasks, models can generate novel solutions that differ entirely from the reference patch — they just need to pass the tests.

The paper also introduces SWE-bench Lite, a curated subset of 300 task instances focused on self-contained functional bug fixes. This provides a faster evaluation loop for researchers developing new approaches while preserving the benchmark's essential difficulty.

What SWE-bench unlocked

  1. 2023

    SWE-bench

    2,294 real GitHub issues as tasks. Best model (Claude 2) solves 1.96%. Establishes that real software engineering is far beyond current model capabilities.

  2. 2024

    SWE-Agent

    Agent-based approach with tool use and iterative debugging. Dramatically improves on the baseline retrieve-and-generate paradigm.

  3. 2024

    SWE-bench Verified

    Human-verified subset of 500 instances to ensure each task is solvable with only the issue description, addressing concerns about ambiguous or underspecified tasks.

  4. 2024

    Agent Era

    Systems like Devin, OpenHands, and CodeAct push SWE-bench Lite scores above 40%, showing that agents that plan, navigate, and iterate outperform single-shot generation.

SWE-bench's deepest impact is not any single score but the paradigm shift it triggered. Before SWE-bench, AI coding was about generating correct short functions. After SWE-bench, the field pivoted toward autonomous software engineering agents — systems that can navigate repositories, plan multi-step fixes, run tests iteratively, and reason about code at the scale that real engineering demands. The benchmark showed that the gap between "writes good code snippets" and "can actually fix bugs in real projects" is enormous, and closing it requires fundamentally different approaches than scaling up the same models on the same benchmarks.

CitationJimenez, Yang, Wettig, Yao, Pei, Press, Narasimhan. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. ICLR, 2024.

Terms in this paper