NLP2021intermediate14 min read
Evaluating Large Language Models Trained on Code
تقييم النماذج اللغوية الكبيرة المُدرَّبة على الشيفرات البرمجية
Chen, M. · Tworek, J. · Jun, H. · Yuan, Q. · Pinto, H.P.O. · Kaplan, J. · Edwards, H. · Burda, Y. · Joseph, N. · Brockman, G. · Ray, A. · Puri, R. · Krueger, G. · Petrov, M. · Khlaaf, H. · Sastry, G. · Mishkin, P. · Chan, B. · Gray, S. · Ryder, N. · Pavlov, M. · Power, A. · Kaiser, L. · Bavarian, M. · Winter, C. · Tillet, P. · Such, F.P. · Cummings, D. · Plappert, M. · Chantzis, F. · Barnes, E. · Herbert-Voss, A. · Guss, W.H. · Nichol, A. · Paino, A. · Tezak, N. · Tang, J. · Babuschkin, I. · Balaji, S. · Jain, S. · Saunders, W. · Hesse, C. · Carr, A.N. · Leike, J. · Achiam, J. · Misra, V. · Morikawa, E. · Radford, A. · Knight, M. · Brundage, M. · Murati, M. · Mayer, K. · Welinder, P. · McGrew, B. · Amodei, D. · McCandlish, S. · Sutskever, I. · Zaremba, W. — arXiv
The problem
By 2021, large language models like GPT-3 had shown impressive capabilities in natural language tasks, yet their ability to write functional code was nearly zero. GPT-3 solved 0% of programming problems on the HumanEval . Meanwhile, the demand for AI-assisted programming was growing rapidly, but there was no reliable way to generate correct code from natural language descriptions, and no principled benchmark to measure progress on this front.
The contribution
Codex: a GPT model fine-tuned on 159 GB of Python code from GitHub. Codex-12B solves 28.8% of HumanEval problems with a single sample (pass@1), compared to 0% for GPT-3. With 100 samples per problem, it solves 72.3%. The paper also introduces HumanEval, a benchmark of 164 hand-written problems measuring , the pass@k metric with an , and Codex-S (supervised fine-tuned variant achieving 37.7% pass@1). A production version of Codex powers GitHub Copilot.
The impact
Codex proved that language models can become practical programming tools when fine-tuned on code, launching the era of AI-powered code assistants. HumanEval became the standard benchmark for , adopted by virtually every subsequent model. The pass@k framework established how to evaluate stochastic code generators. The paper's insights on repeated sampling, tuning, and became foundational techniques. Its direct descendants include Phi-1, StarCoder, and it inspired benchmarks like SWE-bench and applications like Voyager.
Imagine a brilliant linguistics student who has read every novel ever written but has never seen a line of code. You hand them a programming manual and say: "Read this and start writing Python." They might grasp the grammar, but they would produce broken programs that look plausible but don't run.
Now imagine instead you send that student to work at a software company for a year, reading and writing real code all day. When they return, you give them a plain English description of a function and they write working code on the first try — or at least on one of their first few attempts.
That is exactly what this paper does with GPT-3: fine-tune it on millions of real Python files so it transforms from a zero-score novice into a capable code generator.
From language model to code generator: the Codex pipeline
GPT-3 was trained on a massive corpus of internet text, giving it powerful natural language understanding. But code is a fundamentally different domain: it has strict syntax, requires logical consistency, and must pass unit tests to be considered correct. Simply put, knowing how to write an essay does not mean you can write a Python function.
The key insight of this paper is that the same autoregressive architecture that excels at predicting the next word in a sentence can also predict the next in a program — if you train it on enough code. The authors collected 159 GB of Python files from 54 million public GitHub repositories and fine-tuned GPT-3 on this data to produce Codex.
Importantly, even though Codex was fine-tuned from GPT-3 (which already contains natural language representations), the authors found that starting from a pre-trained language model did not improve final performance compared to from scratch. The advantage of starting from GPT-3 was faster , not better final accuracy.
A critical engineering detail: the GPT-3 was designed for natural language, where whitespace is minimal. But Python code relies heavily on indentation — every tab and space matters. The authors added special tokens for whitespace runs of varying lengths, reducing the token count by roughly 30%. This seemingly small change is important: fewer tokens per file means the model can see more code context within its fixed , effectively giving it a wider field of view when generating programs.
HumanEval: measuring functional correctness
How do you know if generated code actually works? Before this paper, code generation was mostly evaluated with — a metric borrowed from machine translation that measures text overlap between the generated code and a reference solution. But two programs that look very different textually can be functionally identical, and two programs that look almost identical can behave completely differently due to a single character change.
The authors introduced HumanEval: 164 hand-written Python programming problems, each with a function signature, a describing the task, and an average of 7.7 unit tests. A solution is correct if and only if it passes all unit tests. This is the same standard human developers use — in test-driven development, success means passing the test suite.
Critically, these problems were hand-written specifically for this evaluation. Since Codex was trained on a large fraction of GitHub, existing benchmark problems (like those from competitive programming sites) might already appear in the training data. Hand-writing new problems avoids this risk.
The key metric is pass@k: generate k code samples for each problem, and the problem is considered solved if any of the k samples passes all unit tests. A naive estimate of pass@k has high variance, so the authors derive an unbiased estimator. In practice, they generate samples (typically ), count the number of correct ones , and compute:
Scaling laws: bigger models write better code
Just as GPT-3 showed that language model performance scales as a smooth with model size, the Codex paper reveals that the same scaling pattern holds after code fine-tuning. Test loss on a held-out Python corpus follows the relationship , where is the number of non-embedding parameters.
This scaling extends to the functional metric that matters most: pass@k. As model size increases from 12M to 12B parameters, pass@1 (with temperature ) rises from 2.0% to 28.8%, and pass@100 (with temperature ) rises from 8.6% to 72.3%. The relationship follows a curve in log-parameters, suggesting that there is a critical scale at which the model transitions from generating mostly broken code to generating mostly correct code.
An important practical insight: the optimal sampling temperature depends on how many samples you generate. For pass@1 (you only get one shot), a low temperature of is best — the model plays it safe and picks the most likely tokens. But for pass@100 (you get many shots), a higher temperature of is optimal, because diverse samples are more likely to include at least one correct solution.
Think of it like a talent show: if you only have one contestant, you want the most reliable performer. But if you have 100 contestants, you want diversity — someone unconventional might give a brilliant performance that a room full of safe choices would never produce.
Selecting the best candidate without running tests
The pass@k metric assumes you have an oracle — you know which sample is correct because you have the unit tests. But in real-world use (like an autocomplete tool), you don't have tests. You generate multiple candidates and must pick one to show the user. How?
The authors compared three strategies. First, random selection: just pick any sample. Second, mean log-probability ranking: pick the sample whose tokens have the highest average log-probability under the model — essentially, the sample the model is most confident about. Third, back-translation: use the Codex-D model (which generates docstrings from code) to pick the sample whose code best describes the original prompt.
Mean log-probability ranking significantly outperformed random selection and slightly outperformed back-translation. For Codex-12B at temperature 0.8 with 100 samples, mean log-prob ranking achieves 44.5% accuracy, compared to the oracle's 77.5%. This means without any test execution, you can recover more than half of the oracle's advantage simply by asking: "which answer does the model believe in most?"
Codex-S: supervised fine-tuning for sharper code
Most Python code on GitHub is not standalone functions with docstrings — it includes class definitions, configuration files, scripts, and even data storage files. This distribution mismatch means Codex spends much of its capacity modeling code patterns that are irrelevant to the specific task of synthesizing functions from docstrings.
To sharpen Codex's focus, the authors built Codex-S by further fine-tuning on a curated set of correctly implemented standalone functions. These came from two sources: competitive programming websites (which test algorithmic reasoning) and open-source projects with continuous integration (which test practical implementation skills). In total, they assembled about 50,000 training problems.
A quality control step was essential: they used Codex-12B itself to generate 100 samples per curated problem and filtered out any problem where no sample passed the unit tests. This removed ambiguous or excessively difficult problems that would inject noise into training.
The results were striking. Codex-S achieved 37.7% pass@1 and 77.5% pass@100, compared to 28.8% and 72.3% for the base Codex — making Codex-S roughly one to two orders of magnitude more parameter-efficient than Codex for the same performance level.
Codex-D: describing code in natural language
The Codex pipeline works in one direction: docstring → code. But the authors also wanted the reverse: code → docstring. This has a safety motivation — if a model generates code, a companion model that describes what that code does helps users verify the output's intent.
They created Codex-D by rearranging the training data: instead of (signature + docstring + body), they used (signature + body + docstring) and trained to predict the docstring. When graded by hand (since there's no automatic way to judge natural language quality), Codex-D achieved pass@1 of 20.3% and pass@10 of 46.5%, comparable to but slightly below Codex-S's code generation rates.
Interestingly, the docstring model reflects the data it was trained on. It sometimes produces gems like "I just found this function online" or "This test is not correctly written and it's not my solution" — mimicking the casual documentation style found in real GitHub repositories.
Limitations: where Codex struggles
Despite impressive results, Codex has clear failure patterns. The most revealing is the chain-of-operations degradation: the authors created synthetic tasks by chaining simple string operations ("convert to lowercase," "remove every third character," "reverse the word order") into longer sequences. Each individual operation is easy for Codex to handle, but as the chain grows, pass rates drop exponentially — roughly halving with each additional operation.
This is strikingly unlike human behavior. A human programmer who can implement a chain of length two can typically implement a chain of arbitrary length. For Codex, each additional step compounds the probability of failure, because the model cannot decompose a long specification into independently solvable subproblems.
A second failure pattern is variable binding errors. When a docstring describes multiple operations on multiple variables, Codex often applies the right operation to the wrong variable. For example, when told to "add 3 to y, subtract 4 from x and w, and return the product of all four numbers," Codex may forget to decrement w and incorrectly compute the product. This reflects a fundamental difficulty with tracking state across a specification — the model sees tokens, not a variable table.
Simplified to show the idea — not the real implementation.
def do_work(x, y, z, w):
"""Add 3 to y, then subtract 4
from both x and w. Return the
product of the four numbers."""
# Codex generates:
t = y + 3
u = x - 4
v = z * w # ERROR: w not decremented
return v # ERROR: not the product of all fourBroader impacts: safety, alignment, and society
The paper dedicates significant attention to the risks of deploying code generation. The most nuanced concern is over-reliance: Codex can generate code that looks correct but contains subtle bugs. Novice programmers are especially vulnerable — they may not have the expertise to spot the errors, leading to a false sense of security.
An issue emerges: when the prompt contains buggy code, Codex tends to generate more buggy code rather than correcting the pattern. This worsens with scale — larger models are better at mimicking the style of the prompt, including its flaws. The model is optimized to predict what comes next in the training distribution, not to be helpful to the user. A highly capable but misaligned model could even produce obfuscated code that looks correct on careful inspection but does something undesirable.
The authors also flag security risks (Codex can generate vulnerable code), concerns (code comments can reflect stereotypes from the training data), economic impacts on the programming labor market, environmental costs of training large models, and legal questions about training on public code.
Key results at a glance
Legacy: the age of AI-powered programming
2020
GPT-3 shows rudimentary code ability
GPT-3 could generate simple programs from docstrings, but solved 0% of HumanEval problems. This hinted that larger models with code-specific training could do much better.
2021
Codex and GitHub Copilot launch
This paper introduced Codex, HumanEval, and the pass@k framework. A production version powered GitHub Copilot, the first widely adopted AI coding assistant.
2022
AlphaCode tackles competitive programming
DeepMind's AlphaCode generated millions of samples for competitive programming problems and filtered them, achieving human-competitor-level performance on Codeforces.
2023
StarCoder and open-source code models
The BigCode project released StarCoder, trained on The Stack, a responsibly sourced multi-language code dataset. Open-source code models became competitive with proprietary ones.
2023
SWE-bench evaluates real-world software engineering
SWE-bench moved beyond function synthesis to evaluate models on real GitHub issues requiring understanding of entire codebases — a natural successor to HumanEval's function-level evaluation.
2023
Phi-1 achieves strong results with small models
Microsoft's Phi-1 showed that high-quality training data (textbook-quality code) could enable a 1.3B parameter model to achieve 50.6% pass@1 on HumanEval — surpassing the original Codex-12B with 10× fewer parameters.
The Codex paper is a watershed moment. Before it, code generation was an academic curiosity with near-zero practical utility. After it, AI-assisted programming became an industry worth billions. HumanEval remains the most commonly cited benchmark for code generation, and the pass@k metric it introduced is now the standard way to evaluate stochastic code generators. Every model that followed — StarCoder, Code Llama, GPT-4, Claude — builds on the foundation that Codex laid.
CitationChen, Tworek, Jun, Yuan, Pinto, Kaplan, Edwards, Burda, Joseph, Brockman, et al.. Evaluating Large Language Models Trained on Code. arXiv, 2021.
Terms in this paper
- Code Generationتوليد الشيفرات
- Fine-Tuningالضبط الدقيق
- Autoregressive Modelالنموذج التوليدي التراجعي
- Benchmarkالمعيار المرجعي
- Nucleus Samplingمعاينة النواة الاحتمالية
- Temperatureالحرارة
- Functional Correctnessالصحة الوظيفية
- Docstringالوصف النصي
- Supervised Learningالتعلم الـمُوجّه (المصحوب ببيانات مرجعية)
- Scaling Lawsقوانين التوسعة
- Tokenوحدة لغوية (رمز)
- Tokenizationتجزئة النصوص
- Autoregressive Generationالتوليد الارتجاعي
- BLEU Scoreمعيار بلو الإحصائي لتقييم الترجمة
- Alignmentالمحاذاة