Research Paper
Peer-ReviewedEvaluating LLMs for Code Generation in Real-World Software Engineering Tasks
An empirical study on the capabilities and limitations of LLMs in generating production-ready code for real-world software engineering scenarios.
Bishal Saha
ACM SIGSOFT Conference on the Foundations of Software Engineering (FSE)
pp. 1021–1033, 2022
Abstract
Benchmarks for code generation mostly measure isolated functions. We evaluate LLMs on tasks drawn from real repositories — changes that touch several files, respect existing conventions and must pass the project's own tests.
Results show a large gap between benchmark scores and real-world task success, and we catalogue the failure modes that explain it.
Cite This Paper
Saha, B. (2022). Evaluating LLMs for Code Generation in Real-World Software Engineering Tasks. ACM SIGSOFT Conference on the Foundations of Software Engineering (FSE), 1021–1033. https://iambishal.com/work/research/papers/evaluating-llms-code-generation