IamBishalIamBishal

Research Paper

Peer-Reviewed

Evaluating LLMs for Code Generation in Real-World Software Engineering Tasks

An empirical study on the capabilities and limitations of LLMs in generating production-ready code for real-world software engineering scenarios.

Bishal Saha

ACM SIGSOFT Conference on the Foundations of Software Engineering (FSE)

pp. 1021–1033, 2022

Abstract

Benchmarks for code generation mostly measure isolated functions. We evaluate LLMs on tasks drawn from real repositories — changes that touch several files, respect existing conventions and must pass the project's own tests.

Results show a large gap between benchmark scores and real-world task success, and we catalogue the failure modes that explain it.

Cite This Paper

Saha, B. (2022). Evaluating LLMs for Code Generation in Real-World Software Engineering Tasks. ACM SIGSOFT Conference on the Foundations of Software Engineering (FSE), 1021–1033. https://iambishal.com/work/research/papers/evaluating-llms-code-generation
Research is not just about new knowledge, but about creating real impact.
— Bishal Saha