Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks

Abstract

Output correctness alone does not tell us whether a large language model actually reasons correctly about code. This work introduces CodeRQ-Bench, a benchmark for assessing LLM reasoning quality in coding tasks beyond output correctness, together with VERA, a two-stage evaluator for detecting flawed reasoning in model solutions. VERA achieves improvements of up to 0.26 in AUCROC and 0.21 in AUPRC over existing evaluation approaches, moving LLM evaluation for code toward transparency and reasoning-aware assessment.

Publication
Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)
Yuangang Li
Yuangang Li
PhD Student at UCI