AI & ML interests

None defined yet.

Recent Activity

Organization Card

Precise Debugging Benchmarking (PDB)

⭐ Neurips'26: Evaluations & Datasets

šŸ“„ Paper  Ā·  šŸ’» Code  Ā·  🌐 Project page  Ā·  šŸ† Leaderboard

PDB is an automatic pipeline that turns any coding dataset into a debugging benchmark with fine-grained metrics. Beyond binary unit-test scores, PDB evaluates a debugger with edit-level precision (did the model touch only the lines it had to?) and bug-level recall (did it fix every fault?). This rewards targeted fixes and penalizes the regeneration behavior frontier LLMs often fall back on.

Frontier models like GPT-5.1-Codex and DeepSeek-V3.2-Thinking top unit-test leaderboards (>76%) but score at or below 45% on precision: they pass tests by rewriting, not repairing. PDB makes that gap measurable.

Released datasets

Dataset Size Bug granularity Notes
PDB-Single 5,751 single line main evaluation set: tasks not easily solved by 7+ of 9 reference models
PDB-Single-Full 7,589 single line full initial pool before easy-case filtering
PDB-Wild 484 multi-line / repository-level PDB-Multi (256) + SWE-smith repository bugs (228)
PDB-Multi 256 2–4 line blocks BigCodeBench/LiveCodeBench part of PDB-Wild; programs with ≄35 LOC
PDB-Results – – model outputs and scores

PDB-Single, PDB-Single-Full and PDB-Multi are derived from BigCodeBench and LiveCodeBench; PDB-Wild adds SWE-smith repositories. All are built by the PDB pipeline and evaluated with precision / recall / unit-test pass rate.

Citation

@article{chai2026pdb,
  title={Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?},
  author={Chai, Miaosen and Zhu, Wang Bill and Wang, Shangshang and Liu, Yejia and Bian, Song and Dong, Honghua and Neiswanger, Willie and Jia, Robin},
  journal={arXiv preprint arXiv:2604.17338},
  year={2026}
}

Contact

Questions / submissions: wangzhu@usc.edu, miaosenc@usc.edu.

models 0

None public yet