--- language: - en - code license: mit task_categories: - text-generation size_categories: - n<1K tags: - code - debugging - benchmark configs: - config_name: default data_files: - split: test path: data/test-*.parquet --- # PDB-Wild: Precise Debugging Benchmarking โ€” multi-line and repository-level bugs ๐Ÿ“„ [Paper](https://arxiv.org/abs/2604.17338)  ยท  ๐Ÿ’ป [Code](https://github.com/Bill1235813/PDB)  ยท  ๐ŸŒ [Project page](https://precise-debugging-benchmark.github.io/)  ยท  ๐Ÿ† [Leaderboard](https://precise-debugging-benchmark.github.io/leaderboard.html) `PDB-Wild` is the **multi-line and repository-level bug set** of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (`gt_diff`) that encodes the minimal correct fix. - **Source datasets:** [BigCodeBench](https://huggingface.co/datasets/bigcode/bigcodebench) + [LiveCodeBench](https://huggingface.co/datasets/livecodebench/execution) (the 256 examples of [PDB-Multi](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi)) and [SWE-smith](https://github.com/SWE-bench/SWE-smith) repositories (228 examples) - **Sibling datasets:** [PDB-Single](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single) ยท [PDB-Single-Full](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single-Full) ยท [PDB-Multi](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi) ## TL;DR Unit tests reward brute-force regeneration equally with minimal targeted fixes. PDB instead evaluates debugging with edit-level **precision** (were unnecessary lines touched?) and bug-level **recall** (were all faults resolved?). On PDB-Wild, the precision gap persists for multi-line and repository-level bugs: models with high unit-test pass rates still over-edit. ## Statistics - **Total examples:** 484 - **Per source dataset:** - `bigcodebench`: 37 - `livecodebench`: 219 - `swebench` (SWE-smith repositories): 228 โ€” 32 files from 6 repositories, bugs generated by Claude-Opus-4.7 with blocks of up to 30 lines - **Bug count distribution:** - `bug_count = 1`: 220 - `bug_count = 2`: 164 - `bug_count = 3`: 100 ## Schema | field | type | notes | |---|---|---| | `task_id` | string | unique identifier per buggy variant | | `source_dataset` | string | provenance of the underlying program (`bigcodebench`, `livecodebench`, `swebench` = SWE-smith repository) | | `source_model` | string | bug-generator model | | `task_prompt` | string | natural-language description of the task / fix target | | `gt_solution` | string | verified correct program (for SWE-smith: the full content of `target_file`) | | `buggy_code` | string | program with injected bug(s) | | `gt_diff` | string (JSON) | `{line_no: {type, original, modified}}` mapping โ€” the fix | | `bug_count` | int | number of independent bug blocks (range: {1, 2, 3}) | | `bug_type`, `bug_subtype` | string \| null | Orthogonal Defect Classification label (populated for `bug_count == 1`) | | `gt_length` | int | line count of `gt_solution` | | `editable_lines`, `deletable_lines`, `frozen_lines` | int \| null | handler-derived line counts | | `is_buggy` | bool \| null | `true` for single-bug examples, null for composed multi-bug examples | | `repo`, `image_name`, `target_file` | string \| null | SWE-smith repository, Docker image and file path (SWE-smith examples only) | SWE-smith examples are evaluated by applying the model's fix to the repository inside its SWE-smith Docker image and running the repository's test suite; see [`dataset/swesmith`](https://github.com/Bill1235813/PDB/tree/main/dataset/swesmith) in the code repository. ## Loading ```python from datasets import load_dataset ds = load_dataset("Precise-Debugging-Benchmarking/PDB-Wild", split="test") example = ds[0] print(example["buggy_code"]) print(example["gt_solution"]) ``` `gt_diff` is a JSON-encoded string; decode with `json.loads(example["gt_diff"])`. ## License MIT.