Datasets:
|
Download README.md from Precise-Debugging-Benchmarking/PDB-Wild: direct link, hf CLI and curl.
- Browser
- Download file 4.09 kB
-
https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Wild/resolve/main/README.md
- Command line
-
hf download hf://datasets/Precise-Debugging-Benchmarking/PDB-Wild/README.md
-
curl -L -o README.md https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Wild/resolve/main/README.md
4.09 kB
| language: | |
| - en | |
| - code | |
| license: mit | |
| task_categories: | |
| - text-generation | |
| size_categories: | |
| - n<1K | |
| tags: | |
| - code | |
| - debugging | |
| - benchmark | |
| configs: | |
| - config_name: default | |
| data_files: | |
| - split: test | |
| path: data/test-*.parquet | |
| # PDB-Wild: Precise Debugging Benchmarking — multi-line and repository-level bugs | |
| 📄 [Paper](https://arxiv.org/abs/2604.17338) · | |
| 💻 [Code](https://github.com/Bill1235813/PDB) · | |
| 🌐 [Project page](https://precise-debugging-benchmark.github.io/) · | |
| 🏆 [Leaderboard](https://precise-debugging-benchmark.github.io/leaderboard.html) | |
| `PDB-Wild` is the **multi-line and repository-level bug set** of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (`gt_diff`) that encodes the minimal correct fix. | |
| - **Source datasets:** [BigCodeBench](https://huggingface.co/datasets/bigcode/bigcodebench) + [LiveCodeBench](https://huggingface.co/datasets/livecodebench/execution) (the 256 examples of [PDB-Multi](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi)) and [SWE-smith](https://github.com/SWE-bench/SWE-smith) repositories (228 examples) | |
| - **Sibling datasets:** [PDB-Single](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single) · [PDB-Single-Full](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single-Full) · [PDB-Multi](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi) | |
| ## TL;DR | |
| Unit tests reward brute-force regeneration equally with minimal targeted fixes. PDB instead evaluates debugging with edit-level **precision** (were unnecessary lines touched?) and bug-level **recall** (were all faults resolved?). On PDB-Wild, the precision gap persists for multi-line and repository-level bugs: models with high unit-test pass rates still over-edit. | |
| ## Statistics | |
| - **Total examples:** 484 | |
| - **Per source dataset:** | |
| - `bigcodebench`: 37 | |
| - `livecodebench`: 219 | |
| - `swebench` (SWE-smith repositories): 228 — 32 files from 6 repositories, bugs generated by Claude-Opus-4.7 with blocks of up to 30 lines | |
| - **Bug count distribution:** | |
| - `bug_count = 1`: 220 | |
| - `bug_count = 2`: 164 | |
| - `bug_count = 3`: 100 | |
| ## Schema | |
| | field | type | notes | | |
| |---|---|---| | |
| | `task_id` | string | unique identifier per buggy variant | | |
| | `source_dataset` | string | provenance of the underlying program (`bigcodebench`, `livecodebench`, `swebench` = SWE-smith repository) | | |
| | `source_model` | string | bug-generator model | | |
| | `task_prompt` | string | natural-language description of the task / fix target | | |
| | `gt_solution` | string | verified correct program (for SWE-smith: the full content of `target_file`) | | |
| | `buggy_code` | string | program with injected bug(s) | | |
| | `gt_diff` | string (JSON) | `{line_no: {type, original, modified}}` mapping — the fix | | |
| | `bug_count` | int | number of independent bug blocks (range: {1, 2, 3}) | | |
| | `bug_type`, `bug_subtype` | string \| null | Orthogonal Defect Classification label (populated for `bug_count == 1`) | | |
| | `gt_length` | int | line count of `gt_solution` | | |
| | `editable_lines`, `deletable_lines`, `frozen_lines` | int \| null | handler-derived line counts | | |
| | `is_buggy` | bool \| null | `true` for single-bug examples, null for composed multi-bug examples | | |
| | `repo`, `image_name`, `target_file` | string \| null | SWE-smith repository, Docker image and file path (SWE-smith examples only) | | |
| SWE-smith examples are evaluated by applying the model's fix to the repository inside its SWE-smith Docker image and running the repository's test suite; see [`dataset/swesmith`](https://github.com/Bill1235813/PDB/tree/main/dataset/swesmith) in the code repository. | |
| ## Loading | |
| ```python | |
| from datasets import load_dataset | |
| ds = load_dataset("Precise-Debugging-Benchmarking/PDB-Wild", split="test") | |
| example = ds[0] | |
| print(example["buggy_code"]) | |
| print(example["gt_solution"]) | |
| ``` | |
| `gt_diff` is a JSON-encoded string; decode with `json.loads(example["gt_diff"])`. | |
| ## License | |
| MIT. | |