PDB-Wild / README.md
Cmsmr's picture
Add PDB-Wild
52e2478 verified
|
Raw History Blame Contribute Delete
4.09 kB
---
language:
- en
- code
license: mit
task_categories:
- text-generation
size_categories:
- n<1K
tags:
- code
- debugging
- benchmark
configs:
- config_name: default
data_files:
- split: test
path: data/test-*.parquet
---
# PDB-Wild: Precise Debugging Benchmarking — multi-line and repository-level bugs
📄 [Paper](https://arxiv.org/abs/2604.17338) &nbsp;·&nbsp;
💻 [Code](https://github.com/Bill1235813/PDB) &nbsp;·&nbsp;
🌐 [Project page](https://precise-debugging-benchmark.github.io/) &nbsp;·&nbsp;
🏆 [Leaderboard](https://precise-debugging-benchmark.github.io/leaderboard.html)
`PDB-Wild` is the **multi-line and repository-level bug set** of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (`gt_diff`) that encodes the minimal correct fix.
- **Source datasets:** [BigCodeBench](https://huggingface.co/datasets/bigcode/bigcodebench) + [LiveCodeBench](https://huggingface.co/datasets/livecodebench/execution) (the 256 examples of [PDB-Multi](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi)) and [SWE-smith](https://github.com/SWE-bench/SWE-smith) repositories (228 examples)
- **Sibling datasets:** [PDB-Single](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single) · [PDB-Single-Full](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single-Full) · [PDB-Multi](https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi)
## TL;DR
Unit tests reward brute-force regeneration equally with minimal targeted fixes. PDB instead evaluates debugging with edit-level **precision** (were unnecessary lines touched?) and bug-level **recall** (were all faults resolved?). On PDB-Wild, the precision gap persists for multi-line and repository-level bugs: models with high unit-test pass rates still over-edit.
## Statistics
- **Total examples:** 484
- **Per source dataset:**
- `bigcodebench`: 37
- `livecodebench`: 219
- `swebench` (SWE-smith repositories): 228 — 32 files from 6 repositories, bugs generated by Claude-Opus-4.7 with blocks of up to 30 lines
- **Bug count distribution:**
- `bug_count = 1`: 220
- `bug_count = 2`: 164
- `bug_count = 3`: 100
## Schema
| field | type | notes |
|---|---|---|
| `task_id` | string | unique identifier per buggy variant |
| `source_dataset` | string | provenance of the underlying program (`bigcodebench`, `livecodebench`, `swebench` = SWE-smith repository) |
| `source_model` | string | bug-generator model |
| `task_prompt` | string | natural-language description of the task / fix target |
| `gt_solution` | string | verified correct program (for SWE-smith: the full content of `target_file`) |
| `buggy_code` | string | program with injected bug(s) |
| `gt_diff` | string (JSON) | `{line_no: {type, original, modified}}` mapping — the fix |
| `bug_count` | int | number of independent bug blocks (range: {1, 2, 3}) |
| `bug_type`, `bug_subtype` | string \| null | Orthogonal Defect Classification label (populated for `bug_count == 1`) |
| `gt_length` | int | line count of `gt_solution` |
| `editable_lines`, `deletable_lines`, `frozen_lines` | int \| null | handler-derived line counts |
| `is_buggy` | bool \| null | `true` for single-bug examples, null for composed multi-bug examples |
| `repo`, `image_name`, `target_file` | string \| null | SWE-smith repository, Docker image and file path (SWE-smith examples only) |
SWE-smith examples are evaluated by applying the model's fix to the repository inside its SWE-smith Docker image and running the repository's test suite; see [`dataset/swesmith`](https://github.com/Bill1235813/PDB/tree/main/dataset/swesmith) in the code repository.
## Loading
```python
from datasets import load_dataset
ds = load_dataset("Precise-Debugging-Benchmarking/PDB-Wild", split="test")
example = ds[0]
print(example["buggy_code"])
print(example["gt_solution"])
```
`gt_diff` is a JSON-encoded string; decode with `json.loads(example["gt_diff"])`.
## License
MIT.