Datasets:
Download README.md from Precise-Debugging-Benchmarking/PDB-Wild: direct link, hf CLI and curl.
- Browser
- Download file 4.09 kB
-
https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Wild/resolve/main/README.md
- Command line
-
hf download hf://datasets/Precise-Debugging-Benchmarking/PDB-Wild/README.md
-
curl -L -o README.md https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Wild/resolve/main/README.md
language:
- en
- code
license: mit
task_categories:
- text-generation
size_categories:
- n<1K
tags:
- code
- debugging
- benchmark
configs:
- config_name: default
data_files:
- split: test
path: data/test-*.parquet
PDB-Wild: Precise Debugging Benchmarking — multi-line and repository-level bugs
📄 Paper · 💻 Code · 🌐 Project page · 🏆 Leaderboard
PDB-Wild is the multi-line and repository-level bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix.
- Source datasets: BigCodeBench + LiveCodeBench (the 256 examples of PDB-Multi) and SWE-smith repositories (228 examples)
- Sibling datasets: PDB-Single · PDB-Single-Full · PDB-Multi
TL;DR
Unit tests reward brute-force regeneration equally with minimal targeted fixes. PDB instead evaluates debugging with edit-level precision (were unnecessary lines touched?) and bug-level recall (were all faults resolved?). On PDB-Wild, the precision gap persists for multi-line and repository-level bugs: models with high unit-test pass rates still over-edit.
Statistics
- Total examples: 484
- Per source dataset:
bigcodebench: 37livecodebench: 219swebench(SWE-smith repositories): 228 — 32 files from 6 repositories, bugs generated by Claude-Opus-4.7 with blocks of up to 30 lines
- Bug count distribution:
bug_count = 1: 220bug_count = 2: 164bug_count = 3: 100
Schema
| field | type | notes |
|---|---|---|
task_id |
string | unique identifier per buggy variant |
source_dataset |
string | provenance of the underlying program (bigcodebench, livecodebench, swebench = SWE-smith repository) |
source_model |
string | bug-generator model |
task_prompt |
string | natural-language description of the task / fix target |
gt_solution |
string | verified correct program (for SWE-smith: the full content of target_file) |
buggy_code |
string | program with injected bug(s) |
gt_diff |
string (JSON) | {line_no: {type, original, modified}} mapping — the fix |
bug_count |
int | number of independent bug blocks (range: {1, 2, 3}) |
bug_type, bug_subtype |
string | null | Orthogonal Defect Classification label (populated for bug_count == 1) |
gt_length |
int | line count of gt_solution |
editable_lines, deletable_lines, frozen_lines |
int | null | handler-derived line counts |
is_buggy |
bool | null | true for single-bug examples, null for composed multi-bug examples |
repo, image_name, target_file |
string | null | SWE-smith repository, Docker image and file path (SWE-smith examples only) |
SWE-smith examples are evaluated by applying the model's fix to the repository inside its SWE-smith Docker image and running the repository's test suite; see dataset/swesmith in the code repository.
Loading
from datasets import load_dataset
ds = load_dataset("Precise-Debugging-Benchmarking/PDB-Wild", split="test")
example = ds[0]
print(example["buggy_code"])
print(example["gt_solution"])
gt_diff is a JSON-encoded string; decode with json.loads(example["gt_diff"]).
License
MIT.