PDB-Wild / README.md
Cmsmr's picture
Add PDB-Wild
52e2478 verified
|
Raw History Blame Contribute Delete
4.09 kB
metadata
language:
  - en
  - code
license: mit
task_categories:
  - text-generation
size_categories:
  - n<1K
tags:
  - code
  - debugging
  - benchmark
configs:
  - config_name: default
    data_files:
      - split: test
        path: data/test-*.parquet

PDB-Wild: Precise Debugging Benchmarking — multi-line and repository-level bugs

📄 Paper  ·  💻 Code  ·  🌐 Project page  ·  🏆 Leaderboard

PDB-Wild is the multi-line and repository-level bug set of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix.

TL;DR

Unit tests reward brute-force regeneration equally with minimal targeted fixes. PDB instead evaluates debugging with edit-level precision (were unnecessary lines touched?) and bug-level recall (were all faults resolved?). On PDB-Wild, the precision gap persists for multi-line and repository-level bugs: models with high unit-test pass rates still over-edit.

Statistics

  • Total examples: 484
  • Per source dataset:
    • bigcodebench: 37
    • livecodebench: 219
    • swebench (SWE-smith repositories): 228 — 32 files from 6 repositories, bugs generated by Claude-Opus-4.7 with blocks of up to 30 lines
  • Bug count distribution:
    • bug_count = 1: 220
    • bug_count = 2: 164
    • bug_count = 3: 100

Schema

field type notes
task_id string unique identifier per buggy variant
source_dataset string provenance of the underlying program (bigcodebench, livecodebench, swebench = SWE-smith repository)
source_model string bug-generator model
task_prompt string natural-language description of the task / fix target
gt_solution string verified correct program (for SWE-smith: the full content of target_file)
buggy_code string program with injected bug(s)
gt_diff string (JSON) {line_no: {type, original, modified}} mapping — the fix
bug_count int number of independent bug blocks (range: {1, 2, 3})
bug_type, bug_subtype string | null Orthogonal Defect Classification label (populated for bug_count == 1)
gt_length int line count of gt_solution
editable_lines, deletable_lines, frozen_lines int | null handler-derived line counts
is_buggy bool | null true for single-bug examples, null for composed multi-bug examples
repo, image_name, target_file string | null SWE-smith repository, Docker image and file path (SWE-smith examples only)

SWE-smith examples are evaluated by applying the model's fix to the repository inside its SWE-smith Docker image and running the repository's test suite; see dataset/swesmith in the code repository.

Loading

from datasets import load_dataset
ds = load_dataset("Precise-Debugging-Benchmarking/PDB-Wild", split="test")
example = ds[0]
print(example["buggy_code"])
print(example["gt_solution"])

gt_diff is a JSON-encoded string; decode with json.loads(example["gt_diff"]).

License

MIT.