Dataset Viewer
Auto-converted to Parquet Duplicate
prompt_id
stringlengths
20
20
task_id
stringlengths
33
33
dataset_index
int64
0
29.5k
skip_bucket
stringclasses
1 value
try
int64
1
3
harness
stringclasses
3 values
model
stringclasses
2 values
prompt
stringlengths
43
1.77k
docker_image
stringlengths
52
52
cwd
stringclasses
69 values
difficulty
stringclasses
3 values
category
stringclasses
43 values
n_turns
int64
0
262
submitted
bool
2 classes
reasons
stringlengths
39
1.32k
reward
stringclasses
9 values
test_sh_returncode
int64
0
137
test_sh_tail
stringlengths
0
2k
tests_total
int64
1
30
tests_passed
int64
0
28
tests_failed
int64
0
23
tests_pass_pct
float64
0
100
messages
listlengths
0
548
turns
listlengths
1
262
tools
stringclasses
4 values
env_cost
float64
0
0.07
token_cost
float64
0
0.37
prompt_hint
stringclasses
1 value
row_uuid
stringlengths
36
36
0003a547c2e9fe94d324
rts_task_76b7d459b0d484a82d60aaf4
13,446
failed_task
1
hermes
deepseek/deepseek-v4-flash
Connect to the local PostgreSQL instance at localhost:5432 (database postgres, trust authentication). The `person` table currently lives in the `public` schema. Reorganise the database according to the specification in `/app/desired_objects.sql` (which defines a view `v_person` and a function `s_person`). Use the compl...
registry-b.deepgrove.ai/rts:76b7d459b0d484a82d60aaf4
/app
medium
database-operations
25
true
["tests failed: {\"reward\": \"0\", \"test_sh_returncode\": 0, \"test_sh_tail\": \"\\n\\nconn = <connection object at 0x75b3d8bf3c40; dsn: 'postgresql://postgres@localhost:5432/postgres', closed: 0>\\n\\n def test_function_fails_without_schema_qualification(conn):\\n with conn.cursor() as cur:\\n c...
0
0
conn = <connection object at 0x75b3d8bf3c40; dsn: 'postgresql://postgres@localhost:5432/postgres', closed: 0> def test_function_fails_without_schema_qualification(conn): with conn.cursor() as cur: cur.execute("SET search_path TO public;") with pytest.raises(psycopg2.errors.Undefin...
14
5
9
35.71
[ { "role": "system", "content": "You are Hermes Agent, an intelligent AI assistant created by Nous Research. You are helpful, knowledgeable, and direct. You assist users with a wide range of tasks including answering questions, writing and editing code, analyzing information, creative work, and executing act...
[ { "turn": 1, "reasoning": "Let me start by carefully enumerating every requirement from the task, then systematically work through them.\n\n## Requirements\n\n1. \"Connect to the local PostgreSQL instance at localhost:5432 (database postgres, trust authentication).\"\n2. \"The `person` table currently lives...
[{"name": "clarify", "description": "Ask the user a question when you need clarification, feedback, or a decision before proceeding. Supports three modes:\n\n1. **Single-select multiple choice** — provide up to 4 choices. The user picks one or types their own answer via a 5th 'Other' option.\n2. **Multi-select multiple...
0.007907
0.009872
Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERIFIED until a tool result proves it. Never assert that something exists, works, or is correct because you wrote it, intended it, or because it seems obvious -- only because a command you ran returned output that shows it. ...
e941d358-b358-40c8-a40f-7cda1be33c91
0003cbc20fd0f8218bc8
rts_task_51f2351df9b6b2a1035765da
9,173
failed_task
1
opencode
deepseek/deepseek-v4-flash
"Reconcile /etc/fstab to match the active mount specification.\nThe specification directory is not h(...TRUNCATED)
registry-b.deepgrove.ai/rts:51f2351df9b6b2a1035765da
/app
easy
system-administration
15
true
"[\"tests failed: {\\\"reward\\\": \\\"0\\\", \\\"test_sh_returncode\\\": 0, \\\"test_sh_tail\\\": \(...TRUNCATED)
0
0
"e = _parse_report_comparison_entries(report)\n> assert parsed_table, \"No comparison table en(...TRUNCATED)
4
1
3
25
[{"role":"system","content":"You are powered by the model named deepseek/deepseek-v4-flash. The exac(...TRUNCATED)
[{"turn":1,"reasoning":"Let me start by carefully enumerating every requirement from the task, then (...TRUNCATED)
"[{\"type\": \"function\", \"function\": {\"name\": \"bash\", \"description\": \"Executes a given ba(...TRUNCATED)
0.004737
0.0085
"Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERI(...TRUNCATED)
d042f45a-22e7-47fc-991b-993016cafaa6
000988dce5f1a79bad06
rts_task_a1f7a50adbe965e42492898f
18,370
failed_task
1
hermes
deepseek/deepseek-v4-flash
"In the `/app` directory, a Judge0 API response is stored at `data/judge0_response.json` containing (...TRUNCATED)
registry-b.deepgrove.ai/rts:a1f7a50adbe965e42492898f
/app
easy
scripting-automation
10
true
"[\"tests failed: {\\\"reward\\\": \\\"0\\\", \\\"test_sh_returncode\\\": 0, \\\"test_sh_tail\\\": \(...TRUNCATED)
0
0
"abcdef' for c in stripped), 'Integrity hash contains non-hex characters'\n \n # --- repai(...TRUNCATED)
4
1
3
25
[{"role":"system","content":"You are Hermes Agent, an intelligent AI assistant created by Nous Resea(...TRUNCATED)
[{"turn":1,"reasoning":"Let me start by reading the necessary files to understand the task requireme(...TRUNCATED)
"[{\"name\": \"clarify\", \"description\": \"Ask the user a question when you need clarification, fe(...TRUNCATED)
0.004177
0.005757
"Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERI(...TRUNCATED)
ce018fa6-55f2-4be7-888b-c55fe74d3741
0009b0f843e653e94c73
rts_task_068a4e6d69bfcbda4228b6d4
720
failed_task
1
pi
deepseek/deepseek-v4-flash
"Edit /etc/fstab to add entries for every storage device listed in /etc/device-inventory.json, so th(...TRUNCATED)
registry-b.deepgrove.ai/rts:068a4e6d69bfcbda4228b6d4
/app
easy
system-administration
8
true
"[\"tests failed: {\\\"reward\\\": \\\"0\\\", \\\"test_sh_returncode\\\": 0, \\\"test_sh_tail\\\": \(...TRUNCATED)
0
0
"ed 3 UUID entries, found {len(uuid_entries)}\"\n \n data_entry = None\n musik_entr(...TRUNCATED)
4
3
1
75
[{"role":"system","content":"You are an expert coding assistant operating inside pi, a coding agent (...TRUNCATED)
[{"turn":1,"reasoning":"Let me start by understanding the task requirements thoroughly:\n\n1. Read /(...TRUNCATED)
"[{\"type\": \"function\", \"function\": {\"name\": \"bash\", \"description\": \"Execute a bash comm(...TRUNCATED)
0.00317
0.001697
"Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERI(...TRUNCATED)
2864b29d-e3d3-41eb-80d2-ca4a22d0f35a
0009b0f843e653e94c73
rts_task_068a4e6d69bfcbda4228b6d4
720
failed_task
2
pi
deepseek/deepseek-v4-flash
"Edit /etc/fstab to add entries for every storage device listed in /etc/device-inventory.json, so th(...TRUNCATED)
registry-b.deepgrove.ai/rts:068a4e6d69bfcbda4228b6d4
/app
easy
system-administration
9
true
"[\"tests failed: {\\\"reward\\\": \\\"0\\\", \\\"test_sh_returncode\\\": 0, \\\"test_sh_tail\\\": \(...TRUNCATED)
0
0
"ed 3 UUID entries, found {len(uuid_entries)}\"\n \n data_entry = None\n musik_entr(...TRUNCATED)
4
3
1
75
[{"role":"system","content":"You are an expert coding assistant operating inside pi, a coding agent (...TRUNCATED)
[{"turn":1,"reasoning":"Let me start by enumerating all the requirements carefully, then explore the(...TRUNCATED)
"[{\"type\": \"function\", \"function\": {\"name\": \"bash\", \"description\": \"Execute a bash comm(...TRUNCATED)
0.002968
0.001593
"Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERI(...TRUNCATED)
d1d55c35-1f50-494b-970b-19f2c2cdfa87
000cfb2405171959a4c7
rts_task_8d6154dc01944a213f8772ae
15,995
failed_task
1
opencode
deepseek/deepseek-v4-flash
"In `/app`, use the `extra-object` npm package to process the nested object in `/app/data/input.json(...TRUNCATED)
registry-b.deepgrove.ai/rts:8d6154dc01944a213f8772ae
/app
easy
scripting-automation
17
true
"[\"tests failed: {\\\"reward\\\": \\\"0\\\", \\\"test_sh_returncode\\\": 0, \\\"test_sh_tail\\\": \(...TRUNCATED)
0
0
"_and_correct ________________________\n\n def test_r3_file_exists_and_correct():\n path =(...TRUNCATED)
5
1
4
20
[{"role":"system","content":"You are powered by the model named deepseek/deepseek-v4-flash. The exac(...TRUNCATED)
[{"turn":1,"reasoning":"The user wants me to be very careful and methodical. Let me start by enumera(...TRUNCATED)
"[{\"type\": \"function\", \"function\": {\"name\": \"bash\", \"description\": \"Executes a given ba(...TRUNCATED)
0.004953
0.025261
"Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERI(...TRUNCATED)
cd55041a-64e9-4dbe-88d5-9fc3fa65a643
00166aa2fafaccaec9e3
rts_task_2231baeae30fac30f0bd6027
3,765
failed_task
1
opencode
deepseek/deepseek-v4-flash
"Upgrade the shopping cart CLI at `/app/shopping_cart.py` to support a `--profile` argument.\nWhen r(...TRUNCATED)
registry-b.deepgrove.ai/rts:2231baeae30fac30f0bd6027
/app
easy
scripting-automation
20
true
"[\"tests failed: {\\\"reward\\\": \\\"0\\\", \\\"test_sh_returncode\\\": 0, \\\"test_sh_tail\\\": \(...TRUNCATED)
0
0
" cmd in required_commands:\n> assert cmd in content, f\"Missing expected command '{cmd}' (...TRUNCATED)
4
1
3
25
[{"role":"system","content":"You are powered by the model named deepseek/deepseek-v4-flash. The exac(...TRUNCATED)
[{"turn":1,"reasoning":"Let me start by understanding the task. I need to:\n\n1. First, read the exi(...TRUNCATED)
"[{\"type\": \"function\", \"function\": {\"name\": \"bash\", \"description\": \"Executes a given ba(...TRUNCATED)
0.00652
0.009712
"Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERI(...TRUNCATED)
bb6bb937-94cb-41ab-8fcd-cc65ee513ea1
0016bbf50f93284c474e
rts_task_cdc8c7d72495a706d9bd1fa4
23,464
failed_task
1
hermes
deepseek/deepseek-v4-flash
"In /app, repair the KSQL validation pipeline.\nFix validate.sh to count only standalone semicolons (...TRUNCATED)
registry-b.deepgrove.ai/rts:cdc8c7d72495a706d9bd1fa4
/app
easy
database-operations
33
true
"[\"tests failed: {\\\"reward\\\": \\\"0\\\", \\\"test_sh_returncode\\\": 1, \\\"test_sh_tail\\\": \(...TRUNCATED)
0
1
"verify_out.read_text()\n assert \"Independent count:\" in vtext, \"verify_output.txt missing(...TRUNCATED)
4
2
2
50
[{"role":"system","content":"You are Hermes Agent, an intelligent AI assistant created by Nous Resea(...TRUNCATED)
[{"turn":1,"reasoning":"Let me start by enumerating every requirement from the task statement as a n(...TRUNCATED)
"[{\"name\": \"clarify\", \"description\": \"Ask the user a question when you need clarification, fe(...TRUNCATED)
0.006958
0.017836
"Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERI(...TRUNCATED)
1103681c-4bb4-4f43-91f1-3d0d62199e6b
0016e86ea9749540863a
rts_task_851dff8b02b5c44123310651
15,153
failed_task
1
opencode
deepseek/deepseek-v4-flash
"The impatientanimedownloader tool is available at `/app/impatientanimedownloader/`.\nRefer to its l(...TRUNCATED)
registry-b.deepgrove.ai/rts:851dff8b02b5c44123310651
/app
easy
scripting-automation
22
true
"[\"tests failed: {\\\"reward\\\": \\\"0\\\", \\\"test_sh_returncode\\\": 0, \\\"test_sh_tail\\\": \(...TRUNCATED)
0
0
"form linux -- Python 3.9.24, pytest-8.4.1, pluggy-1.6.0\nrootdir: /tests\nplugins: json-ctrf-0.3.5\(...TRUNCATED)
5
4
1
80
[{"role":"system","content":"You are powered by the model named deepseek/deepseek-v4-flash. The exac(...TRUNCATED)
[{"turn":1,"reasoning":"Let me start by enumerating all the requirements from the task, then systema(...TRUNCATED)
"[{\"type\": \"function\", \"function\": {\"name\": \"bash\", \"description\": \"Executes a given ba(...TRUNCATED)
0.006219
0.020111
"Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERI(...TRUNCATED)
c870ab02-45d2-4feb-aaa8-c0ffc84594e3
0019a648390fb6d06c7c
rts_task_c36f6bea1af4bade36f6ff98
22,333
failed_task
1
pi
deepseek/deepseek-v4-flash
"Use the scripts and guide in /opt: gen_results.py (buggy generator), check.py (validator), and evid(...TRUNCATED)
registry-b.deepgrove.ai/rts:c36f6bea1af4bade36f6ff98
/app
easy
scripting-automation
30
true
"[\"tests failed: {\\\"reward\\\": \\\"0\\\", \\\"test_sh_returncode\\\": 0, \\\"test_sh_tail\\\": \(...TRUNCATED)
0
0
"ne.strip())\n ]\n fail_lines = [\n line.strip() for line in report.splitli(...TRUNCATED)
14
12
2
85.71
[{"role":"system","content":"You are an expert coding assistant operating inside pi, a coding agent (...TRUNCATED)
[{"turn":1,"reasoning":"Let me break down this task carefully:\n\n1. First, I need to read the files(...TRUNCATED)
"[{\"type\": \"function\", \"function\": {\"name\": \"bash\", \"description\": \"Execute a bash comm(...TRUNCATED)
0.004868
0.006925
"Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERI(...TRUNCATED)
f7c2a16c-39a6-418c-9393-a3c038f2ed4e
End of preview. Expand in Data Studio

RTS × deepseek-v4-flash — REJECTED traces

25,210 rejected agentic attempts from deepseek/deepseek-v4-flash on the Recursive Task Synthesis corpus. The accepted counterpart is RedMod/recursive-v4-flash.

Every row is an attempt that ran and was graded, and the task's own tests returned 0. Harness failures (failed_pre_checks — the model never got a scored shot) are deliberately EXCLUDED, so nothing here misattributes a plumbing bug to the model.

Why these are worth keeping

They are near-misses, not garbage. Measured over a 1,544-attempt sample:

signal rate
garbled / non-Latin tokens 0.8%
empty assistant text 0.2%
looping (same tool call ≥5×) 10.9%
median duplicate-tool-call fraction 0.06
median turns 26

The model typically does the substantive work and then misses part of the deliverable spec. Categorising 25,007 rejected attempts by what the grader objected to:

failure mode share
a required artifact was absent 60.9%
shortcut / forbidden modification detected 19.0%
placeholder (TODO, stub) left in output 13.8%
artifact present but wrong shape 3.1%
value mismatch 2.1%

Per-check fail rates show the same gradient — surface compliance is easy, exactness is not:

check fail rate
required_evidence (did you leave something) 32.7%
intermediate_artifact 56.3%
no_shortcut 72.2%
final_semantics 79.0%

Partial credit — the binary reward hides most of the signal

reward is 0 for every row here, but the graders run several independent test functions per task and these attempts pass most of them. Across 194,615 individual tests in 24,604 rejected attempts:

individual tests run 194,615
tests PASSED 113,057 (58.1%)
tests FAILED 81,558 (41.9%)
median per-attempt pass rate 50.0%
failed exactly ONE test 4,574 attempts (18.6%)

Most common shapes: 2/4 (4,151), 1/4 (4,139), 3/4 (2,848), 0/4 (2,096).

So a ~12% task-level accept rate coexists with a ~58% test-level pass rate, and nearly one rejected attempt in five missed by a single check. Use the tests_* columns to filter for near-misses rather than treating every row as an equal failure.

Schema

One row per attempt (a task retried twice contributes two rows; dedupe on prompt_id for task-level counts).

column
prompt_id, task_id, dataset_index identity; task_id joins to the corpus
skip_bucket always failed_task here
try, harness, model which attempt, which harness ran it
prompt, docker_image, cwd, difficulty, category task metadata
messages the full trajectory as a native list of structs{role, content, reasoning_content, tool_calls[], tool_call_id, name}; tool_calls is itself a list of {id, type, name, arguments} (arguments stays a JSON string, since its shape is per-tool)
turns native list of {turn, reasoning, n_tool_calls}
tools the harness tool schema (JSON string)
n_turns, submitted, reasons episode outcome
reward, test_sh_returncode, test_sh_tail grader output — test_sh_tail names the failing checks
tests_total, tests_passed, tests_failed, tests_pass_pct per-TEST partial credit parsed from the grader output (null when unparseable)
env_cost, token_cost real spend for this attempt
prompt_hint scaffolding appended to the prompt, stripped from messages

Run configuration

max_turns 100, max_tries 1, temperature 0.5, top_p 0.95, reasoning_effort high, harnesses hermes / opencode / pi (no terminus2), Daytona sandboxes, verifier agentic (the task's own tests/test.sh).

prompt_hint is set on these rows (the "hint-v3" self-critical + enumerate + verify-as-stated instruction). It is stripped from messages so the trajectory reads as the unmodified task — check the column before mixing with unhinted data.

Caveats

  • reward is MISSING on ~3% of rows: test.sh never wrote /logs/verifier/reward.txt (usually a grader that died early). Treated as a failure, but it is not a clean 0.
  • ~4% of the underlying corpus is unsolvable regardless of model (the instruction cites rewrite_*.json docs the image never ships); those tasks were filtered out of this run's dataset, but related noise may remain.
Downloads last month
219