prompt_id stringlengths 20 20 | task_id stringlengths 33 33 | dataset_index int64 0 29.5k | skip_bucket stringclasses 1
value | try int64 1 3 | harness stringclasses 3
values | model stringclasses 2
values | prompt stringlengths 43 1.77k | docker_image stringlengths 52 52 | cwd stringclasses 69
values | difficulty stringclasses 3
values | category stringclasses 43
values | n_turns int64 0 262 | submitted bool 2
classes | reasons stringlengths 39 1.32k | reward stringclasses 9
values | test_sh_returncode int64 0 137 ⌀ | test_sh_tail stringlengths 0 2k ⌀ | tests_total int64 1 30 ⌀ | tests_passed int64 0 28 ⌀ | tests_failed int64 0 23 ⌀ | tests_pass_pct float64 0 100 ⌀ | messages listlengths 0 548 | turns listlengths 1 262 ⌀ | tools stringclasses 4
values | env_cost float64 0 0.07 | token_cost float64 0 0.37 | prompt_hint stringclasses 1
value | row_uuid stringlengths 36 36 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
0003a547c2e9fe94d324 | rts_task_76b7d459b0d484a82d60aaf4 | 13,446 | failed_task | 1 | hermes | deepseek/deepseek-v4-flash | Connect to the local PostgreSQL instance at localhost:5432 (database postgres, trust authentication).
The `person` table currently lives in the `public` schema.
Reorganise the database according to the specification in `/app/desired_objects.sql` (which defines a view `v_person` and a function `s_person`).
Use the compl... | registry-b.deepgrove.ai/rts:76b7d459b0d484a82d60aaf4 | /app | medium | database-operations | 25 | true | ["tests failed: {\"reward\": \"0\", \"test_sh_returncode\": 0, \"test_sh_tail\": \"\\n\\nconn = <connection object at 0x75b3d8bf3c40; dsn: 'postgresql://postgres@localhost:5432/postgres', closed: 0>\\n\\n def test_function_fails_without_schema_qualification(conn):\\n with conn.cursor() as cur:\\n c... | 0 | 0 |
conn = <connection object at 0x75b3d8bf3c40; dsn: 'postgresql://postgres@localhost:5432/postgres', closed: 0>
def test_function_fails_without_schema_qualification(conn):
with conn.cursor() as cur:
cur.execute("SET search_path TO public;")
with pytest.raises(psycopg2.errors.Undefin... | 14 | 5 | 9 | 35.71 | [
{
"role": "system",
"content": "You are Hermes Agent, an intelligent AI assistant created by Nous Research. You are helpful, knowledgeable, and direct. You assist users with a wide range of tasks including answering questions, writing and editing code, analyzing information, creative work, and executing act... | [
{
"turn": 1,
"reasoning": "Let me start by carefully enumerating every requirement from the task, then systematically work through them.\n\n## Requirements\n\n1. \"Connect to the local PostgreSQL instance at localhost:5432 (database postgres, trust authentication).\"\n2. \"The `person` table currently lives... | [{"name": "clarify", "description": "Ask the user a question when you need clarification, feedback, or a decision before proceeding. Supports three modes:\n\n1. **Single-select multiple choice** — provide up to 4 choices. The user picks one or types their own answer via a 5th 'Other' option.\n2. **Multi-select multiple... | 0.007907 | 0.009872 | Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERIFIED until a tool result proves it. Never assert that something exists, works, or is correct because you wrote it, intended it, or because it seems obvious -- only because a command you ran returned output that shows it. ... | e941d358-b358-40c8-a40f-7cda1be33c91 |
0003cbc20fd0f8218bc8 | rts_task_51f2351df9b6b2a1035765da | 9,173 | failed_task | 1 | opencode | deepseek/deepseek-v4-flash | "Reconcile /etc/fstab to match the active mount specification.\nThe specification directory is not h(...TRUNCATED) | registry-b.deepgrove.ai/rts:51f2351df9b6b2a1035765da | /app | easy | system-administration | 15 | true | "[\"tests failed: {\\\"reward\\\": \\\"0\\\", \\\"test_sh_returncode\\\": 0, \\\"test_sh_tail\\\": \(...TRUNCATED) | 0 | 0 | "e = _parse_report_comparison_entries(report)\n> assert parsed_table, \"No comparison table en(...TRUNCATED) | 4 | 1 | 3 | 25 | [{"role":"system","content":"You are powered by the model named deepseek/deepseek-v4-flash. The exac(...TRUNCATED) | [{"turn":1,"reasoning":"Let me start by carefully enumerating every requirement from the task, then (...TRUNCATED) | "[{\"type\": \"function\", \"function\": {\"name\": \"bash\", \"description\": \"Executes a given ba(...TRUNCATED) | 0.004737 | 0.0085 | "Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERI(...TRUNCATED) | d042f45a-22e7-47fc-991b-993016cafaa6 |
000988dce5f1a79bad06 | rts_task_a1f7a50adbe965e42492898f | 18,370 | failed_task | 1 | hermes | deepseek/deepseek-v4-flash | "In the `/app` directory, a Judge0 API response is stored at `data/judge0_response.json` containing (...TRUNCATED) | registry-b.deepgrove.ai/rts:a1f7a50adbe965e42492898f | /app | easy | scripting-automation | 10 | true | "[\"tests failed: {\\\"reward\\\": \\\"0\\\", \\\"test_sh_returncode\\\": 0, \\\"test_sh_tail\\\": \(...TRUNCATED) | 0 | 0 | "abcdef' for c in stripped), 'Integrity hash contains non-hex characters'\n \n # --- repai(...TRUNCATED) | 4 | 1 | 3 | 25 | [{"role":"system","content":"You are Hermes Agent, an intelligent AI assistant created by Nous Resea(...TRUNCATED) | [{"turn":1,"reasoning":"Let me start by reading the necessary files to understand the task requireme(...TRUNCATED) | "[{\"name\": \"clarify\", \"description\": \"Ask the user a question when you need clarification, fe(...TRUNCATED) | 0.004177 | 0.005757 | "Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERI(...TRUNCATED) | ce018fa6-55f2-4be7-888b-c55fe74d3741 |
0009b0f843e653e94c73 | rts_task_068a4e6d69bfcbda4228b6d4 | 720 | failed_task | 1 | pi | deepseek/deepseek-v4-flash | "Edit /etc/fstab to add entries for every storage device listed in /etc/device-inventory.json, so th(...TRUNCATED) | registry-b.deepgrove.ai/rts:068a4e6d69bfcbda4228b6d4 | /app | easy | system-administration | 8 | true | "[\"tests failed: {\\\"reward\\\": \\\"0\\\", \\\"test_sh_returncode\\\": 0, \\\"test_sh_tail\\\": \(...TRUNCATED) | 0 | 0 | "ed 3 UUID entries, found {len(uuid_entries)}\"\n \n data_entry = None\n musik_entr(...TRUNCATED) | 4 | 3 | 1 | 75 | [{"role":"system","content":"You are an expert coding assistant operating inside pi, a coding agent (...TRUNCATED) | [{"turn":1,"reasoning":"Let me start by understanding the task requirements thoroughly:\n\n1. Read /(...TRUNCATED) | "[{\"type\": \"function\", \"function\": {\"name\": \"bash\", \"description\": \"Execute a bash comm(...TRUNCATED) | 0.00317 | 0.001697 | "Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERI(...TRUNCATED) | 2864b29d-e3d3-41eb-80d2-ca4a22d0f35a |
0009b0f843e653e94c73 | rts_task_068a4e6d69bfcbda4228b6d4 | 720 | failed_task | 2 | pi | deepseek/deepseek-v4-flash | "Edit /etc/fstab to add entries for every storage device listed in /etc/device-inventory.json, so th(...TRUNCATED) | registry-b.deepgrove.ai/rts:068a4e6d69bfcbda4228b6d4 | /app | easy | system-administration | 9 | true | "[\"tests failed: {\\\"reward\\\": \\\"0\\\", \\\"test_sh_returncode\\\": 0, \\\"test_sh_tail\\\": \(...TRUNCATED) | 0 | 0 | "ed 3 UUID entries, found {len(uuid_entries)}\"\n \n data_entry = None\n musik_entr(...TRUNCATED) | 4 | 3 | 1 | 75 | [{"role":"system","content":"You are an expert coding assistant operating inside pi, a coding agent (...TRUNCATED) | [{"turn":1,"reasoning":"Let me start by enumerating all the requirements carefully, then explore the(...TRUNCATED) | "[{\"type\": \"function\", \"function\": {\"name\": \"bash\", \"description\": \"Execute a bash comm(...TRUNCATED) | 0.002968 | 0.001593 | "Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERI(...TRUNCATED) | d1d55c35-1f50-494b-970b-19f2c2cdfa87 |
000cfb2405171959a4c7 | rts_task_8d6154dc01944a213f8772ae | 15,995 | failed_task | 1 | opencode | deepseek/deepseek-v4-flash | "In `/app`, use the `extra-object` npm package to process the nested object in `/app/data/input.json(...TRUNCATED) | registry-b.deepgrove.ai/rts:8d6154dc01944a213f8772ae | /app | easy | scripting-automation | 17 | true | "[\"tests failed: {\\\"reward\\\": \\\"0\\\", \\\"test_sh_returncode\\\": 0, \\\"test_sh_tail\\\": \(...TRUNCATED) | 0 | 0 | "_and_correct ________________________\n\n def test_r3_file_exists_and_correct():\n path =(...TRUNCATED) | 5 | 1 | 4 | 20 | [{"role":"system","content":"You are powered by the model named deepseek/deepseek-v4-flash. The exac(...TRUNCATED) | [{"turn":1,"reasoning":"The user wants me to be very careful and methodical. Let me start by enumera(...TRUNCATED) | "[{\"type\": \"function\", \"function\": {\"name\": \"bash\", \"description\": \"Executes a given ba(...TRUNCATED) | 0.004953 | 0.025261 | "Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERI(...TRUNCATED) | cd55041a-64e9-4dbe-88d5-9fc3fa65a643 |
00166aa2fafaccaec9e3 | rts_task_2231baeae30fac30f0bd6027 | 3,765 | failed_task | 1 | opencode | deepseek/deepseek-v4-flash | "Upgrade the shopping cart CLI at `/app/shopping_cart.py` to support a `--profile` argument.\nWhen r(...TRUNCATED) | registry-b.deepgrove.ai/rts:2231baeae30fac30f0bd6027 | /app | easy | scripting-automation | 20 | true | "[\"tests failed: {\\\"reward\\\": \\\"0\\\", \\\"test_sh_returncode\\\": 0, \\\"test_sh_tail\\\": \(...TRUNCATED) | 0 | 0 | " cmd in required_commands:\n> assert cmd in content, f\"Missing expected command '{cmd}' (...TRUNCATED) | 4 | 1 | 3 | 25 | [{"role":"system","content":"You are powered by the model named deepseek/deepseek-v4-flash. The exac(...TRUNCATED) | [{"turn":1,"reasoning":"Let me start by understanding the task. I need to:\n\n1. First, read the exi(...TRUNCATED) | "[{\"type\": \"function\", \"function\": {\"name\": \"bash\", \"description\": \"Executes a given ba(...TRUNCATED) | 0.00652 | 0.009712 | "Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERI(...TRUNCATED) | bb6bb937-94cb-41ab-8fcd-cc65ee513ea1 |
0016bbf50f93284c474e | rts_task_cdc8c7d72495a706d9bd1fa4 | 23,464 | failed_task | 1 | hermes | deepseek/deepseek-v4-flash | "In /app, repair the KSQL validation pipeline.\nFix validate.sh to count only standalone semicolons (...TRUNCATED) | registry-b.deepgrove.ai/rts:cdc8c7d72495a706d9bd1fa4 | /app | easy | database-operations | 33 | true | "[\"tests failed: {\\\"reward\\\": \\\"0\\\", \\\"test_sh_returncode\\\": 1, \\\"test_sh_tail\\\": \(...TRUNCATED) | 0 | 1 | "verify_out.read_text()\n assert \"Independent count:\" in vtext, \"verify_output.txt missing(...TRUNCATED) | 4 | 2 | 2 | 50 | [{"role":"system","content":"You are Hermes Agent, an intelligent AI assistant created by Nous Resea(...TRUNCATED) | [{"turn":1,"reasoning":"Let me start by enumerating every requirement from the task statement as a n(...TRUNCATED) | "[{\"name\": \"clarify\", \"description\": \"Ask the user a question when you need clarification, fe(...TRUNCATED) | 0.006958 | 0.017836 | "Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERI(...TRUNCATED) | 1103681c-4bb4-4f43-91f1-3d0d62199e6b |
0016e86ea9749540863a | rts_task_851dff8b02b5c44123310651 | 15,153 | failed_task | 1 | opencode | deepseek/deepseek-v4-flash | "The impatientanimedownloader tool is available at `/app/impatientanimedownloader/`.\nRefer to its l(...TRUNCATED) | registry-b.deepgrove.ai/rts:851dff8b02b5c44123310651 | /app | easy | scripting-automation | 22 | true | "[\"tests failed: {\\\"reward\\\": \\\"0\\\", \\\"test_sh_returncode\\\": 0, \\\"test_sh_tail\\\": \(...TRUNCATED) | 0 | 0 | "form linux -- Python 3.9.24, pytest-8.4.1, pluggy-1.6.0\nrootdir: /tests\nplugins: json-ctrf-0.3.5\(...TRUNCATED) | 5 | 4 | 1 | 80 | [{"role":"system","content":"You are powered by the model named deepseek/deepseek-v4-flash. The exac(...TRUNCATED) | [{"turn":1,"reasoning":"Let me start by enumerating all the requirements from the task, then systema(...TRUNCATED) | "[{\"type\": \"function\", \"function\": {\"name\": \"bash\", \"description\": \"Executes a given ba(...TRUNCATED) | 0.006219 | 0.020111 | "Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERI(...TRUNCATED) | c870ab02-45d2-4feb-aaa8-c0ffc84594e3 |
0019a648390fb6d06c7c | rts_task_c36f6bea1af4bade36f6ff98 | 22,333 | failed_task | 1 | pi | deepseek/deepseek-v4-flash | "Use the scripts and guide in /opt: gen_results.py (buggy generator), check.py (validator), and evid(...TRUNCATED) | registry-b.deepgrove.ai/rts:c36f6bea1af4bade36f6ff98 | /app | easy | scripting-automation | 30 | true | "[\"tests failed: {\\\"reward\\\": \\\"0\\\", \\\"test_sh_returncode\\\": 0, \\\"test_sh_tail\\\": \(...TRUNCATED) | 0 | 0 | "ne.strip())\n ]\n fail_lines = [\n line.strip() for line in report.splitli(...TRUNCATED) | 14 | 12 | 2 | 85.71 | [{"role":"system","content":"You are an expert coding assistant operating inside pi, a coding agent (...TRUNCATED) | [{"turn":1,"reasoning":"Let me break down this task carefully:\n\n1. First, I need to read the files(...TRUNCATED) | "[{\"type\": \"function\", \"function\": {\"name\": \"bash\", \"description\": \"Execute a bash comm(...TRUNCATED) | 0.004868 | 0.006925 | "Be extremely self-critical and grounded. Treat every belief about the state of the system as UNVERI(...TRUNCATED) | f7c2a16c-39a6-418c-9393-a3c038f2ed4e |
RTS × deepseek-v4-flash — REJECTED traces
25,210 rejected agentic attempts from deepseek/deepseek-v4-flash on the
Recursive Task Synthesis
corpus. The accepted counterpart is
RedMod/recursive-v4-flash.
Every row is an attempt that ran and was graded, and the task's own tests
returned 0. Harness failures (failed_pre_checks — the model never got a
scored shot) are deliberately EXCLUDED, so nothing here misattributes a
plumbing bug to the model.
Why these are worth keeping
They are near-misses, not garbage. Measured over a 1,544-attempt sample:
| signal | rate |
|---|---|
| garbled / non-Latin tokens | 0.8% |
| empty assistant text | 0.2% |
| looping (same tool call ≥5×) | 10.9% |
| median duplicate-tool-call fraction | 0.06 |
| median turns | 26 |
The model typically does the substantive work and then misses part of the deliverable spec. Categorising 25,007 rejected attempts by what the grader objected to:
| failure mode | share |
|---|---|
| a required artifact was absent | 60.9% |
| shortcut / forbidden modification detected | 19.0% |
placeholder (TODO, stub) left in output |
13.8% |
| artifact present but wrong shape | 3.1% |
| value mismatch | 2.1% |
Per-check fail rates show the same gradient — surface compliance is easy, exactness is not:
| check | fail rate |
|---|---|
required_evidence (did you leave something) |
32.7% |
intermediate_artifact |
56.3% |
no_shortcut |
72.2% |
final_semantics |
79.0% |
Partial credit — the binary reward hides most of the signal
reward is 0 for every row here, but the graders run several independent test
functions per task and these attempts pass most of them. Across 194,615
individual tests in 24,604 rejected attempts:
| individual tests run | 194,615 |
| tests PASSED | 113,057 (58.1%) |
| tests FAILED | 81,558 (41.9%) |
| median per-attempt pass rate | 50.0% |
| failed exactly ONE test | 4,574 attempts (18.6%) |
Most common shapes: 2/4 (4,151), 1/4 (4,139), 3/4 (2,848), 0/4 (2,096).
So a ~12% task-level accept rate coexists with a ~58% test-level pass rate, and
nearly one rejected attempt in five missed by a single check. Use the
tests_* columns to filter for near-misses rather than treating every row as an
equal failure.
Schema
One row per attempt (a task retried twice contributes two rows; dedupe on
prompt_id for task-level counts).
| column | |
|---|---|
prompt_id, task_id, dataset_index |
identity; task_id joins to the corpus |
skip_bucket |
always failed_task here |
try, harness, model |
which attempt, which harness ran it |
prompt, docker_image, cwd, difficulty, category |
task metadata |
messages |
the full trajectory as a native list of structs — {role, content, reasoning_content, tool_calls[], tool_call_id, name}; tool_calls is itself a list of {id, type, name, arguments} (arguments stays a JSON string, since its shape is per-tool) |
turns |
native list of {turn, reasoning, n_tool_calls} |
tools |
the harness tool schema (JSON string) |
n_turns, submitted, reasons |
episode outcome |
reward, test_sh_returncode, test_sh_tail |
grader output — test_sh_tail names the failing checks |
tests_total, tests_passed, tests_failed, tests_pass_pct |
per-TEST partial credit parsed from the grader output (null when unparseable) |
env_cost, token_cost |
real spend for this attempt |
prompt_hint |
scaffolding appended to the prompt, stripped from messages |
Run configuration
max_turns 100, max_tries 1, temperature 0.5, top_p 0.95,
reasoning_effort high, harnesses hermes / opencode / pi (no terminus2),
Daytona sandboxes, verifier agentic (the task's own tests/test.sh).
prompt_hint is set on these rows (the "hint-v3" self-critical + enumerate +
verify-as-stated instruction). It is stripped from messages so the
trajectory reads as the unmodified task — check the column before mixing with
unhinted data.
Caveats
rewardisMISSINGon ~3% of rows:test.shnever wrote/logs/verifier/reward.txt(usually a grader that died early). Treated as a failure, but it is not a clean 0.- ~4% of the underlying corpus is unsolvable regardless of model (the
instruction cites
rewrite_*.jsondocs the image never ships); those tasks were filtered out of this run's dataset, but related noise may remain.
- Downloads last month
- 219