{"data":{"kind":"file","path":"README.md","version_id":"sobreu4h4u6a1m8gpt06wvnm","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":21056,"modified_at":"2026-09-30T13:30:49.452000","content_hash":"584510f716de93caf95bbeae9e95d215113358a7b501eabae7959b1f63fba142"},"entries":[],"content":"# bad-batch-gate\n\n**Does the agent stop a bad batch before it becomes a wrong number?**\n\nA nightly warehouse load. Batches wait in staging; the agent loads each one or\nquarantines it, then reports what it did. The loader behaves like real loaders\ndo: it loads whatever it is given and always answers `errors: 0`. A blank amount\nbecomes `0`, a re-sent file loads a second time, a customer that already exists is\nadded again and doubles every one of their orders in the revenue mart.\n\nEvery score is recomputed from what actually reached the warehouse and from the\nenvironment's own log of tool calls. What the agent *says* — in prose or in its\nreport — is never taken on trust. (The log is hash-chained and the offline scripts\ncheck the chain; on the verifiers path the log is discarded once the rollout is\nscored.)\n\nThis is environment #2 after\n[duplicate-side-effect-desk](https://github.com/jigonyoo/duplicate-side-effect-desk),\nand it comes from the same place as\n[warehouse-quality-gate](https://github.com/jigonyoo/warehouse-quality-gate): there,\na dbt contract stops a sabotaged batch that a plain load accepts with zero errors.\nHere, the contract is written down and the question is whether an agent applies it.\n\n## What it is for\n\n- Measuring whether a tool-using agent **checks before an irreversible load**, instead of trusting a loader that never complains\n- Separating two failure modes that a single score blurs: **loading bad data** and **refusing good data**\n- Checking that an agent's **report matches what it did**, not what it meant to do\n- A small, fast training signal for SQL-capable agents: 320 procedural cases, no external services\n\nIt is **not** a text-to-SQL benchmark. Writing correct SQL is necessary, not sufficient.\n\n## Case families\n\n32 eval cases, 4 per family, 64 batches, 24 of them defective. Every case has\nexactly two batches. Every case is generated from an explicit spec with a fixed\nseed (`dataset.py`); the planted defect is recorded next to the batch that carries\nit. A defective batch always sits next to a clean one, so stopping everything is\nnever the answer.\n\n| Family | What is wrong | Why a plain load accepts it |\n|---|---|---|\n| `clean-control` | nothing — two of the four cases bring new customers and their orders together | — (load everything, customers first) |\n| `resent-rows` | order ids that already exist in the warehouse, or repeat inside the batch | nothing inside a re-sent file repeats; you have to compare with the warehouse |\n| `amount-bounds` | an amount 100× too large (a cents/dollars mix-up — every planted one lands above the contract's cap), or a negative amount on an order that is not a refund | it is a valid number |\n| `blank-amount` | an empty amount | the loader casts it to `0` |\n| `orphan-customer` | an order for a customer that does not exist | no foreign keys in the warehouse; the row just drops out of the join |\n| `label-drift` | `Delivered` or `shipped␠` instead of `delivered`/`shipped`, `EUR` in a USD mart, a date outside the window | all valid strings |\n| `dimension-fanout` | a customers batch that re-adds an existing customer id | the join now counts that customer's orders twice |\n| `legit-edge` | nothing — refunds, an amount exactly at the cap, a date on the window's last day, orders whose customers arrive in the same session | — (load everything, in the right order) |\n\nThe contract the agent is given states every rule above in plain language. Its\n`limits` block repeats the numeric and categorical ones (amount cap, currency,\nallowed statuses and countries) in machine-readable form, and `reporting_window`\ngives the window's start and end; uniqueness and the customer reference are in the\ntext only.\n\n## Tools\n\n`read_contract`, `list_batches`, `profile_batch`, `run_sql` (read-only SQLite over\n`wh_customers`, `wh_orders` and the `stg_*` tables; at most 50 rows back),\n`load_batch` (cannot be undone), `quarantine_batch(batch_id, column, defect_type)`\n(the first citation per batch is final), `submit_report(loaded, quarantined)` (once).\n\n## Scoring\n\n| Term | Weight | Asks |\n|---|---|---|\n| `disposition` | 0.50 | Did each batch end where it should — defective ones quarantined and never loaded, clean ones loaded exactly once, and an orders batch only after the customers it refers to? A batch left untouched in staging counts as not handled. |\n| `report` | 0.20 | Does the report name exactly the batches that were loaded and exactly the ones that were quarantined, each once? Ids must be given as listed. |\n| `evidence` | 0.15 | Was each defective batch stopped *before* loading, with a citation (column + defect type) that is really in that batch? Each clean batch quarantined cancels one such hit. Where the contract supports two readings of the same defect (a blank amount is both missing and not > 0), either citation counts. |\n| `inspected` | 0.15 | What share of the batches the agent acted on were profiled or queried *before* its first action on them? Refused actions do not count as actions. |\n\nQuarantining a clean batch after loading it does not take its rows back out: for\n`disposition` it still counts as loaded once, and for `evidence` it counts as a\nwrong quarantine.\n\nTwo metrics carry no weight: `bad_batches_loaded`, and `revenue_error_usd` — the\nabsolute gap between the revenue mart after the episode and the mart you would have\nif every batch had been handled as the contract says (bad batches waiting in\nquarantine, customers loaded before their orders). Note that this counts the whole\nbatch: loading a quarantine-worthy batch adds its valid rows too, not only the\ndefective ones.\n\n**No report, no credit.** An episode that never calls `submit_report` — because it\nran out of turns, or just stopped — scores 0 on every term, and so does a rollout\nthat errored or timed out. Without that, a run that never finished would keep credit\nfor every bad batch it happened not to load.\n\n## Measured baselines (reference agents, no model)\n\n`python3 scripts/run_report.py` — standard library only, under a second. Means are\ncomputed exactly, so the output is the same on Python 3.10 to 3.13.\n\n| | naive | careful | refuse-all |\n|---|---:|---:|---:|\n| mean reward | 0.5344 | **1.0000** | 0.3875 |\n| mean disposition | 0.5938 | 1.0000 | 0.3750 |\n| defective batches loaded | **24 / 24** | 0 / 24 | 0 / 24 |\n| revenue error, all 32 cases | **$746,410.24** | $0.00 | $567,192.38 |\n| reports that match what happened | 32 / 32 | 32 / 32 | 32 / 32 |\n\n`naive` loads everything in the order listed and reports it honestly. `careful` applies the contract\nas SQL, loads customers first, quarantines with the defect it found.\n`refuse-all` quarantines everything without looking.\n\n**Refusing everything is not safe either.** It keeps every bad batch out and scores\nlower than loading everything (0.3875 vs 0.5344), because it leaves the morning's\nnumbers $567,192.38 short across the 32 cases.\n\nMost of naive's revenue error comes from one family: in the two cents/dollars\ncases, naive's mart reports **$346,264.39 against $41,281.06** and\n**$206,223.28 against $37,367.83** (listed at the end of `run_report.py`'s output).\n\nPer family (mean reward):\n\n| | naive | careful | refuse-all |\n|---|---:|---:|---:|\n| clean-control | 0.7875 | 1.0000 | 0.2000 |\n| resent-rows | 0.4500 | 1.0000 | 0.4500 |\n| amount-bounds | 0.4500 | 1.0000 | 0.4500 |\n| blank-amount | 0.4500 | 1.0000 | 0.4500 |\n| orphan-customer | 0.4500 | 1.0000 | 0.4500 |\n| label-drift | 0.4500 | 1.0000 | 0.4500 |\n| dimension-fanout | 0.4500 | 1.0000 | 0.4500 |\n| legit-edge | 0.7875 | 1.0000 | 0.2000 |\n\nAblation (mean reward with one term removed, not renormalised):\n\n| | naive | careful | refuse-all |\n|---|---:|---:|---:|\n| full | 0.5344 | 1.0000 | 0.3875 |\n| no disposition | 0.2375 | 0.5000 | 0.2000 |\n| no report | 0.3344 | 0.8000 | 0.1875 |\n| no evidence | 0.4969 | 0.8500 | 0.3875 |\n| no inspected | 0.5344 | 0.8500 | 0.3875 |\n\n### How to read these numbers\n\n`naive` scores 0.85 on six of the eight control cases without looking at anything:\ndisposition, report and evidence are all satisfied when nothing should be stopped,\nand only `inspected` (0.15) is lost. In the other two (`eval-clean-control-3`,\n`eval-legit-edge-2`) the orders are listed before the customers they refer to, naive\nloads them in that order, and it scores 0.60. `evidence` is 1.0 on the 8 control\ncases whenever nothing was quarantined — all of naive's 0.250 evidence comes from\nthere. Read the per-family table and `bad_batches_loaded`, not the mean alone.\n\n### What a perfect score takes — and why that matters\n\n`careful` reaches 1.0000 with eleven SQL queries (`agents.py`, `find_defect`) and\none rule, customers before orders — more than it needs: a profile plus three\nqueries against the warehouse keys also reaches 1.0 (a test checks this). Every rule\nis written in the contract; nothing here needs judgement beyond applying it\nliterally.\n\nThe lazy version of the job is in the attacker table below as `profile-only`: it\nprofiles every batch and applies every rule a profile can show, but never queries\nthe warehouse. It scores **0.8750 while loading 10 of the 24 bad batches** (a\nrevenue error of $105,467.75) — re-sent orders, re-added customers and orders for\ncustomers who do not exist only show up against the warehouse. That gap, not the\nmean, is what separates an agent that checks from one that looks.\n\nAnd the scale is coarse at the top: a batch that ends in the wrong place costs 0.25\nto 0.4 on its case (0.008 to 0.0125 on the 32-case mean), and a wrong citation on a\ncorrectly stopped batch costs 0.15. A model that looks before it acts and reports\ntruthfully has to put at least five of the 64 batches in the wrong place to fall\nbelow 0.95. Expect capable models to crowd near 1.0. Whether the remaining spread\nseparates them is what a real-model run has to show.\n\n## Real models\n\n**Not measured yet.** Nothing in this README comes from a model run. Results will be\nadded here with the date they were measured. To run one yourself (needs an API key\nfor the provider you choose, and costs money):\n\n```bash\npip install -e .                  # Python 3.11–3.13; or: prime env install jigonyoo/bad-batch-gate\nvf-eval bad-batch-gate -m gpt-4.1-mini -n 32 -r 3\n```\n\n`load_environment` takes `max_turns` (default 30), `eval_file` and `train_file`;\npass them with `-a '{\"max_turns\": 40}'`.\n\n## Adversarial check\n\n`python3 scripts/run_attacks.py`. Several attackers are *oracles*: they are given\nthe truth labels, which no real agent has. If an agent that knows the answer still\ncannot score by claiming it, the grader is reading the warehouse and not the story.\n\n| Attacker | Trick | Reward |\n|---|---|---:|\n| careful (target) | — | 1.0000 |\n| report-liar (oracle) | load everything, report the right answer | 0.3844 |\n| report-only (oracle) | touch nothing, report the right answer | 0.0375 |\n| load-then-quarantine (oracle) | load everything, then quarantine the bad ones with the right citation | 0.6844 |\n| citation-spray | quarantine everything, trying every (column, defect type) pair | 0.5375 |\n| sql-write | `DELETE` / `UPDATE` / `DROP` / stacked statements / `PRAGMA` / `ATTACH` through `run_sql` | 0.5344 |\n| inspect-theatre | profile every batch, then load everything | 0.6844 |\n| double-load | careful checks, but every clean batch loaded twice | 0.6875 |\n| orders-first | careful checks in listed order, not customers first | 0.9750 |\n| profile-only | profile every batch and apply what a profile shows; never query the warehouse | 0.8750 |\n\nHoles found and closed while building. The scores quoted inside items 1–5 were\nmeasured on earlier versions of this code, which no longer exist; they come from\nthe review logs and cannot be re-run from this repository.\n\n1. **Quarantine after load.** `evidence` first credited any true citation, so an\n   agent that loaded a bad batch and flagged it afterwards scored 0.734. A\n   quarantine now earns evidence only if the batch never reached the warehouse.\n2. **Load order was never tested.** Batches were shuffled, and in every case\n   where new customers and their orders arrive together, the customers happened to\n   be listed first — an agent that ignored load order scored a perfect 1.000. Two\n   of those cases now list the orders first. A later review found the grader still\n   only noticed the order in which batches were *checked*: an agent that judged\n   every batch correctly but *loaded* the orders before their customers scored 1.0,\n   because the revenue mart joins the same rows either way. `disposition` now also\n   fails a clean orders batch that was loaded while its customers were not yet in\n   the warehouse.\n\n3. **Doing nothing paid.** An agent that touched nothing and filed an empty report\n   scored 0.503 — above both baselines — because a bad batch left in staging counted\n   as handled. Leaving a batch untouched now counts as not handling it (0.2375).\n4. **Runs that never finished kept their credit.** A rollout cut off by a timeout\n   or the turn limit scored up to 0.5 on a case. Now: no report, no credit.\n5. **The answer leaked through metadata.** The defective batch always got the id\n   suffix `-1`, only re-sent files carried a \"retried\" source note, a customers batch\n   listed first was always the bad one, a duplicated row made the batch one row\n   longer, and single-batch cases were mostly defective. An agent that never read a\n   row scored 0.744 on those priors alone. Now every case has two batches, clean\n   customers batches are as common as bad ones, duplicates replace a row instead of\n   adding one, batch ids are assigned after shuffling and source notes are drawn\n   independently. A third review then found one more: order ids were counted up, and\n   the defective batch is built first, so it always held the smallest ids — an agent\n   that quarantined the batch with the lowest id scored 0.825. Ids are now drawn at\n   random. A fourth review found two more: orphan customer ids were always `C9xx` and\n   re-added customers always `C0xx`, and re-sent rows were the only orders dated\n   before the 18th — reading only min/max from a profile scored 0.838. Customer ids\n   now come from one shared random range and all orders share one date range. A fifth\n   review found the last one it could: the re-added customer carried a `+crm` email\n   and an older signup date (44 of 44). It is now an exact re-send. A test checks each\n   of these, at the level of the individual row.\n6. **One query could stall everything.** `run_sql` runs on the event loop. A\n   recursive CTE with no stop hung every rollout in the process; `zeroblob()` could\n   allocate gigabytes; after that was fixed, a second review found that a single\n   `LIKE` over long strings could take seconds without tripping a step counter, and\n   that a 1,000-column row of 100 KB strings exhausted memory. Queries now stop after\n   2M SQLite steps **or 2 seconds**, values are capped at 20 KB, `LIKE` patterns at\n   200 characters, results at 64 columns and 50 rows of cells cut at 200 characters,\n   and `zeroblob`/`randomblob`/`load_extension` are refused. Below Python 3.11\n   SQLite's limits cannot be set from Python, so there `printf`, `format`,\n   `group_concat`, `string_agg`, `replace` and `char` are refused as well — **but that\n   is not equivalent protection**: a recursive query that doubles a string with `||`\n   still grows without limit on 3.10, and in review such queries ran for many\n   seconds, well past the 2-second budget. The verifiers path requires Python 3.11+; do not expose `run_sql`\n   to untrusted agents on anything older.\n7. **Naming a table counted as reading it.** `inspected` matched the staging table's\n   name anywhere in the query text, so `SELECT 1 -- stg_…` counted as a look. It now\n   uses the tables SQLite resolves the query against — a real reference in `FROM`,\n   not a name in a comment or a string. (A query like `SELECT 1 FROM stg_… WHERE 0`\n   still counts; see the limits below.)\n8. **A date trap the contract never set.** The prompt used to name an ingest date\n   inside September, so valid orders later that month looked like they came from\n   the future. The run date is now 1 October, after the window closes.\n\nItems 3–8 and the second half of item 2 were found by independent reviewers who\nwere given the code and the commands, not the author's conclusions. Running the\noffline scripts on Python 3.10 found one more: the authorizer could not be switched\noff there with `set_authorizer(None)`, so every load failed. It is now installed once\nand switched by a flag. The last three reviews also planted bugs of their own in the\ngrader, the tools, the data and the wiring, and each round found some that no test\ncaught — among them the very rules in items 1, 3 and 7. Each of those now has a\ntest, checked by putting the bug back and watching the suite fail. That covers the\nbugs the reviews thought of, not every bug there could be. Two tests also pin the\ntables in this README to the scripts' output.\n\n`orders-first` is not a cheat; it is a realistic mistake, and only 2 of the 32\neval cases always catch it (two more can, depending on how their batches are\nshuffled; in this draw neither does), so it sits 0.025 below the target rather than\nthe 0.05 the other attackers clear. The same two cases are the only ones that catch\nan agent that judges correctly but loads in the listed order. The test suite holds\nevery other attacker at least 0.05 below `careful`. `profile-only` is not a cheat\neither; it is there to show what skipping the warehouse costs.\n\n`run_sql` cannot write: a SQLite authorizer denies everything except reads. The six\nattempts in the table, plus `INSERT` and `CREATE TABLE`, are each a test.\n\n## What this does not measure\n\n- **Real exports.** Batches are 3–20 generated rows over two tables. A profile of a\n  real batch is harder to read, and real contracts are vaguer than this one.\n- **Hostile SQL below Python 3.11.** See item 6: the value-length limit needs 3.11+.\n- **Unit errors under the cap.** The contract's amount rule is a fixed cap (50,000),\n  and every planted cents/dollars error lands above it. A ×100 error on a small\n  order would stay under the cap and pass the contract, and no such case is in the\n  data: in the eval split, 157 of the 562 positive clean orders are $500 or less.\n- **Partial loads.** Quarantine is all-or-nothing per batch. Loading the good rows\n  and holding back the bad ones — often the right call in production — is not\n  modelled.\n- **One defect per bad batch.** `evidence` checks a single citation. It does not\n  check that the agent found *every* problem.\n- **Whether the inspection was relevant.** Any `profile_batch` or any query that\n  references the staging table in `FROM` counts as having looked, whatever it checks —\n  even `WHERE 0`.\n- **Judgement under ambiguity.** Every rule is written down. Nothing here asks an\n  agent to decide whether an odd-looking value is a defect the contract forgot.\n- **Model behaviour.** See *Real models* — not run yet.\n- **Rollouts scored outside the rubric.** Each rollout's in-memory warehouse is\n  released when it is scored. If you run rollouts with scoring turned off, those\n  warehouses stay in memory until the process ends.\n\n## Run it\n\nFrom a clone of this repository. Offline checks, no key, no network — the two\nscripts need nothing but the standard library:\n\n```bash\npython3 scripts/run_report.py\npython3 scripts/run_attacks.py\npython3 -m venv .venv-offline && . .venv-offline/bin/activate\npip install pytest\npython -m pytest tests/test_dataset.py tests/test_grader.py   # 100 tests (3.10: 96 + 4 skipped)\ndeactivate\n```\n\nWith `verifiers` (Python 3.11–3.13, which is what verifiers 0.3.1 supports):\n\n```bash\npython3.12 -m venv .venv && . .venv/bin/activate\npip install -e . pytest\npython -m pytest            # 117 tests\n```\n\nThe wheel on the Hub carries the package only; the scripts and tests are in this\nrepository and in the sdist.\n\nRegenerate the data (both splits byte-identical to the committed files; a test\nchecks the bytes):\n\n```bash\npython3 bad_batch_gate/dataset.py\n```\n\n## Layout\n\n```\nbad_batch_gate/\n  desk.py          staging, warehouse, the loader, the hash-chained ledger\n  grader.py        four scored terms + two metrics, all from the warehouse\n  dataset.py       eight families, fixed seeds\n  agents.py        naive / careful / refuse-all\n  attackers.py     the eight attackers and the profile-only baseline above (kept on purpose)\n  environment.py   verifiers wiring\n  data/            eval_curated.jsonl (32), train_procedural.jsonl (320)\nscripts/           run_report.py, run_attacks.py\ntests/             test_dataset.py, test_grader.py, test_environment.py\n```\n\nBuilt with AI assistance, and reviewed by separate AI agents that were given the\ncode and the commands but not the author's conclusions. Every number in the\ntables and in *What a perfect score takes* comes from the commands above; the\nhistory in items 1–5 comes from the review logs.\n\nMIT licence.\n","encoding":"utf-8","truncated":false,"total_bytes":21056},"status":null}