{"data":{"kind":"file","path":"README.md","version_id":"afcf157yvh73xf4tepc5v05b","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":11646,"modified_at":"2026-10-01T08:53:04.820000","content_hash":"48c4ebdb9c4bd3f866ba5d7647b20b3f169cc86d57084401e9f4663d3a0c914c"},"entries":[],"content":"# proration-bench\n\nA **non-coding** RL environment and benchmark for subscription billing. The agent works as a billing operations analyst: it reads an account, a plan catalog, coupons and a history of mid-cycle plan and seat changes, and returns a structured, machine-checked answer. That answer covers the billing decision, the period and day counts, the prorated credit and charge lines with discounts, the net, and the rules that justify them. Review tasks also give the agent a billing system's (possibly wrong) output to validate, reconcile, diagnose or audit.\n\nThe agent never reads, edits, compiles or runs code. It has no tools: one prompt in, one JSON answer out.\n\n| | |\n|---|---|\n| Environment ID | `proration-bench` |\n| Framework | verifiers **v1** `Taskset` (`verifiers>=0.3.1`), single turn, no container |\n| Harness | `ProrationHarness`, a `NullHarness` with no tools, no shell and no code execution |\n| Rules | `rules@2.0.0`, versioned, shipped in the package and shown to the agent in full |\n| Grading | Strict JSON parse, then an exact comparison with an independent oracle, gated reward in [0, 1] |\n| Splits | train (on the fly), validation (1,000), public test (500, frozen), hidden (1,000, secret seed) |\n\n## Contents\n\n- [How it works](#how-it-works)\n- [Task families and levels](#task-families-and-levels)\n- [Quickstart](#quickstart)\n- [Running evaluations](#running-evaluations)\n- [Configuration](#configuration)\n- [Metrics](#metrics)\n- [Quality evidence](#quality-evidence)\n- [Repository layout](#repository-layout)\n- [Documentation](#documentation)\n\n## How it works\n\n```\nseed ──► scenario generator ──► scenario (facts only)\n                                   │\n                                   ├──► primary oracle (Fraction) ─┐\n                                   ├──► shadow oracle (Decimal) ───┴─► must agree exactly\n                                   ├──► mutators (realistic billing bugs) ─► \"billing system output\" for review tasks\n                                   └──► renderer (JSON export | support ticket | ledger) ─► user prompt\nsystem prompt = role + full rulebook + answer schema + formatting rules + one worked example\nagent ─► one <answer>{json}</answer> ─► strict parser ─► consistency gate ─► field checks ─► gated reward\n```\n\n- **Rules.** [`proration_bench/rulebook.md`](proration_bench/rulebook.md) defines R1–R19: calendar periods, billing-timezone change dates from timestamps written in arbitrary UTC offsets, inclusive remaining days, actual month and year lengths, exact arithmetic, per-line half-away-from-zero rounding to ISO 4217 minor units, plan-scoped compounding coupons (a change's credit and charge lines can be discounted differently), downgrade policy, same-value changes, trials, anchor clamping, multi-change histories, invoice structure, graduated and volume tiered seat pricing (with volume tiers, adding seats can be a downgrade), per-line sales tax with banker's rounding, voided change records that must be ignored, and the reason-code table. The agent sees all of it.\n- **Two oracles.** The primary oracle (`oracle.py`) implements each rule step as a `Policy` method. The shadow oracle (`shadow.py`) was written independently, from the rulebook alone, using `Decimal`. Every generated task must get identical answers from both, and both reproduce 50 hand-worked canonical cases.\n- **Billing bugs.** Twenty-five mutators each override one rule step (UTC date, off-by-one, 30-day month, 365-day year, ignoring seats, a missing credit line, coupon mistakes including ignoring a coupon's plan scope, truncation, net not summed, downgrade policy ignored, same-value invoiced, trial billed, anchor rollover, original-plan credit, netted or reordered lines, swapped tier modes, tax rounded half away, tax on gross, tax on net, voided change processed). Review tasks show their output as the \"billing system output\", and each mutator's rule becomes the answer key for diagnosis.\n- **Grading.** The expected answer is recomputed from the scenario at grading time. Nothing answer-like is stored in the task.\n\n## Task families and levels\n\n| Family | Task | Answer |\n|---|---|---|\n| A | Proration calculation | calculation |\n| D | Billing decision (invoice, credit, defer, none) | calculation |\n| F | Multi-step subscription history | calculation |\n| H | Edge-case-heavy calculation | calculation |\n| B | Invoice validation | review |\n| C | Reconciliation of an issued invoice | review |\n| E | Error diagnosis (which rules were violated) | review |\n| G | Audit / compliance with rule citations | review |\n\nLevels are **measured** from each scenario's features, not assigned by hand: L1 is a single taxed upgrade in a 31-day month; L2 adds month length, seats and downgrades; L3 adds timezone, DST, (plan-scoped) coupons and tiered pricing; L4 histories of 3–6 changes, often with voided records; L5 combines two edge cases (annual, leap, anchor clamping, trials, 0- or 3-decimal currencies, volume-tier traps) on a multi-change history; L6 review tasks with two combined billing bugs. From L3 up (and on half of L2) timestamps are written in assorted UTC offsets, and no task states the current period. See [docs/SPEC.md](docs/SPEC.md) and [docs/SAMPLE_TASKS.md](docs/SAMPLE_TASKS.md).\n\n## Quickstart\n\n```bash\ncd environments/proration_bench\nuv sync                                   # Python 3.11-3.13, pinned deps incl. tzdata 2025.2\nuv run pytest                             # 376 tests\nuv run python scripts/self_check.py       # 102 checks, writes reports/self_check.md and reports/red_team.md\nuv run python scripts/validate_dataset.py # all splits; exits 1 on any critical issue\nuv run python scripts/benchmark.py --policy oracle --split test   # end-to-end, no API key\n```\n\n## Running evaluations\n\n```bash\n# Through verifiers (Prime Inference by default: `prime login` or PRIME_API_KEY)\nuv run eval proration-bench -m <model-id> -n 50 -r 1 --env.taskset.split test\n\n# Any OpenAI-compatible provider through verifiers\nuv run eval proration-bench -m <model-id> --client.base-url https://provider/v1 \\\n  --client.api-key-var MY_KEY --env.taskset.split validation --env.taskset.num_tasks 100\n\n# Re-grade and break down a verifiers run\nuv run python scripts/benchmark.py --traces outputs/<run-dir>\n\n# Or call an endpoint directly with the same prompts and grader\nuv run python scripts/benchmark.py --model <model-id> --base-url https://provider/v1 --api-key-var MY_KEY --split test --num-tasks 50\n\n# Hidden split (eval operators only; the seed never goes in the repo)\nPRORATION_HIDDEN_SEED=<secret> uv run eval proration-bench -m <model-id> --env.taskset.split hidden\n```\n\n`prime eval run` on prime CLI ≤ 0.6.x uses the legacy loader, which requires `load_environment()`. Use `uv run eval`, or upgrade the CLI.\n\n### Windows\n\n`uv run eval` does not run natively on Windows: verifiers v1 imports the Unix-only module `fcntl` when it loads (`ModuleNotFoundError: No module named 'fcntl'`). Everything else in this package is platform-independent. Two options:\n\n- **WSL2 (recommended for `uv run eval`).** Install Ubuntu under WSL2 and run the commands above inside it.\n- **Natively, without verifiers.** `scripts/benchmark.py` sends the same system and user prompts to any OpenAI-compatible endpoint and grades with the same grader. `pytest`, `self_check.py`, `validate_dataset.py`, `generate_tasks.py` and `red_team_attacks.py` also run natively; only the verifiers wiring tests are skipped.\n\n```powershell\n$env:OLLAMA_API_KEY = \"...\"\nuv run python scripts/benchmark.py --model gpt-oss:120b --base-url https://ollama.com/v1 `\n  --api-key-var OLLAMA_API_KEY --split test --num-tasks 20 --concurrency 1 --max-tokens 16384\n```\n\n## Configuration\n\n`--env.taskset.*`:\n\n| Field | Default | Description |\n|---|---|---|\n| `split` | `test` | `train`, `validation`, `test` (frozen public file) or `hidden` (needs `PRORATION_HIDDEN_SEED`) |\n| `num_tasks` | train 2000, validation 1000, test 500, hidden 1000 | Tasks to emit |\n| `seed` | `0` | Train only: a different seed gives a different deterministic pool |\n| `families` | all of A–H | Families to include |\n| `levels` | 1–6 | Levels to include |\n\nThere are no grading knobs (`--env.taskset.task.*`), so scores stay comparable across runs.\n\n## Metrics\n\n| Name | Meaning |\n|---|---|\n| `billing` (reward) | Gated reward in [0, 1]; 1.0 only for an exactly correct answer ([docs/SPEC.md](docs/SPEC.md#reward-specification)) |\n| `exact` | 1 if reward is 1.0 |\n| `parse_failed` | 1 if the answer could not be parsed (no block, bad JSON, wrong types, bad money strings) |\n| `check_*` | Per-check results: decision, period, changes, lines, net, reason codes; for reviews also verdict, discrepancies, violated rules, corrected |\n| `level` | Difficulty level of the task |\n\n## Quality evidence\n\nLast local run (2026-09-29, rules 2.0.0):\n\n| Check | Result |\n|---|---|\n| Canonical hand-worked cases | 50/50 on both oracles |\n| Primary vs shadow oracle | 0 disagreements on 10,000 random scenarios (self-check --full) plus every generated task |\n| Red-team gates | 86/86 (degenerate answers, 25 single-bug agents, 23 malformed-output attacks, replay) |\n| Red-team round 2 | 20,000 fuzzed scenarios and 12,792 edge probes with 0 oracle disagreements; findings RT-01 to RT-05 fixed ([reports/RED_TEAM_REPORT.md](reports/RED_TEAM_REPORT.md), `scripts/red_team_attacks.py`) |\n| Self-check | 102/102 |\n| Dataset validation | 0 errors across train (2,000), validation (1,000), test (500) and hidden (1,000) |\n| Oracle policy on public test | mean reward 1.000 |\n| Real-model runs | Only 2- and 10-task smoke runs on rules 1.0.0 (gpt-oss:20b, nemotron-3-ultra). Rules 2.0.0 is not yet calibrated against real models ([docs/RUNBOOK.md](docs/RUNBOOK.md#calibration)) |\n| Work per task vs rules 1.0.0 (public test) | change records 1.62 → 2.97, invoice lines 2.07 → 3.16 (L5: 1.56 → 3.45), distinct realistic mistakes that break a task 7.9 → 9.7, plus a tax line calculation on every invoice |\n\n## Repository layout\n\n```\nproration_bench/                 # the shipped package (flat, one module per concern)\n  taskset.py                     # verifiers v1 Taskset, Task, no-tools Harness, system prompts\n  rulebook.md                    # rules@2.0.0, shown to the agent in full\n  rules.py                       # currencies and exact money, calendar and tzdata, typed models\n  oracle.py                      # primary oracle (a Policy per rule step) and the 25 billing-bug mutators\n  shadow.py                      # independent shadow oracle\n  generator.py                   # scenario sampler, validity checks, task specs and splits, renderers\n  verifier.py                    # strict parser, consistency and discrepancy checks, gated reward\n  data/                          # public_test.jsonl (frozen), DATASET_MANIFEST.json\ntests/                           # test_*.py, canonical_cases.json, redteam.py, dataset_checks.py (not shipped)\nscripts/                         # generate_tasks, validate_dataset, self_check, benchmark\ndocs/                            # SPEC, DATASET_AND_SECURITY, RUNBOOK, SAMPLE_TASKS, implementation plan\n```\n\n## Documentation\n\n- [SPEC.md](docs/SPEC.md): business rules and decisions, tasks and schemas, reward, verifier\n- [DATASET_AND_SECURITY.md](docs/DATASET_AND_SECURITY.md): generation, splits and validity, isolation, red-team results\n- [RUNBOOK.md](docs/RUNBOOK.md): commands, evaluation, release checklist, calibration, troubleshooting\n- [SAMPLE_TASKS.md](docs/SAMPLE_TASKS.md): 18 complete tasks with expected answers\n","encoding":"utf-8","truncated":false,"total_bytes":11646},"status":null}