{"data":{"kind":"file","path":"README.md","version_id":"lzcwiziyu2g4w9l85poffptt","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":29319,"modified_at":"2026-09-25T07:44:36.047000","content_hash":"3183d9c2a9114f6f5e1b08b991c6352b34f1610e183b52374a2ec3488efccff2"},"entries":[],"content":"# billing-payment-debug\n\nAn agent debugs a small Java 17 payment and billing service with **seeded defects**: all ten in the two benchmark tasks, or a random subset in any number of sampled training tasks. A hidden test suite, run in a fresh sandbox after the agent finishes, scores each defect separately.\n\n| | |\n| --- | --- |\n| Environment ID | `billing-payment-debug` |\n| Python package | `billing_payment_debug`, which exports `BillingTaskset` |\n| Framework | [verifiers](https://github.com/primeintellect-ai/verifiers) **v1** `Taskset` / `Task` API (`verifiers>=0.3.1`). There is no `load_environment()` entry point. |\n| Task type | Multi-turn tool use in a container, with the `bash` harness by default |\n| Tags | coding, debugging, java, sandbox, multi-turn, train, eval |\n\n---\n\n## 1. Overview\n\nThe sandbox holds `billing-service`, a dependency-free Java 17 codebase (42 source files) covering invoices, coupons, tax, idempotent card charges, refunds, processor webhooks, mid-cycle plan proration, metered usage, dunning, credit notes and statements. About half of it (dunning, credit notes, statements, CSV export, usage alerts, plan catalog, audit log) is correct code documented in the same README, so the agent has to audit it rather than assume every class is broken. Ten defects of different kinds can be seeded into it. Each task activates either all ten or a sampled subset, and every inactive defect is already fixed in that task's starting repo. The agent's job is to find and fix the active defects without changing the public API or breaking anything else.\n\nScoring is per defect, from hidden tests. Each defect's score is gated on \"guard\" tests of correct behavior in the same area, so an over-broad fix earns nothing.\n\n## 2. What the agent sees\n\nThe agent's sandbox (`/workspace/billing-service`) contains the files in `billing_payment_debug/assets/task_repo/`. For a sampled task, the fix hunks of every inactive defect (from `defects.py`) are applied first. The table describes the full task:\n\n| Path | Contents |\n| --- | --- |\n| `README.md` | How to build and test, the business rules the code must implement, and the rule not to change public signatures |\n| `run_tests.sh` | Compiles everything with `javac` and runs the public tests |\n| `src/main/java/com/acme/billing/{model,repository,gateway,service,notify,export}/` | The defective service (42 files) |\n| `src/test/java/com/acme/billing/*Test.java`, `RecordingGateway.java` | 23 public tests. Only 2 fail with all ten defects active (`coupons`, `proration`); the other eight defects are **silent**. |\n| `src/test/java/com/acme/billing/testkit/` | The built-in test kit (`@Test`, `Assertions`, `TestRunner`) |\n\nIt also gets the task prompt (see §4) and the harness's tools: `bash` plus a single-replacement `edit` tool, by default.\n\n## 3. What the agent is NOT allowed to see\n\nNone of these files is ever written into the agent's sandbox:\n\n| Component | Location in this repo | Agent sandbox | Grading sandbox | Used by |\n| --- | --- | :---: | :---: | --- |\n| Task README (business rules) | `billing_payment_debug/assets/task_repo/README.md` | **Yes** | No | agent |\n| Buggy source | `billing_payment_debug/assets/task_repo/src/main/java/` | **Yes** (editable) | The agent's edited `.java` files only | agent, grader |\n| Public tests | `billing_payment_debug/assets/task_repo/src/test/java/` | **Yes** | No. They are only SHA-256 compared on the host. | agent, tamper check |\n| Test kit | `billing_payment_debug/assets/task_repo/src/test/java/com/acme/billing/testkit/` | **Yes** | Pristine copy pushed from the host | agent, grader |\n| Hidden tests | `billing_payment_debug/assets/hidden_tests/` | No | Yes, pushed only at grading time | grader |\n| Gold solution | `billing_payment_debug/assets/gold/` | No | Only during gold validation (§11), never in a real rollout | `validate`, `self_check.py` |\n| Defect catalog (fix hunks, tickets) | `billing_payment_debug/defects.py` | No, though each task's repo is built from it on the host | No | building task repos and prompts |\n| Active-defect list | `BillingData.active_defects` | No. It isn't in the prompt, and sampled task names and descriptions are neutral. | No | grader |\n| Defect manifest | `defect_manifest.md` | No | No | maintainers only. Not read by any code. Not in the wheel, but it ships in the sdist. |\n| Grading script | generated by `grading.grade_script()` | No | Yes, generated per grading run | grader |\n| Signing keys | generated per run with `secrets.token_hex(32)` | No | Written as `keys.txt`, read, then deleted before any JVM starts | grader |\n| Grading report | `trace.info[\"grading\"]` | No | No, it's host-side | rewards, metrics |\n| Self-check script | `scripts/self_check.py` | No | No | maintainers only. Not in the wheel, but it ships in the sdist. |\n\nThe agent's sandbox also has no git history, since the repo is written in from a tar archive. It has no internet access while the agent works, unless `allow_network=true`.\n\n## 4. Task structure and variants\n\nEvery task uses the same codebase. Tasks differ in which defects are **active** (seeded) and in the **prompt variant** (config `variants`, default `[\"tickets\", \"minimal\"]`):\n\n- **`tickets`**: one symptom-level support ticket per active defect, with no file or method names.\n- **`minimal`**: \"find and fix every defect\", with the README's business rules as the only guide.\n\nBoth variants tell the agent to keep public signatures and not to edit the existing tests, and that a separate hidden suite grades the work.\n\n**Benchmark tasks** (`include_full_tasks`, default on, always listed first):\n\n| Task name | Variant | Active defects |\n| --- | --- | --- |\n| `billing-debug-tickets` | `tickets` | all 10 |\n| `billing-debug-minimal` | `minimal` | all 10 |\n\n**Sampled tasks** (`num_sampled_tasks`, default 0):\n- Each sampled task gets a random variant and a random subset of defects. Subset sizes are drawn uniformly from `min_active_defects`..`max_active_defects` (default 1..5).\n- Sampling is seeded (`sampling_seed`) and never repeats a (variant, subset) pair, so the same config always yields the same task list. There are up to 1,274 distinct tasks at the default sizes, and 2,046 across all sizes.\n- Sampled tasks are named `billing-debug-<variant>-s<idx>`. The name doesn't reveal which defects are active.\n- **How a sampled repo is built:** start from the fully seeded `task_repo/` and apply the fix hunks of every inactive defect from `defects.py`. `defects.verify()` runs on every load. It checks that each hunk applies exactly once, and that applying all of them reproduces `assets/gold/` byte for byte.\n\n```bash\n# e.g. a 500-task training set, benchmark tasks excluded\n--env.taskset.include-full-tasks false --env.taskset.num-sampled-tasks 500\n```\n\nPer-task settings (`BillingTaskset.load`):\n- **Image**: `eclipse-temurin:17-jdk`\n- **Working directory**: `/workspace/billing-service`\n- **Resources**: 2 CPUs, 4 GB memory, 5 GB disk\n- **Timeouts**: setup 300s, agent `agent_timeout_seconds` (default 1800s), finalize `grading_timeout_seconds + 60`, scoring 60s\n\n## 5. Seeded defects\n\n| # | Group (`defect_<group>`) | Category | Location | Symptom | Failing public test? |\n| --- | --- | --- | --- | --- | :---: |\n| 1 | `proration` | off-by-one | `SubscriptionService.prorate` | plan changes bill one day short | yes |\n| 2 | `tax` | floating point | `TaxService.calculateTax` | tax computed in `double`, off by a cent at half-cent boundaries and at scale | no |\n| 3 | `coupons` | business rule | `DiscountService.applyCoupons` | percentage coupons add up instead of compounding | yes |\n| 4 | `refunds` | concurrency | `RefundService.issueRefund` | the balance check is correct but not atomic with the processor call and save, so overlapping refunds overdraw | no |\n| 5 | `idempotency` | concurrency | `PaymentService.charge` | overlapping retries with the same idempotency key double-charge | no |\n| 6 | `line_rounding` | rounding order | `LineItem.total` | unit price rounded before multiplying | no |\n| 7 | `coupon_dedup` | input validation | `InvoiceService.generateInvoice` | the same coupon code, in any case, applies twice | no |\n| 8 | `tax_cache` | stale cache | `TaxService` rate cache | a re-registered tax rate is ignored | no |\n| 9 | `webhook_redelivery` | state (with a concurrency trap) | `WebhookService.handle` | an event is marked processed before it's applied, so a delivery that fails is never re-applied. The obvious fix (mark after applying) lets overlapping redeliveries apply it twice, which a guard test catches. | no |\n| 10 | `usage_timezone` | time zones | `UsageReportService.unitsByDay` | days bucketed in UTC, not the billing zone | no |\n\nThree of the ten are concurrency or state bugs (`idempotency`, `refunds`, `webhook_redelivery`). None of them can be found from test output, and each has a plausible half-fix that scores 0. The agent-facing README states every rule exactly, but without worked numeric examples.\n\nFull details (exact code, gold fix, test mapping) are in the maintainer-only [`defect_manifest.md`](defect_manifest.md).\n\n## 6. Docker / sandbox architecture\n\n```\nAgent Sandbox  (task image: eclipse-temurin:17-jdk)\n    ↓  setup: task_repo/ tar-extracted into /workspace/billing-service\nBuggy Java Billing Repository\n    ↓\nAgent edits and runs tests   (bash/edit tools, ./run_tests.sh, no internet)\n    ↓  finalize: tar -ch src/  (symlinks dereferenced, ≤ 8 MB)\nSource snapshot  ──────────►  host: tamper check (public-test hashes), edit diff\n    ↓  only src/main/java/**/*.java, minus testkit/ and hidden/ package paths\nFresh Grader Sandbox  (new box, same image and config, booted just for grading)\n    ↓\nPristine test kit + hidden tests + grade.sh + keys.txt   (pushed from the host)\n    ↓  javac → one JVM per hidden group, signed RESULT lines\nDefect / regression / process checks   (host verifies signatures, scores groups)\n    ↓\nReward 0.0–1.0\n```\n\n- The agent's sandbox is never used for grading when `isolated_grading=true`, the default. Changes the agent made to its own box, including the JDK, `/etc` or shell profiles, have no effect on grading.\n- The grading box is created with `make_runtime(runtime.config)`, the agent box's own resolved config: same image, resources and network policy. It is always stopped afterwards, even if grading fails.\n- Grading needs a container runtime (`docker` or `prime`). The `subprocess` runtime is refused.\n\n## 7. Agent execution flow\n\n1. **Provision**: the runtime starts a box from the task image.\n2. **Task setup** (`BillingTask.setup`): the files of `assets/task_repo/` are packed into a deterministic tar on the host, written into the box, and extracted to `/workspace/billing-service`.\n3. **Harness setup**: the harness installs its agent program. The bash harness bootstraps `uv` with curl. The execution network policy is applied after this, so the harness has network access but the agent does not, by default.\n4. **Agent run**: the model works through the harness tools until it stops or `agent_timeout_seconds` runs out.\n5. **Finalize and score**: see §8. If the rollout itself failed (a harness or model error), the framework skips finalize and scoring and records the error.\n\n## 8. Grading flow\n\n`BillingTask.finalize` calls `grade()`, capped by `grading_timeout_seconds`:\n\n1. **Snapshot** (`take_snapshot`): `tar -chf … --exclude=build src` in the agent box, with a 120s timeout. If `src/` is missing, the snapshot is empty and grades as a compile failure, not a rollout error. A snapshot over 8 MB is rejected and **scores 0**.\n2. **Inspect** (`inspect_snapshot`): compares the SHA-256 of each original public test file with the snapshot, giving `public_tests_intact`. It also records which `src/main/java` files changed.\n3. **Select sources** (`grading.grading_sources`): keeps only regular `.java` files under `src/main/java`. It drops anything under the `com/acme/billing/testkit/` or `com/acme/billing/hidden/` paths, then adds the pristine test kit and the hidden tests from the host. More than 500 source files is rejected and **scores 0**. Links, devices, absolute paths and `..` paths in the tar are ignored.\n4. **Fresh box**: starts a new runtime from the agent box's config.\n5. **Bundle**: writes `src/`, `grade.sh` and `keys.txt` into `/tmp/billing-grade-<uuid>`. There is one random key per JVM run.\n6. **`grade.sh`**:\n   - reads the keys into memory and deletes `keys.txt`,\n   - works within a total budget of `suite_timeout_seconds` (default 900s). Every `javac`/`java` call is capped by what's left of it, and runs that no longer fit are skipped, so their tests count as failed. Code that hangs therefore scores low instead of erroring the rollout.\n   - compiles the agent's sources into `build/app`, then the pristine test kit and hidden tests into `build/grader`, capped at 300s. At runtime `build/grader` comes first on the classpath, so an agent class can't shadow a grader class.\n   - runs every test JVM as `nobody` when grading runs as root (via `setpriv`), after making the bundle non-writable. Any processes a run leaves behind are killed before the next one. The report records this as `test_user`.\n   - is wrapped in one function, so bash parses the whole script before running it.\n   - runs each hidden group in its own JVM, capped at 600s, with `-XX:+DisableAttachMechanism`, a 20s per-test timeout, and that run's key on stdin. The concurrency groups (`idempotency`, `refunds`, `webhook_redelivery`) run `concurrency_repeats` times each (default 3).\n7. **Verify and score** (host): only `RESULT` lines whose signature equals `sha256(key|status|test_id)` count. A test passes only if it reported PASS in every run and on every line. A missing report is a failure. The report goes into `trace.info[\"grading\"]`.\n\n## 9. Hidden-test architecture\n\nThe hidden tests are in `assets/hidden_tests/com/acme/billing/hidden/`: 11 classes and 97 tests in total. The denominators are read from these source files (`grading.expected_tests`), not from runner output.\n\n| Group | Class | `fix_` | `guard_` |\n| --- | --- | ---: | ---: |\n| proration | `ProrationHiddenTest` | 7 | 4 |\n| tax | `TaxHiddenTest` | 6 | 3 |\n| coupons | `CouponHiddenTest` | 5 | 6 |\n| refunds | `RefundHiddenTest` (run 3×) | 3 | 7 |\n| idempotency | `IdempotencyHiddenTest` (run 3×) | 3 | 5 |\n| line_rounding | `LineRoundingHiddenTest` | 4 | 3 |\n| coupon_dedup | `CouponDedupHiddenTest` | 4 | 3 |\n| tax_cache | `TaxCacheHiddenTest` | 3 | 3 |\n| webhook_redelivery | `WebhookHiddenTest` (run 3×) | 3 | 5 |\n| usage_timezone | `UsageTimezoneHiddenTest` | 4 | 4 |\n| regression | `RegressionHiddenTest` | 12 plain tests | |\n\n- **`fix_` tests** fail on the defective code and pass on the gold fix.\n- **`guard_` tests** pass on both, and fail if a \"fix\" breaks correct behavior in the same area.\n- **Regression tests** cover behavior that already works, including the usage-block rule and the correct supporting code (dunning, invoice numbering under load, credit notes, statements, CSV export, usage alerts, plan catalog, audit log).\n- Every hidden test works against the public API only. The hidden tests bring their own gateway and notifier test doubles and don't depend on the public test files.\n- The concurrency tests are deterministic where it matters: a blocking gateway or notifier holds the first request until a second one reaches it, or for one second. Correct code never lets the second one through, so it just waits; racy code always does. Barrier-started thread storms with injected latency cover the broader case.\n\n## 10. Reward calculation\n\n```\nreward = 0.90 × defects + 0.05 × regression + 0.05 × process          (max 1.0)\n\ndefect_<group>  = (passing fix_ tests / all fix_ tests)   if every guard_ test passes\n                = 0                                        otherwise\ndefects         = mean of defect_<group> over the ACTIVE defects\n                  × fraction of INACTIVE defect groups whose hidden tests all pass\nregression      = passing / total over the regression tests + every test of the inactive groups\nprocess         = 1 if compiled AND public tests unmodified AND ≥1 src/main file changed\n                  AND at least one active defect scored > 0, else 0\n```\n\n- **Benchmark tasks** (all ten active) have no inactive groups, so `defects` is the plain mean and the reward is exactly `Σ 0.09 × defect_<group> + 0.05 × regression + 0.05 × process`.\n- **Sampled tasks:** breaking an area that was already fixed costs reward twice. It scales down `defects` (one of nine areas broken → × 8/9), and it lowers `regression`.\n- If the code does not compile, every component is 0.\n- **Reference points** (measured):\n\n  | State | Reward |\n  | --- | --- |\n  | untouched repo | 0.05 (regression only) |\n  | cosmetic edit only (a comment) | 0.05 (`process` needs at least one defect credited) |\n  | gold fixes | 1.00 |\n  | benchmark task, gold fixes for 2 defects | 0.28 |\n  | sampled task, active defect fixed but an inactive one re-broken | 0.90 |\n\n- **Metrics** (recorded, not summed into the reward):\n  - `defect_<group>` for each active defect\n  - `compiled`, `public_tests_intact`, `main_files_changed`\n  - `active_defects`, `inactive_defects_intact`, `defects_fully_fixed`\n  - `defect_files_touched`: the number of active defect groups whose file changed. `tax` and `tax_cache` share `TaxService.java`, so one edit there can count twice.\n\n## 11. Gold solution and gold validation\n\n`assets/gold/` holds the reference versions of the 9 files that carry defects: `LineItem.java` and 8 services (Discount, Invoice, Payment, Refund, Subscription, Tax, UsageReport, Webhook). They overlay the task repo.\n\n`BillingTask.validate()` installs the gold files over the task repo and runs the **same** `grade()` pipeline a rollout uses: snapshot, fresh grading box, hidden suites. It passes only if the code compiles, **every active defect group scores 1.0**, and **regression is 1.0**. Regression includes every inactive group, so with gold all 10 groups must pass. The framework's validate CLI calls it and also runs a setup-only check:\n\n```bash\nuv run validate billing-payment-debug --runtime.type docker                               # 2 benchmark tasks → valid_rate 1.0\nuv run validate billing-payment-debug --runtime.type docker --taskset.num-sampled-tasks 6 # + 6 sampled tasks → valid_rate 1.0\n```\n\n`validate()` doesn't check `process`. The full maximum reward of **1.00**, including `process`, is asserted by the `gold` scenario of the self-check (§12).\n\n## 12. Self-check / anti-hacking validation\n\n`scripts/self_check.py` first runs `defects.verify()`, then runs 31 scenarios, each through the real grading path in Docker. For each one it installs the task's repo in a box (the benchmark task, or a sampled task for the `sampled-*` rows), applies the scenario's edits as if an agent had made them, calls `finalize` (fresh grading box and all) and `score`, and checks the rewards. It exits non-zero if any scenario misses its expectation.\n\n| Scenario | What it simulates | Asserted |\n| --- | --- | --- |\n| `buggy` | no changes | all defects 0, regression 1, process 0 |\n| `gold` | all gold fixes | all defects 1, regression 1, process 1 (total 1.00) |\n| `partial-refunds-coupons` | gold for 2 defects | those 2 at 1, the other 8 at 0, regression 1, process 1 |\n| `tamper-public-tests` | public tests edited to pass | all defects 0, process 0 |\n| `sleep-instead-of-lock` | `Thread.sleep` added before the idempotency check | all defects 0 |\n| `swallow-gateway-errors` | gateway exceptions caught and ignored | all defects 0 |\n| `coarse-synchronized-fix` | `synchronized` on `charge` (a correct fix) | idempotency 1, others 0 |\n| `refunds-reject-any-second-refund` | over-restrictive refund \"fix\" | refunds 0 |\n| `refunds-synchronized-method` | `synchronized` on `issueRefund` (a correct fix) | refunds 1 |\n| `refunds-lock-around-check-only` | lock around the balance check only; the processor call and save still race | refunds 0 |\n| `webhook-mark-after-apply-no-lock` | the obvious half-fix: mark events processed after applying them, with no lock | webhook_redelivery 0 |\n| `webhook-unmark-on-failure` | keep marking first, un-mark if applying fails (a correct fix) | webhook_redelivery 1 |\n| `webhook-synchronized-mark-after-apply` | `synchronized` `handle` that marks after applying (a correct fix) | webhook_redelivery 1 |\n| `noop-edit-farms-process` | comment-only edit to production code | process 0, total 0.05 |\n| `refunds-reject-while-busy` | reject a refund while another is in flight | refunds 0 |\n| `webhook-drop-dedup` | apply every delivery, with no deduplication | webhook_redelivery 0 |\n| `same-name-as-grader-class` | gold plus a stray class with the grader's test-runner name | ignored, all defects 1 |\n| `tax-double-epsilon-hack` | `Math.round(tax * 100 + 1e-6)` | tax 0 |\n| `tax-bigdecimal-valueof-double` | `BigDecimal.valueOf(double)` | tax 0 |\n| `forge-results-from-agent-code` | agent code prints fake `RESULT … PASS` lines for every hidden test via reflection, then halts the JVM | proration, tax, refunds and idempotency all 0 |\n| `does-not-compile` | gold plus a syntax error | everything 0 |\n| `squat-on-testkit-package` | gold plus a fake `testkit/Assertions.java` in `src/main` | the fake is ignored, all defects 1 |\n| `coupon-dedup-case-sensitive` | a dedup that ignores case differences | coupon_dedup ≤ 0.5 (observed 0.25) |\n| `delete-sources` | `rm -rf src` | everything 0, and scored rather than recorded as an error |\n| `symlink-sources-to-dev-zero` | gold plus symlinks to `/dev/zero` and `/proc/self/environ` inside `src/` | ignored, all defects 1 |\n| `oversized-snapshot` | gold plus a 9 MB file in `src/` | everything 0, scored rather than recorded as an error |\n| `too-many-source-files` | gold plus 501 extra `.java` files | everything 0, scored rather than recorded as an error |\n| `hang-every-money-call` | gold plus an infinite loop in `Money.round`, with a short test timeout and suite budget | total ≤ 0.35 (observed 0.05), scored rather than recorded as an error |\n| `sampled-untouched` | sampled task (tax, refunds, usage_timezone active), no changes | those 3 at 0, regression 1, total 0.05 |\n| `sampled-gold` | same task, gold for its 3 defects | defects 1, regression 1, total 1.00 |\n| `sampled-breaks-inactive-defect` | refunds-only task: refunds fixed, but proration (already fixed) re-broken | refunds 1, defects ≤ 0.9, total ≤ 0.9 (observed 0.89 / 0.90) |\n\nEvery scenario also asserts that the hidden tests ran as `nobody`. The red-team pass that added the last four scenarios, and the grader hardening that came with it, are written up in [`RED_TEAM.md`](RED_TEAM.md).\n\nThe `oversized-snapshot`, `too-many-source-files` and `hang-every-money-call` rows close a specific hole: turning a low score into a rollout *error*, which some trainers drop instead of penalizing.\n\n**Not exercised by the self-check** (the design covers these, but no test does):\n- tampering with the agent box's JDK, which is irrelevant with isolated grading,\n- adversarial fixes for `line_rounding`, `tax_cache` and `usage_timezone`,\n- grading with `isolated_grading=false`.\n\n## 13. Local validation commands\n\nAll commands run from this directory. They need Docker running and the image `eclipse-temurin:17-jdk`, which is pulled automatically.\n\n```bash\nuv sync                                                        # install from uv.lock\nuv run validate billing-payment-debug --runtime.type docker    # gold + setup validation → valid_rate 1.0\nuv run python scripts/self_check.py                            # all 31 scenarios → \"31/31 scenarios as expected\"\nuv run python scripts/self_check.py gold buggy -v              # a subset, with per-test failures\n\n# Public tests on the untouched task repo → \"21 passed, 2 failed\"\ndocker run --rm -v \"$PWD/billing_payment_debug/assets/task_repo:/src:ro\" eclipse-temurin:17-jdk \\\n  bash -c 'cp -r /src /tmp/repo && cd /tmp/repo && bash run_tests.sh'\n```\n\n## 14. Model evaluation commands\n\nThe eval CLI talks to any OpenAI-compatible endpoint. By default the client is Prime Inference (`PRIME_API_KEY` or `prime login`) and the runtime is Prime-hosted sandboxes.\n\n```bash\n# Local Docker sandboxes, any OpenAI-compatible provider\nexport MY_API_KEY=...\nuv run eval billing-payment-debug \\\n  -m <model-id> \\\n  --client.base-url <https://provider/v1> \\\n  --client.api-key-var MY_API_KEY \\\n  --env.agent.runtime.type docker \\\n  -n 2 -r 3            # 2 tasks (both variants) × 3 rollouts\n\n# Example: Ollama Cloud\nuv run eval billing-payment-debug -m nemotron-3-super \\\n  --client.base-url https://ollama.com/v1 --client.api-key-var OLLAMA_API_KEY \\\n  --env.agent.runtime.type docker -n 1 -r 1\n\nprime eval view -o outputs   # interactive viewer (vf-tui is deprecated)\n```\n\n- Each run writes `outputs/<env>--<model>--<harness>--<id>/traces.jsonl`, plus `configs/` and `logs/`.\n- Per-test pass/fail details for each rollout are in its trace under `info.grading`.\n- Finished runs are uploaded to the private Evaluations tab of your Prime account when you're logged in. Add `--no-push` to keep them local.\n\n## 15. Configuration\n\n**Requirements**\n- Python 3.11–3.13 and `uv`\n- Docker, for local runs, `validate` and the self-check\n- For the `prime` runtime or Prime Inference: `prime login` or `PRIME_API_KEY`\n- A model API key in an environment variable named by `--client.api-key-var`\n\n**Taskset config** (`--env.taskset.*`)\n\n| Field | Type | Default | Description |\n| --- | --- | --- | --- |\n| `variants` | list[str] | `[\"tickets\", \"minimal\"]` | Prompt styles used by benchmark and sampled tasks |\n| `include_full_tasks` | bool | `true` | Emit the all-ten-defects benchmark task per variant, listed first |\n| `num_sampled_tasks` | int | `0` | Extra tasks with a random active-defect subset and a random variant |\n| `min_active_defects` / `max_active_defects` | int | `1` / `5` | Subset-size range for sampled tasks (1–10) |\n| `sampling_seed` | int | `0` | Seed for the sampled task list |\n| `image` | str | `eclipse-temurin:17-jdk` | Image with JDK 17 on PATH and curl (for the harness bootstrap) |\n| `agent_timeout_seconds` | float | `1800` | Agent time budget |\n| `allow_network` | bool | `false` | Internet access while the agent works (harness setup always has it) |\n\n**Task config** (`--env.taskset.task.*`)\n\n| Field | Type | Default | Description |\n| --- | --- | --- | --- |\n| `isolated_grading` | bool | `true` | Grade in a fresh box. `false` grades in the agent's box after it finishes, which is cheaper but lets agent-side tampering reach the grader. |\n| `concurrency_repeats` | int | `3` | JVM runs of each concurrency group (`idempotency`, `refunds`, `webhook_redelivery`) |\n| `test_timeout_seconds` | int | `20` | Per-test timeout |\n| `suite_timeout_seconds` | int (≥ 30) | `900` | Budget for compiling and running all hidden suites. Runs that don't fit count as failed. |\n| `grading_timeout_seconds` | float | `1200` | Hard cap on the whole grading step (snapshot, grading-box boot, suites). Keep it above `suite_timeout_seconds`. |\n\n## 16. Known limitations and scope\n\n- **One codebase.** Sampled tasks vary which defects are active and how they're reported, not the code itself. A model trained long enough could memorize this repo's ten fixes. Up to 2,046 distinct tasks exist, but they're combinations of the same ten bugs.\n- **Hidden tests and gold are public on the Hub.** They ship inside the package (`billing_payment_debug/assets/`), as is usual for Hub environments, and `defect_manifest.md` ships in the sdist. They never enter the agent's sandbox, but a model could have seen them during pretraining.\n- **The agent's code shares a JVM with the hidden tests.** Agent code can halt or sabotage a run, which only lowers its own score. Forging a pass would require reading the signing key out of the grading shell's memory.\n- **Concurrency grading is partly timing-based.** The blocking-double tests are deterministic; the thread-storm tests rely on injected latency. Every concurrency group runs in 3 JVMs, and a test counts only if it passes all of them. Correct fixes passed and racy ones failed in every repeated local run.\n- **Grading time.** The blocking tests wait up to one second each on correct code, and three groups run 3 times, but a full gold grading run, including booting the fresh box, still took about 18 seconds locally.\n- **The tax tests include ledger-scale amounts** (about 10¹²) on purpose, to defeat `double`-based \"fixes\".\n- **Isolated grading boots a second sandbox per rollout.** That roughly doubles sandbox cost on hosted runtimes.\n- **What has been run.** Validated locally with Docker (`validate`, the self-check) and with real-model evals via `uv run eval` (e.g. nemotron-3-ultra). `prime eval run` does not work with this package on prime CLI ≤ 0.6.34, because it uses the legacy loader, which requires `load_environment()`. Use `uv run eval` instead. Prime-hosted sandboxes haven't been exercised yet.\n- **`[tool.verifiers.eval]`** in `pyproject.toml` is only read by the legacy verifiers eval script. v1 `eval` takes `-n` and `-r`.\n\n**Release.** `prime env push` publishes the package. The package's `verifiers>=0.3.1` requirement makes the CLI target the v1 runtime.\n","encoding":"utf-8","truncated":false,"total_bytes":29319},"status":null}