{"data":{"kind":"file","path":"README.md","version_id":"hv63cjpaib1v3k73iww11mf9","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":19593,"modified_at":"2026-09-24T10:10:03.746000","content_hash":"1160a9c3b94f4fb10e073c8fc6ea462f5828b69621543fb480f189e0dbf22535"},"entries":[],"content":"# tax-calculation\n\nA local Prime Intellect / Verifiers environment that evaluates whether an agent can investigate and repair a billing calculator across fictional tax jurisdictions. The initial implementation runs, but contains eight interacting business-logic defects. The task is source repair, not a question asking for one tax total.\n\n**Developer documentation:** this README, `developer/`, `tests/`, and `reports/` disclose evaluation design. They must not be mounted into an evaluated agent's filesystem or included in its prompt. The supported agent interface consists only of the four registered tools below.\n\n## Setup and local commands\n\nPython 3.11–3.14 is declared; local validation used CPython 3.14.7. The project remains pinned to **Verifiers 0.1.5**. A narrow compatibility adapter is also tested against **Prime CLI 0.6.35’s installed Verifiers 0.2.0** using its real evaluation lifecycle with an in-memory model transport. No dependency upgrade was needed. Live provider calls, the environment-server transport, other Verifiers releases, and the separate v1 API are not covered by that offline validation.\n\n```bash\nUV_CACHE_DIR=\"$PWD/.uv-cache\" uv sync --locked --extra test\n.venv/bin/python -m pytest -q --tb=short\n.venv/bin/python -m tax_calculation.local_eval\n.venv/bin/python -m developer.audit\n```\n\nThe default local evaluation runs the seeded implementation through `vf.load_environment(\"tax-calculation\")`, tool registration, actual `a_generate` rollout, and rubric scoring. It uses the real OpenAI client with `httpx.MockTransport` and scripted responses. **No model service, API key, cloud resource, or paid inference is used.** It is an integration test, not evidence that a real model has solved the task. Expected default reward: **0.0**.\n\n`developer.audit` evaluates the gold repair and adversarial candidates, tests the broken and gold framework rollouts, repeats gold verification three times, and writes `reports/results.json` and `reports/reward-hacks.md`. Run `.venv/bin/python -m developer.validate_upgrade` for all 256 repair subsets, independent defect detection, original-case preservation, state isolation, and replay of saved historical submissions; it writes evaluator-only evidence to `reports/upgrade-validation.json`. Expected gold reward: **1.0**. An optional candidate file can be tested with `tax-evaluate --source path/to/candidate.py`; the default `tax-evaluate` command is equivalent to the local module command above.\n\nBuild distributable packages locally:\n\n```bash\nUV_CACHE_DIR=\"$PWD/.uv-cache\" uv build\n.venv/bin/python -m developer.check_clean\n```\n\n`uv.lock` records all resolved transitive versions and hashes. `pyproject.toml` pins Verifiers, Datasets, Hatchling, and test dependencies. Package installation needs a registry or a populated cache; scoring itself uses only local code and standard-library arithmetic. No publication command is needed or has been run.\n\n## Prime CLI compatibility\n\nDataset rows contain `prompt`, an empty `answer`, and `info = {\"env_id\": \"tax-calculation\"}`. They omit `task`: Prime’s Verifiers 0.2.0 treats that field as a complete structured task payload, so the old route string is invalid. The same single public incident is explicitly registered as both the default and evaluation dataset; hidden cases remain entirely separate.\n\nThe adapter accepts legacy OpenAI nested tool calls and newer typed `vf.ToolCall` messages. It preserves every argument/security check. On Verifiers 0.1.5, `env_response` returns `(messages, state)`; on Prime’s 0.2.0, it returns typed messages only, with state mutated in place. `call_tool` returns a single message in both versions, so dispatch still uses `append`, not `extend`.\n\nRun the compatibility check with the Prime uv tool’s interpreter, from this folder:\n\n```bash\n\"$(uv tool dir)/prime/bin/python\" -m developer.prime_compat\n```\n\nThis checks the actual task normalizer, typed message handling, strict argument checks, source-candidate rewards, and `env.evaluate` with `MockTransport`. It does not contact a model provider. See [compatibility evidence](reports/compatibility.md).\n\nFor a real-model run, first set `OLLAMA_API_KEY` in your terminal, then run:\n\n```bash\nprime eval run tax-calculation \\\n  --model \"gpt-oss:20b\" \\\n  --api-client-type openai_chat_completions \\\n  --api-key-var OLLAMA_API_KEY \\\n  --api-base-url \"https://ollama.com/v1/\" \\\n  --num-examples 1 \\\n  --rollouts-per-example 1 \\\n  --max-concurrent 1 \\\n  --timeout 300 \\\n  --max-retries 1 \\\n  --disable-env-server \\\n  --disable-tui \\\n  --save-results \\\n  --skip-upload\n```\n\nThe real-model command is provided for the owner to run; it was not executed during this compatibility work. It contacts the configured Ollama cloud endpoint. The local offline command continues to use the project’s `.venv` and its pinned 0.1.5 API.\n\n## Scenario and data model\n\nA merchant fulfills individual lines in different jurisdictions. The customer billing jurisdiction can differ from any line's fulfillment jurisdiction. The agent sees a business incident and can inspect products, customers, carts, rules, and source before applying a repair.\n\nThe model is intentionally small but relational: jurisdiction/category rule tables; products with category, name, and unit price; customers with billing jurisdiction; carts referencing customers; line items referencing products with quantity and fulfillment jurisdiction; and invoice results containing ordered lines, subtotal, tax, and total. Dictionary keys enforce unique entity IDs, and validation rejects duplicate line IDs and invalid references. Structured in-memory records are used instead of SQLite because read queries and database mutation are unnecessary for this task.\n\nThere are five visible products and three visible carts: a mixed destination order, an exempt-product order, and a small-value order. Hidden cases introduce new product and cart identities, prices, quantities, ordering, cart sizes, and configured rates. Quantities range from 1 to 1000; carts have 1–24 lines. Price validation permits nonnegative plain decimal strings with at most two decimal places and a maximum unit price of 1,000,000. Rates are nonnegative plain decimals with at most six places, at most 1.0. No real-world tax law is used.\n\n### Fictional rules\n\n| Jurisdiction | Default rate | STANDARD | DIGITAL | GROCERY | MEDICAL |\n|---|---:|---|---|---|---|\n| STATE_A | 7.5% | taxable | taxable | exempt | exempt |\n| STATE_B | 5.25% | taxable | exempt | exempt | exempt |\n| STATE_C | 9% | taxable | taxable | taxable | exempt |\n\nRules are supplied at execution time. The rule table, rather than constants embedded in source, is authoritative; merchant fixtures can configure different rates and taxable flags per category.\n\n1. Use each line's fulfillment jurisdiction.\n2. Read taxability and rate for that jurisdiction and product category.\n3. Compute line net as unit price × quantity.\n4. Compute tax on the complete line net; exempt lines have zero tax.\n5. Round once per line to cents using `ROUND_HALF_UP`.\n6. Include exempt merchandise in subtotal, sum tax across all jurisdictions, include every line exactly once in input order (including repeated product IDs), and compute total = subtotal + tax.\n7. Preserve the full supplied rate precision, up to six decimal places. Rates are dimensionless and are not rounded to currency cents.\n\nPrices exclude tax. There are no discounts, shipping charges, refunds, currency conversion, minimum-tax thresholds, or compound taxes. Exempt lines and taxable lines with a zero configured rate both have zero tax; both remain in the invoice and subtotal. Amounts use fictional USD-like credits with two decimal places.\n\n## Agent task and tools\n\nThe exact prompt is `tax_calculation.tasks.TASK_PROMPT` and is the source of truth for the incident text. It describes customer symptoms without naming faulty expressions, hidden cases, expected totals, or the repair. Dataset rows contain no gold answer; the `answer` field is empty.\n\n| Tool | Parameters visible to model | Purpose / exposed information |\n|---|---|---|\n| `inspect_business` | `section: str` (required) | One of `rules`, `products`, `customers`, `carts`, `contract`; visible records or written business/language contracts |\n| `read_calculator` | none | Current source and supported repair language |\n| `replace_calculator` | `source: str` (required) | Atomically replace the calculation function after syntax validation; returns acceptance or bounded source-validation diagnostics; unexpected failures remain generic |\n| `calculate_cart` | `cart_id: str` (required) | Execute current source on a visible cart; return computed invoice or a generic error |\n\nThe trusted `state` parameter is injected by the framework adapter and removed from every tool schema. Model-supplied `state` arguments are rejected. Tools do not return hidden assertion results, hidden inputs, expected outputs, verifier source, baseline hashes, or a gold patch. There is no SQL, shell, file-access, or arbitrary execution tool.\n\nThe typical flow is: inspect the contract and tables, reproduce invoice errors, inspect implementation, replace source, rerun visible calculations, then finish with an explanation. Only the **last successfully stored calculation source** is scored. Rejected source updates leave the previous source intact. Syntactically valid code that fails during execution is stored and fails verification safely.\n\n## Repair language and invoice contract\n\nRepairs replace `def calculate(cart, products, rules):` using an explicitly interpreted Python subset. It supports local assignments, `for` over lists/dictionaries, `if/else`, `return`, dictionary/list/tuple literals (tuple literals behave as lists), indexing, arithmetic `+ - * /`, comparisons, boolean expressions, and conditional expressions. Decimal values use `D(\"0.00\")`; `money(value)` rounds half up to cents, and `round_even(value)` rounds half even. `len(collection)` is also available. Lists can be extended with `lines = lines + [entry]`.\n\nImports, attributes, method calls, comprehensions, assignment to input fields, arbitrary calls, recursion, `while`, floating-point literals, and exponentiation are rejected. This is **not unrestricted Python**. Candidate text is never executed using Python `eval`, `exec`, or `compile`. The interpreter receives only copies of cart, products, and rules, with fresh locals and Decimal context for each invocation.\n\nThe returned invoice must have exactly:\n\n```text\n{cart_id: str, lines: list, subtotal: Decimal, tax: Decimal, total: Decimal}\nline = {id: str, jurisdiction: str, net: Decimal, tax: Decimal, total: Decimal}\n```\n\nMonetary outputs must be finite, nonnegative Decimal values rounded to cents. Tool responses serialize them as fixed two-place strings. Missing/extra fields, wrong types, duplicate dictionary keys, malformed lines, fractional cents, NaN, and infinity fail safely. Equivalent calculations need not match the reference source text.\n\nThere is **no final-answer parser**. Empty responses, explanatory text, malformed JSON prose, duplicate JSON fields in final text, fake `PASS` claims, and long final messages cannot change reward. Tool arguments are separate: they must be JSON objects with no duplicate fields and valid function parameters.\n\n## Verification and developer validation\n\nThe seeded source contains exactly eight intentional business-logic defects.\nThe evaluator-only [final report](reports/final-report.md) records their coverage\nand independent detection. The gold repair and attack generators are excluded\nfrom wheel and source distributions and are never imported by the loader or\nmodel tools. Do not supply developer reports to an evaluated model.\n\nThe trusted verifier uses independent fixed-point arithmetic and compares the\ncomplete invoice, including ordered line identities, jurisdictions, component\namounts, and totals. It runs 146 hidden invoices: all 116 original cases plus\n30 additional cases. Ten verification groups cover the documented business\ncontract and generalization to unseen records and runtime rules.\n\nReward is bounded in `[0, 1]`. Full correctness receives exactly `1.0`;\nincomplete repairs receive at most `0.45`. Fixture-integrity violations and\ntrusted verification failures receive zero. Candidate execution errors fail\nthe affected cases. Only persisted calculation source is scored; final prose\ncannot change reward. Tools never return hidden verification results.\n\nBusiness records are checked against a fresh trusted baseline before and after\nverification. Hidden inputs are generated from trusted code independently of\nsession state. The current measured results are 340 passing internal tests,\ngold 146/146 with reward 1.0, and the seeded source 8/146 with reward 0.0.\nAll eight single-defect omissions are detected. All 255 incomplete repair\nsubsets and all tested incorrect/adversarial candidates score below 1.0.\n\n## Visibility, isolation, and limitations\n\n| Boundary | Contents |\n|---|---|\n| Model-visible | Incident prompt, tool schemas, visible business tables, business/language contracts, current source, visible calculation results |\n| Model-writable | Calculation source through the replacement tool; interpreter-local variables |\n| Trusted | Adapter, AST interpreter, fixture factory, integrity comparison, integer oracle, reward function |\n| Hidden | Generated verifier cases and seeds, expected results, developer gold and attacks, reports and tests |\n\nThe tested boundary is an **application-level capability boundary**, not OS/container isolation. Candidate programs have no mechanism in the supported language to read files, import packages, inspect objects, mutate input records, or reach host state. Automated attacks exercise these prohibitions and inspect serialized model requests for known hidden markers.\n\nTrusted and developer files still exist on the evaluator host. A process with filesystem/Python access could read or change them. Do not add a shell tool, expose the repository as an agent workspace, or supply this README to the evaluated model. Arbitrary-Python or shell-agent variants require a separate locked-down process/container and a distinct trusted scoring service. The interpreter and dependencies have not received an independent security audit. Finite attack tests cannot prove that every future exploit or every possible hardcoded program is impossible.\n\n## Resource controls\n\nEnforced by this environment:\n\n- 30 model turns by default; configurable from 2 through 100. The pinned framework stops before executing tool calls on the terminal turn, so repairs must be submitted earlier.\n- At most 8 tool calls per batch and 160 dispatchable calls per rollout.\n- JSON tool arguments at most 32,000 UTF-8 bytes; source at most 16,000 bytes.\n- At most 2,500 AST nodes and AST depth 60.\n- 30,000 interpreter steps per calculation; no unbounded `while`, recursion, or arbitrary calls.\n- Collections at most 256 entries; expanded nested values at most 4,096 nodes/depth 32; strings at most 16,000 characters; numeric magnitude at most `10**15`.\n- Validated output schema and bounded line counts/identifiers.\n\n**Enforced by runner, not by environment:** model request timeout and retries, maximum generated tokens, wall-clock rollout timeout, process memory/CPU quotas, concurrency admission, and host isolation. This environment does not set OS resource limits or a hard wall-clock tool timeout. Step/value limits reduce candidate computation but are not a claim of a process-level quota. The offline client uses zero retries. A network failure during an actual model request remains a runner/framework concern; scoring never needs a network service.\n\n## Reset, separation, and reproducibility\n\n`setup_state` installs a fresh session, deep independent fixture records, the original source, and a zeroed tool counter for each rollout. Sessions share no candidate state. Rejected patches are atomic; repeated verification does not change records or source.\n\nThere is one public debugging incident and **no traditional training/evaluation dataset split**. Public diagnostic carts and hidden verifier inputs are the separation. Hidden product/cart IDs are disjoint from public IDs; hidden cases use a fixed local seed and are not provided as dataset rows, tool results, or state snapshots in model requests. Line-order tests intentionally permute equivalent business combinations. Do not train on developer reports or release hidden seeds/gold to the evaluated model.\n\nReproducibility comes from the lockfile, fixed fixtures/seeds, integer oracle, controlled Decimal context, deterministic reset, and a single scored final state. There is no best-of-N selection. Real-model stochasticity is runner-owned. Tests validate on the local Python/platform only; other supported Python versions need CI confirmation before a release.\n\n## Layout and evidence\n\n```text\ntax_calculation/  loader, environment adapter, fixtures, task, tools, interpreter,\n                  state, initial source, hidden cases, verifier, offline harness\ndeveloper/        gold repair, adversarial candidates, acceptance and compatibility audits\ntests/            fixture, runtime, verifier, and framework integration tests\nreports/          evaluator-only rewards, attack tables, test/build logs, final audit\npyproject.toml    pinned direct dependencies, build configuration, CLI entrypoint\nuv.lock           resolved dependency versions and hashes\n```\n\nSee [reward-hacking results](reports/reward-hacks.md), [machine-readable results](reports/results.json),\n[acceptance evidence](reports/upgrade-validation.json), [README claim audit](reports/readme-audit.md),\nand the [final report](reports/final-report.md). Files under `reports/baseline/`\nand older compatibility/build logs are historical evidence; use the final report\nand `upgrade-*` logs for this revision. Packaging, offline rollout, tool dispatch,\nrubric registration, reset, malformed inputs, integrity gates, and alternative\nvalid repairs are tested.\n\nBefore Hub submission: confirm the target Hub/framework version, test its packaging/runner contract, choose process isolation and model-resource settings, review hidden asset handling, and calibrate actual model difficulty only with explicit authorization for any paid inference. Nothing has been committed, pushed, published, or trained in the cloud.\n\nArchitecture references: the official [Verifiers environment guide](https://docs.primeintellect.ai/verifiers/environments) and [API reference](https://docs.primeintellect.ai/verifiers/reference). Implementation compatibility is established by the pinned installed source and local tests, rather than assumed from the latest documentation.\n\n### Validation feedback\n\nRejected replacements leave the stored calculator unchanged. Only `{\"updated\":true}`\nconfirms installation; acceptance does not establish tax correctness. Expected\nvalidation failures return a fixed message and one of `syntax_invalid`,\n`unsupported_statement`, `unsupported_expression`, `unsupported_helper`,\n`invalid_signature`, `invalid_assignment`, `source_too_large`, or\n`source_too_complex`. The first failing validator check determines the category.\nSyntax diagnostics may include a one-based line and column; unsupported AST\nsyntax may include a line. Responses contain no source excerpts or raw exception\nmessages and are at most 256 UTF-8 bytes. Unexpected failures remain generic.\n\nFor future real evaluations intended for analysis, use `--save-results` together\nwith `--skip-upload`: skipping upload alone does not request saved results.\n","encoding":"utf-8","truncated":false,"total_bytes":19593},"status":null}