{"data":{"kind":"file","path":"README.md","version_id":"s41pd72x99yh706fv3vnftl8","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":4430,"modified_at":"2026-09-29T15:19:45.055000","content_hash":"26cf92f2773d74a6683638c6d1e1fda728e03d7acfb7ae0500cc77ed08eb3343"},"entries":[],"content":"# cad-spec (environment package)\r\n\r\nRL reward environment: write CadQuery for a dimensioned mounting plate; the\r\nbuilt solid is **measured** and scored against 9 requirements behind 5\r\nanti-cheat gates. Source, evidence and limits:\r\n[github.com/azzbilal/cad-spec](https://github.com/azzbilal/cad-spec#readme).\r\n\r\n## What it has shown so far\r\n\r\n![16-model ranking](https://raw.githubusercontent.com/azzbilal/cad-spec/main/results/leaderboard/ranking.svg)\r\n\r\n- **16 models, one greedy answer per spec, 30 held-out specs per tier.** The\r\n  best (deepseek-chat-v3-0324) passes every requirement on 78% of answers\r\n  [72, 85]; 11 of the 16 are below 35%. No model clearly beats a regex\r\n  template parser (75%).\r\n  [Full leaderboard](https://github.com/azzbilal/cad-spec/blob/main/results/leaderboard/leaderboard.md)\r\n- **Knowledge or reasoning?** A pre-registered experiment: seven lines of\r\n  general CadQuery facts (no spec numbers) lift mistral-small-3.2-24b from\r\n  32% to 83% and gemma-3-27b from 12% to 68%, and remove stacked drilling\r\n  entirely. Reasoning failures (a change order not carried into the hole\r\n  pitch, a margin counted twice) move by less than 5 points with either the\r\n  cheat-sheet or error feedback, in all ten model-arm pairs.\r\n  [Pre-registration and outcome](https://github.com/azzbilal/cad-spec/blob/main/docs/experiments/hint-feedback.md)\r\n- **The scorer is validated against itself:** 1,230 mutants of correct\r\n  answers, 0/600 false full credit, 0/600 false rejection.\r\n  [Validation report](https://github.com/azzbilal/cad-spec/blob/main/results/scorer-validation-0.4.0.md)\r\n\r\n## Load\r\n\r\n```python\r\nfrom cad_spec import load_environment\r\n\r\nenv = load_environment()                                  # L0 template task, as in 0.2\r\nenv = load_environment(tier=[\"L2\", \"L4\"], hints=True)     # cheat-sheet in the system prompt\r\nenv = load_environment(tier=[\"L1\", \"L2\", \"L3\"],           # train on harder tiers,\r\n                       eval_tier=[\"L0\", \"L1\", \"L2\", \"L3\", \"L4\"])  # evaluate on all\r\nenv = load_environment(tier=[\"L2\", \"L4\"], hints=True,     # binary training reward:\r\n                       reward=\"binary\")                   # 1.0 only if all nine pass\r\n```\r\n\r\n| Argument | Default | Meaning |\r\n|---|---|---|\r\n| `tier` | `\"L0\"` | training tier(s): L0 template, L1 table, L2 derive pitch, L3 prose, L4 change order |\r\n| `eval_tier` | same as `tier` | eval tier(s); rows carry `info[\"tier\"]` so results split per tier |\r\n| `metrics` | `True` | zero-weight diagnostics `built`, `gates_passed`, `m_R1_length` ... `m_R8_z_datum` |\r\n| `hints` | `False` | append the packaged CadQuery cheat-sheet (`cad_spec/hints.md`) to the system prompt of training AND eval rows; byte-identical to the hint arm of the experiment |\r\n\r\n200 train and 30 held-out eval specs per tier. L3 eval prompts use wording\r\ntemplates that never appear in train. A further **locked test split** (60\r\nspecs, `cad_spec.tasks.make_test_split`, pinned by fingerprint) is never\r\nloaded by the environment: it is reserved for one final comparison after\r\ntraining.\r\n\r\nWhy `hints`: the experiment showed that missing API knowledge is fixed by a\r\nprompt, not by training. With the cheat-sheet in the prompt, the reward\r\nconcentrates on what prompting cannot fix: carrying a change order through,\r\nderiving a dimension from a stated margin.\r\n\r\n## Reward\r\n\r\n```\r\n1.0    all 9 requirements met\r\nk/9    partial compliance, gates permitting\r\n0.05   code builds a solid but fails a gate or meets nothing\r\n0.0    code does not build a solid, times out, or is absent\r\n```\r\n\r\nScorer version: `cad_spec.rubric.SCORER_VERSION` (0.4.0). Scores from\r\ndifferent scorer versions are not comparable.\r\n\r\n## Execution\r\n\r\nModel code only hands back BREP geometry, which a trusted process measures.\r\nEach rollout runs in a forked, rlimited, env-scrubbed child on Linux/macOS\r\n(`CAD_SPEC_SANDBOX=fork`, about 60 ms overhead); Windows uses a persistent\r\nworker (`reuse`) that does not isolate state between rollouts.\r\n`CAD_SPEC_EXEC_TIMEOUT` (default 10 s) and `CAD_SPEC_MEM_MB` (default 2048)\r\nbound each rollout. See `SECURITY.md` in the repository before running\r\nuntrusted output outside Prime's sandbox.\r\n\r\n## Scorer-only use\r\n\r\n`import cad_spec` does not import verifiers. With only cadquery installed:\r\n\r\n```python\r\nfrom cad_spec.rubric import score\r\nfrom cad_spec.tasks import TASKS\r\nprint(score(open(\"answer.py\").read(), TASKS[0]).summary)\r\n```\r\n","encoding":"utf-8","truncated":false,"total_bytes":4430},"status":null}