{"data":{"kind":"file","path":"README.md","version_id":"ozmd21wwq57xk6mz3ddgtxif","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":9906,"modified_at":"2026-10-09T16:54:49.163000","content_hash":"b5fd557aa169c9f1a9c07edde716c69e07314a6c1ade8abac95d14544572fdda"},"entries":[],"content":"# periapical-lesions\n\nA periapical reporting worklist. The agent is the radiographer on shift at an offline reporting clinic: it reads the referral, reports every periapical lesion on the dental pantomograph with its Periapical Index grade (3, 4 or 5), and files the report with the clinic's `radreport` tool. Tasks come from [tirandazdylan/periapical-lesions-pai](https://huggingface.co/datasets/tirandazdylan/periapical-lesions-pai) (pinned revision; train 3,087 / validation 389 / test 387). The source publishes no patient IDs, so the splits are made per film, not per patient: the dataset keeps only the source's original (not augmented) films, drops near-identical ones, and puts look-alike films (average hash within 30 of 512 bits) in the same split. Two films of one patient that do not look alike can still fall in different splits.\n\n## The workspace\n\n`setup` writes four files into the sandbox: `pantomograph.jpg` (the film), `referral.txt` (the referring clinician's note: patient code, age, sex, reason — fixed per film, and free of clinical hints), `findings-codebook.md` (the PAI grades, the report format, the filing command), and `radreport.py` (the clinic's reporting tool). The prompt is the worklist entry: it names no grades and defines no format.\n\nThe agent writes `report.json` and validates it with `python3 radreport.py report.json`. After the agent stops, the environment scores the final file; chat replies are not scored. The validator does not create a separate submission record, and scoring does not require proof that it ran. A missing report counts as no findings. Version 0.3.4 changes only the licence to Apache-2.0 (0.3.1 to 0.3.3 carried a commercial licence). Version 0.3.3 scores a lesion-free film as 1 for an empty report and 0 for any finding (earlier versions raised an error; no film in the pinned dataset is lesion-free), matches boxes at IoU ≥ 0.3 exactly (an epsilon in the IoU denominator had required slightly more), and makes the validator reject NaN and infinite box values, which the scorer already dropped. Version 0.3.2 reads the filed report through the runtime's file API instead of a guest shell command, and applies the validator's types when scoring: boxes must be real numbers and grades whole numbers (4.0 counts, 3.9 and booleans do not). Version 0.3.1 corrected the report command; 0.3.0 had omitted `.py`. Historical results below identify their original setup and are not new benchmarks; rescoring every saved filed report (399 across 12 runs) with the 0.3.3 scorer changes no score.\n\nReward = 0.5 × `lesion_f1` (boxes matched one-to-one at IoU ≥ 0.3) + 0.5 × `pai_f1` (the same F1, counting only matches with the reference grade). Both count extra boxes against the answer, so a label-free answer (two fixed boxes graded PAI 3) scores about 0.04 on the test set. Before 0.2.0 the grade term was the share of reference lesions found and graded correctly, which let a dense grid of boxes score 0.31; every number below is rescored from the saved traces, and no pass changed.\n\nTasks run in a sandbox with no network, like the reporting workstation. The image (`scripts/periapical-lesions/Dockerfile`) ships numpy and Pillow. The config runs a coding agent (`claude_code` harness, up to 50 steps, web tools removed) that crops and enhances the film with Python and looks at the results; `bash` cannot show image crops back to the model.\n\nSandbox image: `prime/dylantirandaz/periapical-lesions:v1` (`scripts/periapical-lesions/Dockerfile`), Prime image id `x4jknf0dsroftl9fqsbpg5b4`, 106,795,139 bytes, pushed 2026-09-25; Prime's registry reports no content digest, so the id is the pin.\n\n## Agent runs on the 0.3.0 worklist\n\nThe first 200 test films, 1 try each, environment 0.3.0 and Claude Code 2.1.280. The current config pins this harness version. The old report command needed correction by the agent; these results predate the command fix.\n\n| Model | n | Reward (95% CI) | `lesion_f1` | `pai_f1` | Pass (reward = 1) |\n| --- | --- | --- | --- | --- | --- |\n| anthropic/claude-opus-5.5 | 200 | 0.12 (0.08-0.15) | 0.16 | 0.07 | 3.5% |\n| anthropic/claude-haiku-4.5 | 200 | 0.00 | 0.00 | 0.00 | 0% |\n\nIntervals are 95% bootstrap intervals over films. On the same 200 films, the earlier chat-answer setup scored Opus 5.5 at 0.19 and Haiku 4.5 at 0.00. In the worklist run, Opus wrote an empty report on 129 films, found 42 of 338 reference lesions, and gave the correct grade to 21 of those 42. It scored 1 on seven films. These counts do not establish clinical reliability or the cause of the difference between setups. The earlier trace audit checked score recomputation and image/reference identity, but it did not detect the incorrect report command.\n\nAgent runs on the chat-reply harness (0.2.0 and earlier; the config above, first n test films, 1 try each):\n\n| Model | n | Reward | Pass (reward = 1) |\n| --- | --- | --- | --- |\n| openai/gpt-6-astra | 25 | 0.26 | 4% |\n| anthropic/claude-opus-5 | 3 | 0.17 | 0% |\n| anthropic/claude-fable-5.1 | 3 | 0.17 | 0% |\n\nOn the same 25 films, single calls from GPT-6 Astra scored reward 0.21 and pass 2.7% (3 tries each); on the first 3, its agent scored 0.28. Claude Fable 5.1 ran with `--env.agent.harness.version 2.1.251`, since its API rejects the pinned Claude Code 2.1.232.\n\nThe results below are single model calls (`null` harness, no tools, unless listed as `bash`).\n\n| Model (test, 1 rollout) | Harness | n | Reward | Pass (reward = 1) |\n| --- | --- | --- | --- | --- |\n| openai/gpt-6-astra | null | 387 | 0.19 | 5.7% |\n| google/gemini-3.8-flash | null | 20 | 0.08 | 0% |\n| google/gemini-3.1-pro-preview | null | 100 | 0.07 | 0% |\n| openai/gpt-5.6-sol | null | 20 | 0.06 | 0% |\n| anthropic/claude-fable-5.1 | null | 20 | 0.05 | 0% |\n| anthropic/claude-opus-5 | null | 387 | 0.03 | 0.3% |\n| openai/gpt-5.6-terra | null | 20 | 0.00 | 0% |\n| anthropic/claude-sonnet-5 | null | 20 | 0.00 | 0% |\n| openai/gpt-4.1-mini | bash | 50 | 0.01 | 0% |\n| openai/gpt-5.4-mini | null | 50 | 0.00 | 0% |\n| anthropic/claude-haiku-4.5 | null | 50 | 0.00 | 0% |\n| qwen/qwen3-vl-235b-a22b-instruct | null | 50 | 0.00 | 0% |\n| alibaba/qwen3.5-flash | bash | 50 | 0.00 | 0% |\n\nCurrent prompt, single model calls, first 100 test films, 3 attempts each:\n\n| Model | Reward | pass@1 | pass@3 | Solved 3/3 |\n| --- | --- | --- | --- | --- |\n| openai/gpt-6-astra | 0.20 | 4.0% | 8% | 2 |\n| anthropic/claude-fable-5.1 | 0.07 | 1.0% | 3% | 0 |\n| openai/gpt-5.6-sol | 0.06 | 1.3% | 4% | 0 |\n| anthropic/claude-opus-5 | 0.03 | 0.0% | 0% | 0 |\n\nThe single-call table predates the prompt's PAI grade definitions. On the first 100 test films the definitions left GPT-6 Astra unchanged within noise (reward 0.22 vs 0.23 and 0.25 in two earlier runs; 6 vs 7 passes).\n\nSource inter-annotator result: the dentists annotated 59 films twice; scored against each other, the annotations reach reward 0.83 and pass 61%. This is not a hard upper bound on model performance. Grades use exact matching; with grades 3-5, a tolerance of one grade would always accept grade 4.\n\nSeparate training experiments used Tinker with LoRA rank 32. The model saw the tooth-bearing area as 3×2 overlapping tiles at 2× zoom, with six calls merged into one answer per film. It was fine-tuned on the train films, then trained with RL on 500 held-out train films. Checkpoints were selected on validation, and the test split was scored once with greedy decoding. These training scripts and weights are not part of this environment package:\n\n| thinkingmachines/Inkling-Small, tiled harness (test, 387 films) | Reward | Pass (reward = 1) |\n| --- | --- | --- |\n| fine-tuned | 0.12 | 2.3% |\n| fine-tuned + RL (50 steps) | 0.18 | 3.1% |\n\nRL added 0.06 reward (paired 95% CI 0.03 to 0.08) and raised the share of lesions found from 20% to 33%. This is a different harness from the rows above. Earlier RL runs from base Inkling scored the test split during training; they are not reported.\n\n## Run\n\nFrom the repository root. The config and Dockerfile are outside the Prime Hub source download:\n\n```bash\nuv sync --locked\nuv pip install -e environments/periapical-lesions\n.venv/bin/eval @ configs/periapical-lesions/eval.toml -m <model> -n 1 -r 2\n```\n\n`--env.taskset.split` selects `train`, `validation` or `test` (default).\nThis requires a supported model endpoint, Prime credentials, and access to the configured sandbox image.\n\n## Tests\n\n```bash\n.venv/bin/python -m pytest -q environments/periapical-lesions/tests\n```\n\nCovers the report parser (fenced reports, prose around them, drafts, malformed boxes and grades), the scorer\n(F1 arithmetic, the IoU threshold, one-to-one matching, extra boxes), the referral generator, the `radreport` tool,\nand the rule that only the final report file is scored. The reporting regression executes the prompt's command\nthrough the real host runtime and checks the reward. Run these tests from the repository root after installation.\nThey test host-side behavior, not model quality.\n\nVersion 0.3.3 was also checked in a real Prime sandbox with `anthropic/claude-haiku-4.5` and\nClaude Code 2.1.280. On test film `00011`, the agent ran `python3 radreport.py report.json`,\nvalidated a two-box report, and received reward 0. The prompt image and reference labels\nmatched the pinned dataset; the recorded scores matched scores calculated again from the\nfinal report. This one-film execution check does not measure model accuracy. The file-API\nreport collection was checked in a sandbox without a model: a guest-replaced `cat` does not\nchange the collected report.\n\n## Licence\n\nThe code, including `configs/periapical-lesions/` and `scripts/periapical-lesions/`, is licensed under the\nApache License, Version 2.0. See `licenses/LICENSE` in this directory.\n\nThe films and lesion labels are third-party data under CC BY 4.0. See `licenses/NOTICE` for attribution.\nThe data is not exclusive and can be obtained from the source.\n","encoding":"utf-8","truncated":false,"total_bytes":9906},"status":null}