{"data":{"kind":"file","path":"README.md","version_id":"lgk009dlcv6sz91vf5akyqf0","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":1550,"modified_at":"2026-08-23T22:38:04.220000","content_hash":"08c3fd1f83019f8af2d65981d27e9b9601563bfc617bcd22e7853c441ebb6df6"},"entries":[],"content":"# billable-fraction\n\n### Overview\n- **Environment ID**: `referee-lab/billable-fraction`\n- **Short description**: What percent of the frames are usable footage? The number a buyer bills on.\n- Part of the **referee-lab capture-QA family** (see also `referee-lab/hands-visible`).\n\n### Task\nSingle-turn, multimodal. Gold by construction: a known number of frames per clip were corrupted (blackout, Gaussian blur, blown exposure); usable% = untouched/12. Scored correct within +/-10 points.\n\n### Quickstart\n```bash\nprime eval run referee-lab/billable-fraction\nprime eval run referee-lab/billable-fraction -m <model> -n 10 -r 1\n```\n\n### Metrics\n| Metric | Meaning |\n| ------ | ------- |\n| `reward` / `correct_answer` | 1.0 when the predicted percentage is within +/-10 of gold |\n\n### Baselines\n| Model | Accuracy | Scored | Errors |\n| ----- | -------- | ------ | ------ |\n| gemma3:27b | 40.0% | 30 | 0 |\n| qwen2.5vl:7b | 26.7% | 30 | 0 |\n\nTemperature 0, one pass, run on local hardware (GX10).\n**Read:** both models score BELOW the 43.3% constant-answer ceiling — numeric usable-footage estimation is a real capability gap today.\n Constant-answer\nceiling per validate_family.py output — a useful model must beat it.\n\n\n### Provenance\nAll media derives from DROID (CC BY 4.0) via lerobot/droid_1.0.1 (some items reuse\nthe graded hands-visible corpus). Per-item provenance is machine-readable in\n`assets/dev/manifest.jsonl`; credits in [ATTRIBUTION.md](ATTRIBUTION.md).\nLabels are by construction or from measured model runs — no human grading claimed.\n","encoding":"utf-8","truncated":false,"total_bytes":1550},"status":null}