{"data":{"kind":"file","path":"README.md","version_id":"gcbfozir6veb4tino7w4qyze","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":1936,"modified_at":"2026-08-24T16:46:25.876000","content_hash":"6b15368a0eb9e8712182cef1c9d8e4c81e4cb0250091ea0e353fd137aadaba6c"},"entries":[],"content":"# clip-twin\n\n### Overview\n- **Environment ID**: `referee-lab/clip-twin`\n- **Short description**: Same footage twice (a resell/duplicate, however disguised) or genuinely different clips?\n- Part of the **referee-lab capture-QA family** (see also `referee-lab/hands-visible`).\n\n### Task\nSingle-turn, multimodal. Gold by construction: twins are programmatic transforms (brightness, crop, mirror, re-encode noise) of the same clip; non-twins are different clips from the same camera domain (so scene similarity alone cannot answer).\n\n### Quickstart\n```bash\nprime eval run referee-lab/clip-twin\nprime eval run referee-lab/clip-twin -m <model> -n 10 -r 1\n```\n\n### Metrics\n| Metric | Meaning |\n| ------ | ------- |\n| `reward` / `correct_answer` | 1.0 on exact match with the constructed/measured gold |\n\n### Baselines\n| Model | Accuracy | Scored | Errors |\n| ----- | -------- | ------ | ------ |\n| gemma3:27b | 100.0% | 30 | 0 |\n| qwen2.5vl:7b | 100.0% | 30 | 0 |\n\nTemperature 0, one pass, run on local hardware (GX10).\n**v0.2 (this version): the eval regrew itself and un-saturated.** The corpus was\nregenerated by the self-hardening loop (`tools/harden_clip_twin.py`): subtle\ngraded-strength twins plus SAME-EPISODE, DISJOINT-TIME hard negatives that kill\nscene-matching shortcuts. Probe results on the new corpus: gemma3:27b 5/40\n(12%), qwen2.5vl:7b 0/40 (0%) — down from 100%/100% on\nv0.1. Every item sits at the measured difficulty frontier (>=1 fleet model wrong).\nGold remains by construction; no human labels.\n Constant-answer\nceiling per validate_family.py output — a useful model must beat it.\n\n\n### Provenance\nAll media derives from DROID (CC BY 4.0) via lerobot/droid_1.0.1 (some items reuse\nthe graded hands-visible corpus). Per-item provenance is machine-readable in\n`assets/dev/manifest.jsonl`; credits in [ATTRIBUTION.md](ATTRIBUTION.md).\nLabels are by construction or from measured model runs — no human grading claimed.\n","encoding":"utf-8","truncated":false,"total_bytes":1936},"status":null}