{"data":{"kind":"file","path":"README.md","version_id":"c7p30vjvje5lxny0jirsb91v","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":849,"modified_at":"2026-09-13T04:06:06.930000","content_hash":"6c7c05da6a1520a11973d56d63046a031c7f14d250b43c69f6b03b23f49a8906"},"entries":[],"content":"# mrcr-v2\n\nMRCR v2 long-context coreference tasks solved by an agent in a sandbox: each task uploads a long conversation transcript to the workspace, and the agent scans it to retrieve the requested content and writes its answer to a file. Answers are scored with the official MRCR v2 metric — a `difflib` `SequenceMatcher` ratio against the reference, gated on the 12-character hash prefix.\n\n## Taskset\n\n- **Source:** [google-deepmind MRCR v2 (GCS bucket `mrcr_v2`)](https://storage.googleapis.com/mrcr_v2)\n- **Size:** 310 tasks (default config: 8 needles, 1m-2m token bucket)\n\n## Changelog\n\n- 2026-08-31: Yield task records on demand so bounded evaluations construct only the requested prefix.\n- 2026-08-30: Keep transcript-search command output bounded so long-context tasks do not overflow the model request.\n- 2026-06-24: Initial v1 taskset.\n","encoding":"utf-8","truncated":false,"total_bytes":849},"status":null}