{"data":{"kind":"file","path":"README.md","version_id":"dulzec9xq17te4op4yknm30g","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":813,"modified_at":"2026-09-13T04:06:06.929000","content_hash":"3fab36b3e2491324c4861be55ca8f5e669ceef5ede6ddd8aa66777a12b916293"},"entries":[],"content":"# longcot-mini\n\nLongCoT-mini long-horizon reasoning tasks spanning five domains (logic, cs, chemistry, chess, math), each a self-contained multi-step task solved by an agent in a sandbox. The agent writes its final answer to `/workspace/answer.txt` and is scored against gold by the upstream `longcot.verify` template dispatch (plus a local math numeric-equivalence fallback), yielding a rule-based `correct` reward.\n\n## Taskset\n\n- **Source:** [LongHorizonReasoning/longcot](https://github.com/LongHorizonReasoning/longcot)\n- **Size:** ~500 tasks (the upstream `easy` difficulty split across all five domains, minus 21 broken easy-math IDs excluded by default)\n\n## Changelog\n\n- 2026-08-31: Yield task records on demand so bounded evaluations construct only the requested prefix.\n- 2026-06-24: Initial v1 taskset.\n","encoding":"utf-8","truncated":false,"total_bytes":813},"status":null}