{"data":{"kind":"file","path":"README.md","version_id":"qtnl3fsb2nnlbl5ys8hfg2j6","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":945,"modified_at":"2026-09-13T04:06:06.929000","content_hash":"983aabf3a3bc5770ca5c2aedf3150f15a04b3cb31522d55fd3374e1e491e616c"},"entries":[],"content":"# longcot_env\n\nLong-horizon reasoning tasks spanning five domains (logic, cs, chemistry, chess, math) at medium and hard difficulty. Each self-contained prompt embeds the full task; the agent writes its final answer to `/workspace/answer.txt`, which is scored via the upstream `longcot.verify` template dispatch (with a local numeric-equivalence fallback for math templates).\n\n## Taskset\n\n- **Source:** [LongHorizonReasoning/longcot](https://github.com/LongHorizonReasoning/longcot)\n- **Size:** ~2,000 tasks (medium + hard across all five domains), loaded from the bundled JSON in the `longcot` package via `load_questions`\n\n## Changelog\n\n- 2026-09-03: Restore default solver network access by reverting the `network_allow=[]` default-deny policy introduced in #780; training rollouts need outbound network.\n- 2026-08-31: Yield task records on demand so bounded evaluations construct only the requested prefix.\n- 2026-06-24: Initial v1 taskset.\n","encoding":"utf-8","truncated":false,"total_bytes":945},"status":null}