{"data":{"kind":"file","path":"README.md","version_id":"s5zkfdf6v8zwtpf6wi46tke2","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":930,"modified_at":"2026-09-13T04:06:06.953000","content_hash":"1a365f4eff62c1543781ca7ebe669b41933f6125d2573b66a2b271da8eaec32d"},"entries":[],"content":"# s1-deepresearch\n\nS1-DeepResearch closed-ended deep-research QA: the agent answers in chat (bring\nyour own web search) and is graded by a binary LLM judge with [CORRECT]/[INCORRECT]\nsemantics. Only the verifiable \"Closed-ended Multi-hop Resolution\" rows with a\nnon-empty gold answer are kept.\n\n## Taskset\n\n- **Source:** [`ScienceOne-AI/S1-DeepResearch-15k`](https://huggingface.co/datasets/ScienceOne-AI/S1-DeepResearch-15k)\n- **Size:** the closed-ended verifiable subset of ~15,000 trajectory rows (open-ended exploration rows have no gradable gold and are dropped)\n\n## Changelog\n\n- 2026-08-31: Yield task records on demand so bounded evaluations construct only the requested prefix.\n- 2026-07-31: Moved the semantic-grading prompt into a packaged, environment-owned reference judge for `verifiers>=0.2.2.dev65`; saved configs now carry the judge ID instead of a checkout-specific prompt path.\n- 2026-07-02: Initial v1 taskset.\n","encoding":"utf-8","truncated":false,"total_bytes":930},"status":null}