{"data":{"kind":"file","path":"README.md","version_id":"ovxpkd662vrh20e8dnm1srxr","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":853,"modified_at":"2026-09-13T04:06:06.951000","content_hash":"2bbe547f25c3cc530c9771f758a46b8aac53b4f29114d71897f38a7b8941fa92"},"entries":[],"content":"# openseeker\n\nOpenSeeker web-research QA: the taskset ships only the questions and scoring (no search tool — the agent brings its own web search) and the agent answers in chat. The last reply is graded against the gold answer by a binary reference LLM judge using OpenSeeker's [CORRECT]/[INCORRECT] prompt with A/B verdict labels.\n\n## Taskset\n\n- **Source:** [PolarSeeker/OpenSeeker-v1-Data](https://huggingface.co/datasets/PolarSeeker/OpenSeeker-v1-Data) (`train` split)\n- **Size:** 11677 tasks\n\n## Changelog\n\n- 2026-08-31: Yield task records on demand so bounded evaluations construct only the requested prefix.\n- 2026-07-31: Moved the OpenSeeker prompt into a packaged, environment-owned reference judge for `verifiers>=0.2.2.dev65`; saved configs now carry the judge ID instead of a checkout-specific prompt path.\n- 2026-06-25: Initial v1 taskset.\n","encoding":"utf-8","truncated":false,"total_bytes":853},"status":null}