{"data":{"kind":"file","path":"README.md","version_id":"z0exizr3sks2m7dxasnn3cuf","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":915,"modified_at":"2026-09-13T04:06:06.952000","content_hash":"5b8c4f81f6768b7398dd34d0994f650103ce3ad3f7de8f55a6c4858838ea5239"},"entries":[],"content":"# redsearcher\n\nREDSearcher long-horizon web-research QA: the taskset ships only the questions and scoring (no search tool — the agent brings its own web search) and the agent breaks the question into search subgoals, cross-checks sources, then answers in chat. The last reply is graded against the gold answer by a reference LLM judge using the BROWSECOMP [CORRECT]/[INCORRECT] prompt with A/B verdict labels.\n\n## Taskset\n\n- **Source:** [Zchu/REDSearcher_RL_1K](https://huggingface.co/datasets/Zchu/REDSearcher_RL_1K) (`train` split)\n- **Size:** 1000 tasks\n\n## Changelog\n\n- 2026-08-31: Yield task records on demand so bounded evaluations construct only the requested prefix.\n- 2026-07-31: Moved the BROWSECOMP prompt into a packaged, environment-owned reference judge for `verifiers>=0.2.2.dev65`; saved configs now carry the judge ID instead of a checkout-specific prompt path.\n- 2026-06-25: Initial v1 taskset.\n","encoding":"utf-8","truncated":false,"total_bytes":915},"status":null}