{"data":{"kind":"file","path":"README.md","version_id":"s5ik53cwo9y2jzjiprk23ec1","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":942,"modified_at":"2026-09-13T04:06:06.952000","content_hash":"9a017761bc469dc4d4682532c263cbc86c4adf1c29877aa2a3377c68fef15799"},"entries":[],"content":"# papersearchqa\n\nPaperSearchQA biomedical literature-search QA: the taskset ships only the questions and scoring (no search tool — the agent brings its own web search) and the agent researches over the biomedical literature, then answers in chat wrapping its final answer in `\\boxed{...}`. The last reply is graded by a strict reference LLM judge against the task's acceptable answers (primary answer plus golden variations; any match counts).\n\n## Taskset\n\n- **Source:** [jmhb/PaperSearchQA](https://huggingface.co/datasets/jmhb/PaperSearchQA) (`test` split)\n- **Size:** 5000 tasks\n\n## Changelog\n\n- 2026-08-31: Yield task records on demand so bounded evaluations construct only the requested prefix.\n- 2026-07-31: Moved the custom judge prompt into a packaged, environment-owned reference judge for `verifiers>=0.2.2.dev65`; saved configs now carry the judge ID instead of a checkout-specific prompt path.\n- 2026-06-24: Initial v1 taskset.\n","encoding":"utf-8","truncated":false,"total_bytes":942},"status":null}