{"data":{"kind":"file","path":"README.md","version_id":"ufhyyoncg7o7688qtro266l9","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":1171,"modified_at":"2026-09-13T04:06:06.927000","content_hash":"2a3cc5fbd255487e6127a114e42962cb57b2880aaa803ab6a50e5999b3cce795"},"entries":[],"content":"# clbench\n\nTencent CL-bench long-context QA tasks: the model answers a long-context conversation and a strict, all-or-nothing LLM judge grades its final reply against the task's rubrics, returning an `Overall Score` of 0 or 1 that becomes the reward.\n\n## Taskset\n\n- **Source:** [tencent/CL-bench](https://huggingface.co/datasets/tencent/CL-bench)\n- **Size:** 1899 tasks\n\n## Notes\n\n- The judge grades all-or-nothing: every rubric must be satisfied for `Overall Score` 1; an empty completion or judge error scores 0.\n- Prompts preserve the dataset's message structure, so they require a message-prompt harness (e.g. `default`). This also lets Verifiers transfer very large single-turn contexts without exceeding process argument limits.\n\n## Changelog\n\n- 2026-08-31: Yield task records on demand so bounded evaluations construct only the requested prefix.\n- 2026-08-30: Standardized the binary all-or-nothing grader on `ReferenceJudge` while retaining the official rubric prompt and judge diagnostics.\n- 2026-08-27: Preserve single-turn prompts as messages so Verifiers can transfer long contexts without exceeding process argument limits.\n- 2026-07-10: Initial v1 taskset.\n","encoding":"utf-8","truncated":false,"total_bytes":1171},"status":null}