{"data":{"kind":"file","path":"README.md","version_id":"h8qbe11pjebqzrkl8sjvgt5z","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":983,"modified_at":"2026-09-13T04:06:06.931000","content_hash":"2daa22e062eb57371707201ee1a2c43fcb8a386b2b87a7406f1f7bdf9e59a1a1"},"entries":[],"content":"# oolong-synth\n\nOolong synthetic long-context tasks solved by an agent in a sandbox: each task's long context window is uploaded to a file so the agent can scan it from a REPL and write a single-token answer to `/workspace/answer.txt`. Tasks are scored with the deterministic official Oolong synth rules (exact match, with partial credit for numeric and date answers), or by an optional host-side binary LLM judge when configured.\n\n## Taskset\n\n- **Source:** [oolongbench/oolong-synth](https://huggingface.co/datasets/oolongbench/oolong-synth)\n- **Size:** 1300 tasks (`validation` split), filtered at load time to a single `context_len` token bucket (default 262144)\n\n## Changelog\n\n- 2026-08-31: Yield task records on demand so bounded evaluations construct only the requested prefix.\n- 2026-08-30: Standardized binary verdict parsing and failure handling on `ReferenceJudge` while retaining deterministic scoring and answer-file response selection.\n- 2026-06-24: Initial v1 taskset.\n","encoding":"utf-8","truncated":false,"total_bytes":983},"status":null}