{"data":{"kind":"file","path":"README.md","version_id":"q5jvfx3tbeytdw976fa43xpt","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":993,"modified_at":"2026-09-13T04:06:06.930000","content_hash":"554206ab9fbab235106944a94e351b7ea5485b655ad75ecf856af7e1dfde4443"},"entries":[],"content":"# oolong-pairs\n\nOolong-Pairs long-context pairwise-aggregation tasks solved by an agent in a sandbox. Each task presents a long context of thousands of general-knowledge questions (each tied to a non-unique User ID whose implicit TREC coarse category must be inferred), and the agent must compute exact aggregate label statistics over pairs of users and write the matching `(id1, id2)` pairs to an answer file. Scored deterministically by precision / recall / F1 over the gold pair set (F1 is the reward).\n\n## Taskset\n\n- **Source:** [mit-oasys/oolong-pairs](https://huggingface.co/datasets/mit-oasys/oolong-pairs) (questions + gold pairs) with context windows from [oolongbench/oolong-synth](https://huggingface.co/datasets/oolongbench/oolong-synth)\n- **Size:** 20 tasks (the 20 questions for the selected `context_len` bucket, default 32K)\n\n## Changelog\n\n- 2026-08-31: Yield task records on demand so bounded evaluations construct only the requested prefix.\n- 2026-06-24: Initial v1 taskset.\n","encoding":"utf-8","truncated":false,"total_bytes":993},"status":null}