{"data":{"kind":"file","path":"README.md","version_id":"zzf5ji4c1kx7evblmo6y2r5e","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":12593,"modified_at":"2026-10-05T01:25:28.472000","content_hash":"6ba966d074e54062a270196eda16fd560c0f746f649d11f432122a8fbc1082e3"},"entries":[],"content":"# Curator\n\n> Look at a set of things and work out why they belong together, then find the ones that only look like they do.\n\n**Tags:** `single-turn` `reasoning` `eval` `train` · **verifiers:** v1 (`verifiers.v1`)\n\n**On the Hub:** [1999-labs/curator](https://app.primeintellect.ai/dashboard/environments/1999-labs/curator)\n\n## What this is\n\nMost reasoning environments test **answer-finding**. A question has a determinate\nanswer and the model's job is to reach it: solve the equation, find the fact,\npass the tests.\n\nCurator tests something else. Someone shows you eight to twelve things and asks\nyou to say what they have in common. Then they ask which ones don't really\nbelong. There is no stated rule, and the two questions are coupled, because the\ncollection was built to make the second question depend on getting the first one\nright.\n\nHere is a real row from the dataset, verbatim:\n\n```text\nThe collection:\n\nA. Caesar salad: Romaine with croutons, parmesan, and a coddled-egg dressing.\nB. Oysters Rockefeller: Baked oysters on the half shell under a rich green herb topping.\nC. Bananas Foster: Bananas flambeed in butter, brown sugar, and rum, served over ice cream.\nD. Chicken Marengo: Chicken braised with tomatoes, garlic, and wine, garnished with crayfish.\nE. Pavlova: Meringue with a crisp crust and soft center, topped with fruit and cream.\nF. Chicken Tetrazzini: Baked pasta and chicken in a creamy mushroom-and-sherry sauce.\nG. Beef Stroganoff: Sauteed strips of beef in a sour-cream sauce.\nH. Peach Melba: Poached peaches with raspberry sauce over vanilla ice cream.\nI. Salisbury steak: Seasoned ground-beef patty served in brown gravy.\nJ. Carpaccio: Paper-thin slices of raw beef dressed with oil and shaved cheese.\nK. Lobster Thermidor: Lobster in a cognac-cream sauce, returned to the shell and browned.\nL. Nachos: Tortilla chips baked with melted cheese and sliced jalapenos.\n```\n\nThe answer is `D` and `K`, because the principle is *dishes named after a\nspecific real person*. Chicken Marengo is named after the Battle of Marengo,\na village in Piedmont. It is tied to Napoleon but not named for him. Lobster\nThermidor is named after the play *Thermidor*, which is named after a month in\nthe French Revolutionary calendar. Nobody is involved.\n\nBoth intruders are real dishes with proper-noun names, so the obvious answer,\n\"restaurant dishes with capital letters in the name,\" explains all twelve items\nand is wrong. The decoy is authored into every row on purpose.\n\nThe model has to do three things:\n\n1. **Generate hypotheses over a set.** Propose candidate principles that explain\n   most of the items.\n2. **Discriminate between near-equivalent hypotheses.** The decoy explains\n   *every* item, so the model has to find the tighter rule under which exactly\n   one or two items fail.\n3. **Use facts that are not in the text.** The principle usually turns on\n   something about each item: where it came from, what it is made of, how it\n   got its name. None of that appears in the description.\n\n## Why I built it\n\nI keep coming back to the judgment a curator or editor makes. Someone looks at a\nshelf and sees why those things are on the same shelf. They notice the one that\nwas put there by mistake. That judgment is not retrieval and it is not\ndeduction. It is the ability to hold a lot of loose facts at once and notice\nwhich ones bind.\n\nThat ability is what I want from the models I work with, and it is close to\nabsent from the Hub. The reward functions on these environments are almost all\nbuilt around convergent correctness, which makes them excellent at measuring\nwhether a model can find a known answer and poor at measuring whether it can\nnotice what is missing from a set.\n\nThere is a practical reason to want a signal like this measured now. Software\ngeneration is getting cheap and fast. Generating a competent first draft of\nalmost anything is not going to stay scarce for long. Knowing what is worth\nkeeping, and being able to say why, is a different problem and I think it stays\nvaluable. If a training signal for that judgment is going to be built, it is\nbetter built on purpose than stumbled into.\n\nI want to be straight about what this environment is and is not. It is not a\nbenchmark of art criticism or a proxy for taste in any broad sense. It is a\nnarrow, mechanical version of one skill: over a set of short factual descriptions,\nfind the principle and the exceptions. Every row is hand-authored, every intruder\nis documented, and the hard cases are hard because the decoy is good, not\nbecause the question is vague.\n\n## Why the signal is underrepresented\n\n- **No keyword route.** Descriptions never contain the word for the property the\n  principle is about. The dataset validator rejects any theme word that appears\n  in member text but not intruder text. Shortcut solvers that pick the lexical\n  odd-one-out or the description-length outlier score at or below random chance.\n- **Intruders are adversarial by construction.** Every intruder is written to\n  satisfy the decoy, and is usually the item people misremember as belonging.\n  Ordinary odd-one-out puzzles do not do this.\n- **Set-level, not item-level, judgment.** No single item can be judged alone.\n  Whether \"Tokyo Tower\" is an intruder depends on whether the collection is about\n  lattice towers or about structures built for world's fairs.\n- **Graded, trainable reward.** Jaccard overlap on the intruder set gives dense\n  partial credit, and the difficulty tiers spread the reward so the signal does\n  not pile up at 0 or 1.\n\n## Task format\n\nThe system prompt explains the task. The user message lists the collection above.\nThe reply must be exactly:\n\n```text\n<intruders>D, K</intruders>\n<theme>Dishes named after a specific real person.</theme>\n```\n\n## Reward\n\n| Signal | Weight | Definition |\n|---|---|---|\n| `intruders` | 0.7 | Jaccard overlap between the predicted and gold intruder ID sets. An exact set scores 1.0, partial overlap earns partial credit, a missing tag scores 0. |\n| `theme` | 0.3 | An LLM judge grades the `<theme>` against the hidden principle: **A** same principle = 1.0, **B** partially right or too broad = 0.5, **C** wrong or just the decoy = 0.0. |\n\nMetrics (not rewarded): `exact_match`, `format_ok`, and `judged` (1 if the judge ran).\n\n**Graceful degradation.** The judge's key comes from `JudgeConfig.api_key_var`\n(default `PRIME_API_KEY`, or the Prime CLI login). If no key resolves, the judge\nis skipped and the `theme` slot mirrors the intruder score, so the total reward\nequals the intruder Jaccard and still spans 0 to 1. The `judged` metric records\nwhich mode ran. v1 has no `vf.ensure_keys`, so this check lives inside the reward\nusing the same key resolution the judge client uses.\n\n## Dataset\n\nThe dataset is bundled at `curator/data/curator.jsonl` and uses Hugging\nFace-compatible columns:\n\n| Column | Content |\n|---|---|\n| `prompt` | Chat messages: `[system, user]` |\n| `answer` | Sorted, comma-joined intruder IDs, e.g. `\"D,K\"` |\n| `info` | `spec_id`, `theme`, `decoy`, `difficulty`, `domain`, `split`, and `items` (each with `id`, `title`, `description`, `is_intruder`, and `why` for intruders) |\n\n**Size:** 203 rows, one per theme. 7 domains x 29 rows (design objects, music,\nfood, architecture, internet culture, tools, games).\n\n| Tier | Rows | train / eval | 1 intruder / 2 intruders |\n|---|---|---|---|\n| obvious | 67 | 53 / 14 | 34 / 33 |\n| moderate | 72 | 57 / 15 | 33 / 39 |\n| subtle | 64 | 51 / 13 | 27 / 37 |\n\nCollections hold 9 to 12 items. No item appears in more than one row. Validator\nresults on the full set:\n\n| Check | Result |\n|---|---|\n| Theme-word leaks | 0 |\n| Intruder position (first / middle / last third) | 0.31 / 0.38 / 0.32 |\n| Random guess, told the true intruder count | 0.118 mean Jaccard |\n| TF-IDF odd-one-out solver | 0.071 |\n| Description-length outlier solver | 0.056 |\n\n**Tiers**\n\n- **obvious:** the principle is a familiar category, such as woodwinds or games\n  with no element of chance. The intruder violates it in a way most people would\n  see on reflection.\n- **moderate:** the principle needs a specific fact about each item, such as\n  whether it was built for a world's fair or began as a mod.\n- **subtle:** the principle is second-order, such as dishes named after the\n  person who actually invented them. The decoy explains every item, the intruders\n  included.\n\n**Splits.** `train` and `eval` are theme-disjoint: 20% of each tier is held out by\nhashing `spec_id`. `split = \"all\"` (the default) uses everything.\n\n### How it was generated\n\nEach row comes from a hand-authored theme spec (`data_gen/specs/*.json`, rules in\n`data_gen/AUTHORING.md`). A spec holds the principle, the decoy, 8 to 10 members,\nand 1 to 3 near-miss intruders, each with a written `why`.\n\n1. `data_gen/build_dataset.py` samples 8 to 12 items and one or two intruders per\n   spec with a seeded RNG. It shuffles them, assigns letter IDs, and assigns splits.\n2. `data_gen/validate_dataset.py` gates the result:\n   - structural checks and near-duplicate theme detection\n   - discriminative theme-word leaks\n   - intruder-position balance\n   - three shortcut baselines (random, TF-IDF odd-one-out, description-length\n     outlier), which must not beat random\n3. Every spec went through an adversarial fact-check pass. Members, intruders, and\n   `why` claims had to be well documented and not time-sensitive. Anything\n   uncertain was cut.\n\nRebuild after editing specs:\n\n```bash\npython data_gen/build_dataset.py data_gen/specs/*.json -o curator/data/curator.jsonl\npython data_gen/validate_dataset.py curator/data/curator.jsonl\n```\n\n## Install\n\nFrom the Hub:\n\n```bash\nprime env install 1999-labs/curator@latest\n```\n\nOr with uv, straight from the index:\n\n```bash\nuv pip install curator --extra-index-url https://hub.primeintellect.ai/1999-labs/curator/install/simple/\n```\n\nOr from source, in a checkout of this repository:\n\n```bash\nuv venv && uv pip install --prerelease=allow -e .\n```\n\n## Evaluate\n\nUse the tool-less `null` harness, since the task is single-turn chat:\n\n```bash\n# quick smoke run\nuv run vf-eval curator -m openai/gpt-4.1-mini -n 20 -r 3 --env.agent.harness.id null\n\n# the bundled config: 30 held-out tasks x 3 rollouts\nuv run vf-eval @ configs/eval.toml\n\n# useful overrides\n--env.taskset.split eval                  # train | eval | all\n--env.taskset.difficulty '[\"subtle\"]'     # filter tiers\n--env.taskset.task.judge.model openai/gpt-4.1-mini\n--env.taskset.task.judge.api-key-var OPENAI_API_KEY\n--env.taskset.task.judge.base-url https://api.openai.com/v1\n--env.agent.runtime.type subprocess       # run locally without a Prime tunnel\n```\n\n`pyproject.toml` also carries `[tool.verifiers.eval]` defaults (30 examples x 3\nrollouts) for tooling that reads them. v1 `vf-eval` reads the TOML config above.\n\nTo use the taskset directly in v1, construct it with its own config:\n\n```python\nfrom curator import CuratorTaskset\nfrom curator.taskset import CuratorConfig\n\ntasks = CuratorTaskset(CuratorConfig(split=\"eval\", difficulty=[\"subtle\"])).load()\n```\n\n## Tests\n\n```bash\nuv pip install pytest pytest-asyncio\nuv run pytest\n```\n\nThe tests cover dataset integrity, splits and filters, the validator, answer\nparsing, Jaccard scoring (exact, partial, over-flagging, unformatted), the\nno-key fallback, and the judge path with a mocked judge.\n\n## The reward is non-degenerate\n\nClaude Haiku 4.5 answered 90 puzzles blind (30 per tier, stratified sample,\nseed 7). They were posed as the exact system and user prompts above, and the\nanswers were scored with this package's own `parse_intruders` and `jaccard`.\nThis is intruder reward only, which is the no-judge mode:\n\n| Tier | Mean reward | Exact set | Scored 0 / partial / 1 |\n|---|---|---|---|\n| obvious | 0.572 | 40% | 7 / 11 / 12 |\n| moderate | 0.328 | 27% | 18 / 4 / 8 |\n| subtle | 0.150 | 7% | 23 / 5 / 2 |\n| **overall** | **0.350** (sd 0.416) | 24% | 48 / 20 / 22 |\n\nThe reward falls monotonically with tier. It does not saturate at 0 or 1, and\nwithin-tier variance is high, which is what GRPO-style training needs. Every reply\nwas well formatted. Subtle rows sit close to the 0.118 random floor for a small\nmodel, which leaves headroom for stronger models and for training.\n\nThe container used to build this environment had no inference API access, so this\nrun went through a sandboxed subagent rather than `vf-eval`. `vf-eval` was\nseparately checked end to end, with the `null` harness and subprocess runtime,\nagainst a local OpenAI-compatible stub. To reproduce with a real endpoint:\n\n```bash\nuv run vf-eval curator -m <model> -n 90 -r 1 --env.agent.harness.id null\n```\n","encoding":"utf-8","truncated":false,"total_bytes":12593},"status":null}