{"data":{"kind":"file","path":"README.md","version_id":"n0si3j6jp4448rfds5kdd1r2","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":21116,"modified_at":"2026-10-09T04:08:11.062000","content_hash":"8b492ea7e7524dbf5ebd9cd36314f219087751dfb7c0f70d6f6b4bc126267fe3"},"entries":[],"content":"# AEVION Chess Mate — verifiable reasoning tasks\n\nForced-mate chess positions with **deterministic, dependency-free verification**: every\naccepted move is precomputed by exhaustive search, so checking an answer is a set lookup,\nnot an engine call.\n\n## Why this environment is unusual\n\nMost environments verify with a judge model or a fuzzy rubric. Here the outcome is binary by\nthe nature of the subject: the move either delivers mate or it does not. There is nothing to\ntune and nothing to disagree about.\n\n**There is no chess engine and no chess library inside this package.** All legal moves that\nreach the goal were enumerated in advance on AEVION's side using `chess.js`\n(BSD-2-Clause). That keeps this package free of copyleft dependencies — relevant because\n`python-chess` is GPL-3.0+, and importing it would make any environment that uses it a\nderivative work.\n\n## What is in the bank\n\nMeasured on the shipped `puzzles.json`, 2026-10-05:\n\n| | |\n|---|---|\n| positions | **2,491** |\n| themes | mate in 1 — **1,702**, mate in 2 — **688**, mate in 3 — **101** |\n| difficulty (Glicko rating) | **400 … 2522**, median **1007** |\n| positions with a single solution | **2,450 of 2,491 (98.4 %)** |\n| positions with alternative solutions | **41** — every one of them is accepted, which a one-string comparison would reject |\n\nDifficulty is a number, not a label, so you can build a curriculum (`rating_min`,\n`rating_max`) and see exactly where a model breaks.\n\n### The traceable set — 200 positions, added 2026-10-07\n\n`load_traceable()` returns a second, smaller set that carries the two things the bank above\ncannot:\n\n| | traceable set | bank above |\n|---|---|---|\n| positions | **200** (mate-in-1: 100, mate-in-2: 100) | 1,237 |\n| difficulty | **615 … 1,982** | 400 … 2,488 |\n| exactly one accepted move | **196 of 200** | 1,213 of 1,237 |\n| **source id on every position** | **yes — 200 of 200**, `li_<PuzzleId>` | **yes — 2,491 of 2,491** |\n| **accepted-move set proven complete** | **yes, exhaustive search** | yes — all three depths, 101 of 101 mate-in-3 sets re-derived twice by independent implementations |\n\nWhy it exists. An **incomplete** accepted-move set scores a genuine alternative mate as 0 —\nthe signal punishes a correct move, invisibly. Completeness can only be *proven* where the\nsearch finishes: one ply for mate-in-1, three plies for mate-in-2. Mate-in-3 is five plies,\ntens of millions of positions per task, which we do not complete — so the traceable set\ncontains none, and proven completeness is never mixed with assumed completeness.\n\nVerification of this set (2026-10-07): reference line legal and mating with the declared\nlength — **200 of 200**; complete mate-in-1 sets re-derived by a **second, different**\npredicate (mate marked in move notation rather than play-then-test) — **100 of 100**, exact\nset equality; declared mate length confirmed by Stockfish 18 Lite at depth 14 — **200 of\n200**; four deliberate data corruptions — **4 of 4 rejected**.\n\n```python\nfrom aevion_chess_env import load_traceable\ntask = load_traceable()[0]\ntask.source_id          # \"li_0Rrzp\" → row of the open Lichess database\n```\n\n### Grading has three outcomes, not two\n\n`grade()` separates \"the model was wrong\" from \"we could not parse what the model wrote\":\n\n```python\nfrom aevion_chess_env import grade\ngrade(task.solutions[0], task.solutions)  # verdict \"correct\",      reward 1.0\ngrade(\"h1h2\",            task.solutions)  # verdict \"wrong\",        reward 0.0\ngrade(\"I think Ra8\",     task.solutions)  # verdict \"unverifiable\", reward None\n```\n\nCollapsing those two makes every formatting failure count as a reasoning failure: comparing\na model that answers `a1a8` against one that answers \"I think the answer is Ra8 because…\"\nthen measures format compliance, not chess. What to do with `unverifiable` — score zero,\ndrop, or re-ask — is your decision, and the environment will not make it silently.\n\nInput is UCI, with unambiguous decoration removed (`a1a8#`, `1. a1a8`, `12...a1a8`). SAN is\n**not** accepted: SAN→UCI without a board is guesswork. A leading move number is stripped\nonly when a separator follows, so `1a1a8` stays `unverifiable` — a false `correct` inflates\na score silently and is worse than a refusal. `reward_mate()` keeps its original binary\nbehaviour for existing callers.\n\n## Install and run\n\nFrom the Hub, which is what you most likely want:\n\n```bash\nprime env install aevion/aevion-chess-mate\n```\n\nFrom a checkout:\n\n```bash\npip install -e .\npython test_env.py                 # 46 self-checks, exit 0 on success\npython test_letter_claims.py       # every number we publish, re-measured\n```\n\n**Check us instead of trusting us.** `test_installed_package.py` walks the path you just took —\nit imports the *installed* package, counts the bank, re-derives the composition, scores a\ncorrect move and a fabricated one, and refuses to run (exit 2) if it finds itself importing\nfrom a source tree instead of site-packages:\n\n```bash\npython -m venv /tmp/buyer\n/tmp/buyer/bin/pip install aevion_chess_mate   --extra-index-url https://hub.primeintellect.ai/aevion/aevion-chess-mate/install/simple/\n/tmp/buyer/bin/python test_installed_package.py    # 13 checks, exit 0\n```\n\nThat refusal is not decoration. Every defect listed in this file as \"fixed in 0.3.3\" was\ninvisible to checks that ran against the source tree and visible within seconds to checks that\nran against the installed wheel.\n\n```python\nfrom aevion_chess_env import ChessMateEnv\n\nenv = ChessMateEnv(seed=42, theme=\"Мат в 1\", rating_max=1200)\nobs  = env.reset()        # {\"fen\": ..., \"goal\": \"mate1\", \"rating\": 491, \"side\": \"black\", ...}\n_, reward, done, info = env.step(\"g6g3\")\n# reward 1.0, info[\"accepted\"] lists every move that would have been correct\n```\n\nEpisodes are reproducible: the same `seed` yields the same order of positions, and sampling\nis **without replacement** within an epoch — otherwise an agent sees far fewer distinct\npositions than episodes, and part of the bank never participates in training.\n\n## Using it on the Environments Hub\n\nThe core package has no dependencies. The Hub wrapper lives in `aevion_chess_env/taskset.py`\nand is installed separately, so the core stays dependency-free:\n\n```bash\npip install -e \".[hub]\"\n```\n\nIt exposes `ChessMateTaskset` (verifiers v1 API: `vf.Taskset` / `vf.Task` / `@vf.reward`).\nConfig options: `theme`, `rating_min`, `rating_max`, `limit`.\n\n**Ignore the usage snippet the CLI prints after installing.** `prime env install` ends with\n\n```python\nfrom verifiers import load_environment      # v0 API — not how this package loads\nenv = load_environment('aevion-chess-mate')\n```\n\nwhich the CLI prints unconditionally for every environment regardless of the runtime it\ndeclares (checked in the CLI source: the two lines have no branch on runtime). This package\ntargets the v1 API, so load the taskset class directly:\n\n```python\nfrom aevion_chess_env.taskset import ChessMateTaskset, ChessMateConfig\ntaskset = ChessMateTaskset(ChessMateConfig(theme=\"Мат в 2\", limit=256))\n```\n\nTwo measured notes on the extra, both of which were our own defects before 0.3.3:\n\n* the extra pinned `verifiers>=1.0`, which **cannot be satisfied** — that package has never\n  reached 1.0 (latest on PyPI is 0.3.1). `pip install \"aevion-chess-mate[hub]\"` failed with\n  \"No matching distribution found\". The v1 in `verifiers.v1` is a submodule name, not a\n  version number. Now `verifiers>=0.3`, tested against 0.3.1.\n* up to and including 0.3.2 the Hub listed this package as **Unclassified**, because it infers\n  the runtime from a *required* dependency and ours is deliberately optional. Declared\n  explicitly from 0.3.3 on; the dependency is still optional, so the core stays at zero.\n\nUnlike most environments, verification needs no isolated runtime: the reward is a pure\nfunction of the trace, because every accepted move is already precomputed. `validate()`\nchecks each shipped position against its own solution set, with no model involved.\n\n## Answer format\n\nMoves are expected in **UCI** (`g6g3`, `e7e8q`). Decorations are stripped (`G6G3#` works).\nSAN (`Qg3#`) is **not** accepted: converting SAN to UCI requires a board, and this package\ndeliberately has no board. Say so in your system prompt.\n\n## How this differs from the other chess environments on the Hub\n\nThere are already several, and one is close to this one — so here is the honest comparison\nrather than a claim of novelty.\n\n`albertklorer/chess-rlvr` is single-turn, takes a FEN **plus the list of legal SAN moves**, and\nrewards the chosen move with a **precomputed Stockfish regret** (0.0 for the best move,\nnegative for worse ones). It covers arbitrary positions, not just mates.\n\nThis environment is narrower on purpose, and differs in three ways that matter for training\nreasoning rather than move selection:\n\n1. **The reward is a proven fact, not an evaluation.** A forced mate was established by\n   exhaustive search; it does not depend on an engine or its depth. A regret score does —\n   change the depth and the numbers change, and so does what the model learns.\n2. **No list of legal moves is given.** The model has to find the move itself, which is a\n   harder task than picking from a supplied list.\n3. **All correct moves are accepted.** 41 of the 2,491 positions have more than one solution,\n   and every one of them scores 1.0; a single-answer comparison would punish a correct\n   alternative.\n\nWhat the other environment does better, plainly: it covers the whole board of positions rather\nthan mates only, and it accepts SAN, which is closer to how a model naturally writes moves.\n\n### On \"proven complete\" for mate-in-3\n\nTwo different claims were made about this inside the team, and both were wrong in opposite\ndirections, so here is the measured answer.\n\nThe claim \"completeness cannot be proven for mate-in-3 — five plies, tens of millions of\npositions\" was an **upper bound, not a measurement**: in mate positions the tree collapses as\nsoon as the defender has a reply with no mate, so the worst case never materialises. The\nopposite claim, \"proven by exhaustive search\", was equally unsafe as a blanket statement,\nbecause the search was not run on every position.\n\nMeasured on real positions from this bank, recomputing each accepted-move set from scratch:\n\n| | |\n|---|---|\n| mate-in-3 positions re-verified from scratch | **101 of 101 — all of them** |\n| differences found | **0** |\n| cost | 45 min total (~27 s per position) |\n| a deliberately easy control (starting position, five plies) | **1.2 s** — which is why an easy control misleads here |\n\nEvery mate-in-3 accepted-move set has now been recomputed from scratch, and every one matched\nthe shipped set. The check is two-sided, which is the part that matters: not only does the\nlisted move force mate, but **no other legal move does** — that second half is what\ncompleteness means, and it is the half that protects a model from being scored 0 for a correct\nalternative.\n\nA second, **independent implementation** — different code, same data, written by another team\nmember — has now finished on every position:\n\n| | |\n|---|---|\n| verified by **two independent implementations** | **101 of 101**, zero differences |\n| accepted-move sets of size 1 | **101 of 101** — each position has exactly one forcing move, and that is a search result, not an assumption |\n| cost of the second pass | median **14.9 s** per position, max **157 s**, **47.7 min** for 99 positions |\n\nThe two implementations were kept apart on purpose. \"My search agrees with itself\" and \"two\nseparately written searches agree\" are different claims, and only the second rules out a shared\nbug in a single implementation. Both halves of the two-sided check were repeated by the second\npass, and its controls were run before the work: a mate-in-1 is found by a depth-3 search, a\nposition with no mate yields an empty set, and the starting position has no mate in 3.\n\nOne measurement detail, because it would otherwise travel as a false number: two of the 101\ntimings (54 min and 781 min) were taken by wall clock across an interrupted run and a machine\nsleep, and one of them accounted for 95% of the naive total. They are artefacts of measurement,\nnot hard positions, so the cost above is stated over the 99 clean timings. Both implementations\nland in the same range — the first pass took 45 min — and an earlier claim that one was ~6×\nslower was itself an artefact of measuring under a concurrent heavy run.\n\n\nWhy this matters beyond pedantry: an **incomplete** accepted-move set scores a genuine\nalternative mate as 0. The signal then punishes a correct move, and it does so invisibly.\n\n## Provenance — every position traceable to its source\n\n⚠️ **Before 0.3.3 this was true of the data and false of the API.** The two data files carry\nthe identifier under different keys — `source_id` in the main bank, `id` in the traceable set —\nand the loader read only `id`, so `load_puzzles()[0].source_id` came back **empty for all\n2,491 positions** while the file held an identifier for every one of them. The one test that\nchecked `source_id` ran against the traceable set, where the key happened to match, so it\nstayed green. It was caught by installing the published wheel into a clean environment and\nwalking through what a reader of this file would actually type. Fixed in 0.3.3, and the bank\nnow has four checks of its own that print the denominator, plus a control that goes red on an\nempty identifier.\n\n\n**All 2,491 positions carry a `source_id`** of the form `li_<PuzzleId>`, matched against the\nsource bank by **full FEN** — halfmove counters included, so each match is an exact record,\nnot merely \"the same position\". Matching ran against the complete mate slice of the bank:\n158,045 records.\n\n**The match is reproducible, and that was verified rather than assumed.** Two independent\npasses over the bank, run back to back, produced **identical mappings for all 2,491 positions**:\nsame positions, same ids, zero differences. The comparison itself was control-tested — a\ndeliberately altered id is detected, so \"identical\" does not mean \"unable to tell apart\".\n\nThis section previously said the opposite, and the history is worth stating because it explains\nwhat to trust. The first pass reported 97 positions with no origin. Those 97 were an artefact:\nthe paginated API ordered rows by a **non-unique** field, so pages silently skipped records,\nand a skipped record looked like \"not in the bank\". After the ordering was fixed at the source,\ntwo consecutive passes agreed exactly and every position matched. Nothing about the positions\nchanged — only our ability to count them did.\n\n## Multi-step episodes — the agent plays the mate out\n\nA one-move answer tests one decision. Training reasoning needs a horizon: does the model hold\na plan when the opponent replies in a way it did not expect? So the forcing tree is\nprecomputed, and the agent plays the line to the end — up to three of its own moves, with\nthe defender answering in between.\n\n```python\nfrom aevion_chess_env.lines import ChessMateLineEnv\n\nenv = ChessMateLineEnv(seed=7)\nobs = env.reset()                     # {\"fen\": ..., \"moves_left\": 2, \"plies_left\": 3, ...}\nobs, reward, done, info = env.step(\"g4h3\")   # reward 1.0, done False\n                                             # info[\"defender_reply\"] = \"h2g1\"\nobs, reward, done, info = env.step(\"h3g2\")   # reward 1.0, done True, info[\"mate\"] = True\n```\n\n**Every node is re-checked by code that does not share the search which built it**, and that\ntook two predicates rather than one, because a single one would have left a gap we found only\nby counting nodes:\n\n* at a **leaf** the mate is one move away, so \"which moves mate\" is answered by applying each\n  legal move and asking `isCheckmate` — no search at all;\n* at an **intermediate** node (mate in two from there) the answer is composed of those same\n  `isCheckmate` calls — a move qualifies when, after it, every defender reply leaves a mate in\n  one — with no shared recursion;\n* a **root** set is compared against the single-step data shipped in this package, which was\n  itself re-derived by two independently written implementations.\n\nThe gap is worth naming because it was ours. After the first pass only the leaves had been\nchecked independently, and that sounded like the whole tree. Counting nodes showed the\nmate-in-3 trees hold 481 nodes — 101 roots, **218 intermediate** and 162 leaves — so 218 of\nthem still rested on the very search that produced them, which is not evidence. They are now\nchecked separately: 218 of 218 matched, with zero misses in either direction.\n\nBoth directions are counted apart, because they fail differently: a move that forces mate but\nis not accepted punishes a correct answer, and an accepted move that does not force mate\ncredits a wrong one.\n\n**Verifiability is unchanged, and that is the whole point.** Every node of the tree carries the\n*complete* set of moves that force mate in the remaining number of moves, enumerated in advance\nby exhaustive search. Grading stays a set lookup: no engine, no judge model, no dependencies.\n\nThree design decisions worth stating, because each could reasonably have gone the other way:\n\n* **The defender's reply is chosen by seed, not by strength.** After a forcing move *every*\n  reply loses, so \"best reply\" is not a meaningful notion here, while reproducibility is: the\n  same seed yields the same episode, or two runs of one model are not comparable.\n* **A wrong move ends the episode with 0.0.** A wrong move can destroy the forced mate, and\n  from there the task is no longer verifiable — continuing would mean scoring something we\n  never proved. `info[\"reason\"]` says so explicitly rather than leaving it to be inferred.\n* **Reward is 1.0 per correct move, not only at mate.** Otherwise a long line gives the same\n  signal as a short one, and the training signal cannot tell at which step the model lost the\n  plan.\n\nMeasured, on the data actually shipped:\n\n| | |\n|---|---|\n| positions with a precomputed tree | **789** — every mate-in-2 (688) and mate-in-3 (101) position |\n| nodes in those trees | **1,990**: 789 roots, 218 intermediate, 983 leaves |\n| root accepted-move sets matching the single-step data | **789 of 789**, zero differences |\n| nodes with a missing or empty accepted set | **0** across the whole tree |\n| leaf nodes re-checked by a search-free predicate | **983 of 983** matched exactly |\n| intermediate nodes re-checked by an independent depth-2 predicate | **218 of 218** matched exactly |\n| moves that force mate but were not accepted (would punish a correct answer) | **0** |\n| moves accepted that do not force mate (would credit a wrong answer) | **0** |\n| checks | `python test_lines.py` — 16, including a denominator per depth and a control that the tree walker detects a broken node |\n\n## Honest limits\n\n* **Only forced-mate tasks.** Positions whose goal is \"find the best move\" need an engine\n  evaluation, which would reintroduce a GPL dependency; they are excluded.\n* **Chains are short.** The underlying bank holds ~500k positions but only about 0.4 % have\n  a 3+ move forced line. This package ships a sample, not the whole bank.\n* **Multi-step episodes cover mate-in-2 and mate-in-3** (789 positions). Mate-in-1 is\n  single-step by nature, so 1,702 of the 2,491 positions stay one move long.\n* **No multi-turn Hub wrapper.** The multi-step environment is a plain Python API, verified by\n  its own 13 checks. A `verifiers` rollout wrapper for it is *not* shipped, because\n  `verifiers.v1` needs `fcntl` and cannot run on the machine this was built on — writing a\n  wrapper we cannot execute would mean shipping untested code and calling it a feature.\n* **No containerisation.** Plain Python package.\n\n## Data provenance and licence\n\nPositions come from the open Lichess puzzle database, released under **CC0**:\n\n> \"Database exports are released under the Creative Commons CC0 license. Use them for\n> research, commercial purpose, publication, anything you like.\"\n> — https://database.lichess.org/ (checked 2026-10-05)\n\nPrecomputed solution sets are derived from those positions by exhaustive search.\n\nPackage code: **Apache-2.0** (see `LICENSE`).\n\n## Verification of the bank itself\n\nEvery shipped position was replayed against its own solution set: **2,491 of 2,491** reference\nsolutions are accepted, and all 2,491 pass the environment's own `validate()`. For the 41\npositions with alternatives, every alternative is accepted too.\n\nSelf-checks include negative controls (wrong move, empty answer, SAN input, empty pool) and\nwere validated by mutation. Mutating each branch separately — not just one — is what found the\ngaps: an earlier `validate()` was a tautology (it asked a set whether its own members belonged\nto it) and survived until the mutation `return True` was tried; a missing length check survived\nbecause both malformed-move samples were caught by the square-range check instead. Both are\nfixed and the mutations now fail the suite.\n","encoding":"utf-8","truncated":false,"total_bytes":21116},"status":null}