{"data":{"kind":"file","path":"README.md","version_id":"yo7zvdft3q1wsq1qkdfm2q9j","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":3521,"modified_at":"2026-10-04T08:14:51.274000","content_hash":"e08dc018ef422e15a8c8730b7910ca436d3d85a29d971a8e5e9fc028df7ab8ed"},"entries":[],"content":"# Exact Math Curriculum — V1.0.1\n\nA single-turn mathematical reasoning environment for reinforcement learning\nand evaluation. Package release: `1.0.1`. Hub: `lema/exact-math-curriculum`.\n\n## Tasks\n\nTen balanced families: fraction arithmetic, linear equations, two-variable\nlinear systems, quadratic root sets, combinations, sampling probabilities,\nmodular exponentiation, polynomial definite integration, squared coordinate\ndistance, and arithmetic progression sums. Each has three difficulty levels.\nAll answers are computed analytically using Python integers and `Fraction`;\nno external datasets, network calls, model judge, or floating-point tolerance.\n\nDefault generation: 1,000 training and 200 evaluation questions at level 2.\nQuestion SHA-256 modulo 5 assigns the evaluation partition (20% of the task\nspace). The requested sizes are sampled from those partitions independently.\nThus train/eval do not share identical question text, even across seeds.\nThis is instance-level separation, not held-out task-template generalization.\nGeneration is seeded, unique within each split, and balanced across topics\n(with count differences at most one). A bounded rejection loop raises an error\nif a small task space is exhausted rather than hanging indefinitely.\n\n## Answer contract and reward\n\nExplain if useful, then emit exactly one `<answer>...</answer>` block.\nScalar examples: `<answer>-3/7</answer>` or `<answer>0.5</answer>`.\nSystems: `<answer>[1/2, -3]</answer>` in [x, y] order.\nQuadratics: `<answer>[-2, 3]</answer>`; root order is ignored, duplicates fail.\n\nReward is 1.0 only for an exactly equivalent numeric answer and 0.0 otherwise.\nEquivalent fractions and finite decimals are accepted; approximations, units,\nexpressions, NaN/Infinity, code, duplicated answer blocks, and answer spraying\nare rejected. No format-only or verbosity bonus can earn reward without a\ncorrect solution. Output length, numeric length, exponent and list size are\nbounded. Parsing never executes submitted text.\n\nThis tests final-answer correctness, not the validity of a written proof.\nProcedural generators and ground truth are public; this is not a secret judge.\nReference parameters are evaluator metadata and are not inserted into prompts.\n\n## Usage\n\n```python\nfrom exact_math_curriculum import load_environment\n\nenv = load_environment(num_train=1000, num_eval=200, seed=42,\n                       difficulty=2, topics=None)\n```\n\nArguments:\n- `num_train`, `num_eval`: integer 1..20000, default 1000 and 200.\n- `seed`: integer 0..2**32-1, default 42.\n- `difficulty`: 1, 2, or 3, default 2.\n- `topics`: nonempty unique list of names from the ten topics above; default all.\n  Machine names: fractions, linear, systems, quadratics, combinatorics,\n  probability, modular, calculus, geometry, sequences.\n- Unknown arguments are rejected.\n\n```sh\nprime env install lema/exact-math-curriculum@1.0.1\nprime eval run lema/exact-math-curriculum -m YOUR_MODEL --env-args '{\"difficulty\":2}'\n```\n\n## Reproduce validation\n\n```sh\nuv venv --python 3.12\nuv pip install --python .venv/bin/python -e '.[test]'\n.venv/bin/python -m pytest tests -q\nuv build --wheel\n```\n\nTests include independent mathematical oracles for all families at every\nlevel, split integrity, configuration errors, parser abuse, message types,\nand real verifiers rollout/rubric integration with a deterministic test client.\nThe deterministic client is a plumbing test, not a model benchmark or evidence\nof model accuracy. No paid training or inference is started by loading.\n","encoding":"utf-8","truncated":false,"total_bytes":3521},"status":null}