{"data":{"kind":"file","path":"README.md","version_id":"hpfpg2j5voojucny8qpkpz68","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":6756,"modified_at":"2026-10-02T10:18:35.675000","content_hash":"7caf8eda6dca616955b82f0a3c06831557602e74ec1ed64db10d3296fe7856c2"},"entries":[],"content":"# Nuclear Physics Lab — V0.1.0\n\nA single-turn Verifiers environment for academic nuclear physics calculations.\nPublished package version: `0.1.0`; release label: `V0.1.0`.\nTarget namespace: `bryant/nuclear-physics-lab`.\n\n## What it evaluates\n\nTwelve task families with procedurally generated, fully specified problems:\n\n1. Remaining radioactive fraction.\n2. Activity after exponential decay.\n3. Half-life inferred from a measured remaining fraction.\n4. Decay constant.\n5. Total binding energy using neutral atomic masses.\n6. Binding energy per nucleon.\n7. Signed reaction Q-value using nuclear rest masses.\n8. Empirical nuclear radius.\n9. Narrow-beam photon attenuation without scattering buildup.\n10. Poisson uncertainty of a background-subtracted count rate.\n11. Daughter ingrowth in a closed two-member decay chain.\n12. Detected photon rate with branching probability and absolute efficiency.\n\nThis is textbook science, not an operational reactor simulator or weapons-design\nbenchmark. No criticality, enrichment, device construction, or weapon optimization\nis included. All necessary numerical constants are given in each question; no\nexternal data, API, sandbox, or LLM judge is required to calculate the rewards.\n\n## Dataset and reproducibility\n\nDefaults: 360 training examples and 120 evaluation examples; seed 42. A local RNG\nuses SHA-256 domain separation between training and evaluation. Both splits are\nbalanced over 12 families and 3 numerical curricula (when sample count permits).\nTrain and eval share problem templates but use different numerical instances;\nthis is an interpolation benchmark, NOT unseen-topic generalization.\n\n`difficulty` is a numerical-range curriculum, not a validated ranking of reasoning\ncomplexity. Levels 1–3 alter decay durations, initial populations, attenuation\nthickness, and ingrowth durations. Some families have identical parameter ranges\nacross levels. No benchmark result is claimed for the difficulty labels.\n\nAtomic masses and reaction masses are synthetic and explicitly labelled in the\nquestions. The environment does not pretend to supply measured isotope tables.\nThe supplied constants define the exercise. Atomic-mass binding calculations use\nhydrogen-atom and neutron masses and neglect electronic binding energy; Q-value\nquestions use nuclear masses consistently. Two-member chains avoid coincident\nhalf-lives and use `expm1` for stable daughter-population evaluation.\n\n`answer` holds a JSON reference; `info` holds derivation and parameters for audits.\nNeither is inserted into the model's prompt. Prompts contain one system message\nand one user question. Reference data is accessible to the evaluator, not secret\nfrom anyone who has the source; do not expose the dataset's answer/info to agents.\n\n## Output and reward\n\nShow reasoning, then put this exact shape on one final line without Markdown:\n\n    {\"value\": 123.456, \"unit\": \"Bq\"}\n\nThe numerical accuracy reward is binary, 0 or 1. Full credit requires compatible\nphysical dimensions and numerical agreement (relative tolerance 0.001; absolute\ntolerance 1e-10 in the reference unit). Common equivalent units are converted:\nBq/kBq/MBq, eV/keV/MeV, s/min/h, fm/m, s^-1/min^-1 and 1/%.\n\nNumbers of nuclei and photons use `count`; rates use `s^-1`. Binding energy per\nnucleon uses `MeV` with dimensionless nucleon count, explicitly described in the\nprompt. The scorer checks the final line only, rejects duplicate keys, booleans,\nstring-valued numbers, NaN/infinity, unknown units, extra keys and answer dumps.\nIt never rewards mere keywords or numbers somewhere in the reasoning.\n\n`valid_final_answer` is a weight-zero diagnostic; correct JSON alone cannot earn\ntraining reward. Reasoning quality is NOT independently graded, and a numerically\ncorrect answer can earn reward even if its accompanying explanation is flawed.\n\n## Installation and usage\n\nPython 3.11–3.13; tested with Python 3.12 and verifiers 0.1.14.\n\n    prime env install bryant/nuclear-physics-lab\n\n    import verifiers as vf\n    env = vf.load_environment(\"nuclear-physics-lab\")\n\nFor local development:\n\n    uv venv --python 3.12 .venv\n    uv pip install --python .venv/bin/python -e '.[test]'\n    .venv/bin/python -m pytest -q\n    .venv/bin/python scripts/verify_release.py\n    uv build --wheel\n\n## Arguments\n\n- `num_examples`: training size, integer 1–10000, default 360.\n- `num_eval_examples`: eval size, integer 1–10000, default 120.\n- `seed`: integer 0 through 2^64-1, default 42.\n- `difficulty`: null (balanced) or integer 1, 2, 3.\n- `topics`: null (all) or a nonempty list of unique task-family names above,\n  using the snake_case identifiers exported as `nuclear_physics_lab.tasks.TOPICS`.\n\n    env = vf.load_environment(\n        \"nuclear-physics-lab\",\n        num_examples=72,\n        num_eval_examples=36,\n        seed=17,\n        difficulty=2,\n        topics=[\"activity\", \"binding_energy\", \"daughter_ingrowth\"],\n    )\n\nUnknown arguments and invalid ranges fail fast. Generated datasets are capped to\navoid accidental excessive memory allocation. Example count controls balance:\nsmall datasets may not include all family/level combinations.\n\n## Verification and limitations\n\nTests check analytical identities independently, generation determinism, split\nseparation, configuration validation, equivalent units, malformed responses,\nanti-gaming cases and actual `Rubric.score_rollout` integration. The release audit\nchecks all default train/eval samples with correct and incorrect completions and\n7,200 distinct generated examples across ten seeds and both splits. Its output is\n`artifacts/release-verification.json`.\n\nOffline oracle checks are software verification, not evidence of model capability.\nNo real-model evaluation or hosted training result is claimed in this release.\nThe dataset is template-based, numerical, and intentionally limited: it does not\ncover nuclear structure theory, spectroscopy inference, experimental apparatus\ncontrol, or complete radiation transport. Reasoning assessment requires a separate\nhuman or validated evaluator. The explicit tolerances also mean tiny numeric\nmistakes within tolerance receive full credit.\n\n## Physics references\n\nThe mathematical conventions follow standard undergraduate nuclear physics:\nKenneth S. Krane, *Introductory Nuclear Physics*, Wiley (1987), for radioactive\ndecay, masses, binding energy and nuclear size; Glenn F. Knoll, *Radiation Detection\nand Measurement*, fourth edition, Wiley (2010), for counting statistics and photon\ndetection; Bateman's coupled exponential decay equations for daughter ingrowth.\nNo textbook problem statements were copied. Exercise-specific constants are\nexplicitly supplied rather than silently fetched from a changing data source.\n\n## License\n\nMIT. See LICENSE.\n","encoding":"utf-8","truncated":false,"total_bytes":6756},"status":null}