{"data":{"kind":"file","path":"README.md","version_id":"ybiarejl5d7a2c69duvyjj0e","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":5161,"modified_at":"2026-09-28T19:17:16.964000","content_hash":"c45372db985f31ed4970c18fd3bb61574a6a01a8a2f820c3f576cbe778883f96"},"entries":[],"content":"# neuro-reasoning\n\nA procedurally generated RL environment of EEG BCI reasoning tasks: sampling and filtering, power\nanalysis, train/test split hygiene, epoch rejection, and preprocessing pipeline order. No LLM\njudge. Every answer is computed by the generator and checked by exact match.\n\n- **Environment ID:** `neuro-reasoning`\n- **verifiers API:** v1 (`verifiers>=0.3.1`)\n- **Task type:** single-turn, one final answer per item\n- **Tags:** neuroscience, eeg, bci, science, reasoning, single-turn, eval, train\n\n## Task families\n\n| Family | Templates | Answer type | What it tests |\n|---|---|---|---|\n| `filter_design` | 7 | integer, number, yes/no | Nyquist limits with transition bands, decimation, aliasing, FFT resolution, FIR length and group delay |\n| `power` | 7 | integer, number | sample size for two-group, paired, and Bonferroni designs; exclusion and rejection inflation; trial counts; session length; achieved power |\n| `split_hygiene` | 7 | letter (a to d) | whether an evaluation design leaks: (a) temporal leakage (b) stimulus-level overlap (c) no problem (d) subject identity leakage |\n| `epoch_rejection` | 7 | integer | how many epochs survive a stated rule applied to a table of amplitudes |\n| `pipeline_order` | 7 | letter sequence | the one valid order of 4 to 7 preprocessing steps given stated failure reasons |\n\nEvery template chains at least two constraints, so that recalling a single fact is not enough.\nExamples include the Nyquist limit applied to the filter's attenuation edge rather than its\npassband edge, and the one bad channel that must be excluded before counting.\n\n## Guarantees\n\n- **Deterministic.** The same `seed` and `split` always give the same items. Each item records the\n  sub-seed that rebuilds it.\n- **Exact verification.** Numbers must match exactly, or within a stated tolerance where the\n  question asks for rounding. Choices, yes/no, and orderings must match exactly. Replies are\n  parsed only from a final `ANSWER: <value>` line.\n- **Unambiguous.** Any sampled value near a rounding boundary is resampled. Every\n  `pipeline_order` item is checked by brute force over all permutations and has exactly one valid\n  order.\n- **Balanced.** Items round-robin over families and templates, so any prefix is balanced. Items\n  are unique on (template, parameters).\n\n## Splits\n\n| Split | Purpose |\n|---|---|\n| `train` | default parameter ranges |\n| `heldout` | on one axis per family, values disjoint from `train`: high-gamma frequencies 55 to 120 Hz, small effect sizes 0.10 to 0.25, cohorts of 50 to 200 participants, strict thresholds 30 to 60 uV, and 6 to 7 pipeline steps |\n| `anchor` | exact values from THINGS-EEG2 (Gifford et al., 2022) and textbook defaults, used for the contamination spot-check; a template may yield only a few distinct items |\n\nThe ranges are defined and locked in `neuro_reasoning/splits.py`.\n\n## Quickstart\n\n```bash\nprime env install <owner>/neuro-reasoning\nuv run eval neuro-reasoning -m <model> -n 175 \\\n  --env.agent.harness.id null --env.agent.max-turns 1 --env.agent.runtime.type subprocess\nuv run eval neuro-reasoning -m <model> -n 175 --env.taskset.split heldout \\\n  --env.agent.harness.id null --env.agent.max-turns 1 --env.agent.runtime.type subprocess\n```\n\nUse the `null` harness, a plain chat loop. The default `bash` harness gives the model a shell,\nwhich this environment does not need. `-n 175` gives five items per template.\n\n## Taskset config\n\n| Field | Type | Default | Description |\n|---|---|---|---|\n| `num_tasks` | int | `350` | items to generate |\n| `seed` | int | `0` | generation seed |\n| `split` | `train` \\| `heldout` \\| `anchor` | `train` | parameter ranges to sample from |\n| `families` | list[str] \\| null | `null` | restrict to these families |\n\n## Metrics\n\n| Metric | Meaning |\n|---|---|\n| `correct` (reward) | 1.0 if the final answer matches exactly, else 0.0 |\n| `format_ok` (metric) | 1.0 if the reply contained an `ANSWER:` line |\n\nEach trace also records the generator's parameters, template, split, and derivation in its task\ndata, so accuracy can be broken down by family and template.\n\n## Baselines\n\nAccuracy (%) on 175 train items (35 per family), temperature 0, 4-bit local weights.\n\n| Model | filter | power | split | epoch | pipeline | All |\n|---|---|---|---|---|---|---|\n| constant-answer floor | 26 | 3 | 46 | 26 | 0 | 20 |\n| Qwen2.5-0.5B-Instruct | 6 | 6 | 0 | 0 | 3 | 3 |\n| Qwen2.5-1.5B-Instruct | 6 | 11 | 0 | 11 | 0 | 6 |\n| Llama-3.2-3B-Instruct | 20 | 23 | 9 | 26 | 14 | 18 |\n\nThe \"constant-answer floor\" row always gives each template's most likely answer. Models up to 3B\nscore at or below it. Held-out accuracy for the 3B model is 10%. Full results, held-out tables, and\nthe contamination check are in `docs/RESULTS.md` in the source repository.\n\n## Scope and limits\n\nQuestions state their formulas and conventions, such as the normal-approximation sample size,\nthe z-values, and the FIR design rules. They test whether a model can apply and chain constraints,\nnot whether it can recall a textbook table. The anchor values come from one dataset paper and\nshould be checked against the paper before the contamination result is reported.\n","encoding":"utf-8","truncated":false,"total_bytes":5161},"status":null}