{"data":{"kind":"file","path":"README.md","version_id":"bl4r451yvwzc6uf3x2znryxx","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":1791,"modified_at":"2026-10-09T07:48:50.340000","content_hash":"2435bc6b858f0a50e35f5ff91e92b9a08862c0a2bb9ecc62a85e23694ca02d78"},"entries":[],"content":"# DNA & Genetics Reasoning Environment\n\nA verifiers environment for training and evaluating LLMs on core\nmolecular-genetics tasks with deterministic, rule-based verification.\nVersion 0.0.1 — PUBLIC release.\n\n## Task types\n\n| Task | Question | Answer |\n|------|----------|--------|\n| `revcomp` | Reverse complement of a DNA strand | DNA sequence, 5'->3' |\n| `transcribe` | mRNA from a 3'->5' template strand | RNA sequence |\n| `reverse_transcribe` | cDNA coding strand from an mRNA genome | DNA sequence |\n| `translate` | Protein from a coding strand | dash-separated 3-letter amino-acid codes |\n| `mutation` | Classify a single-base substitution | synonymous / missense / nonsense / frameshift |\n\nSequences are generated procedurally with a seeded RNG, so the dataset is\nreproducible and unlimited in size. Verification is exact-match on\nnormalized strings for sequence tasks, partial-credit overlap for\ntranslation, and keyword match for mutation classification.\n\n## Reasoning quality rubric\n\nBeyond answer correctness, the environment scores four multi-signal\nreasoning metrics that resist keyword-stuffing:\n\n* `metric_reasoning_structure` — explicit step markers, ordering language,\n  sufficiently long non-repetitive sentences.\n* `metric_domain_reasoning` — multiple distinct domain reasoning verbs\n  used together with the relevant biological rules.\n* `metric_explanation_specificity` — concrete anchors: sequences,\n  positions, amino-acid codes, numbers.\n* `metric_answer_alignment` — the final answer also appears inside the\n  reasoning trace.\n\n## Usage\n\n```python\nfrom dna_genetics import load_environment\n\nenv = load_environment(n_per_task=20, seed=42)\n```\n\n```bash\nprime env push -p .\n```\n\n## Evaluation\n\n```bash\nprime eval run javier/dna-genetics -m Qwen/Qwen3-0.6B\n```\n","encoding":"utf-8","truncated":false,"total_bytes":1791},"status":null}