{"data":{"kind":"file","path":"README.md","version_id":"mkf34mv54qrhgxfccpx2o1yf","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":764,"modified_at":"2026-07-16T07:15:27.344000","content_hash":"46e4db46c4d75302ea97c9718611006b2ebe2458f8bb24e15950d096be761d55"},"entries":[],"content":"# compressed-gsm8k\n\nA v1 GSM8K taskset for training short, scratch-like reasoning with the base harness.\n\nThe model may write brief scratch work but must end with:\n\n```text\n####\nAnswer: INTEGER\n```\n\nEach rollout receives `1.0` for an exact answer. A uniquely shortest correct rollout\nin each group receives another `1.0`. A tie for shortest receives no compression\nbonus, and incorrect rollouts never receive it. The compression bonus is disabled on\nthe `test` split, so evaluation reward is exact-match accuracy. The\n`completion_tokens` metric records model output length on both splits.\n\nThe taskset defaults to GSM8K's `train` split. Use the `test` split for evaluation.\nBoth splits are packaged with the environment so workers do not download data at startup.\n","encoding":"utf-8","truncated":false,"total_bytes":764},"status":null}