{"data":{"kind":"file","path":"README.md","version_id":"ww4hsduizs50kib2lf3o13cc","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":4722,"modified_at":"2026-10-06T22:01:42.897000","content_hash":"119d8829df59b0db9374e73c50c212a8953806973f6a9b4601b74a10cc179fd4"},"entries":[],"content":"# Compact Language — Prime stack\n\nOne episode runs two isolated model roles:\n\n1. Sender sees English, and emits as compact a code as it can in `<code></code>` tags.\n2. Receiver sees only that code and a fixed instruction, and reconstructs the English in `<text></text>` tags.\n3. Both roles receive the same reward: `compression` **only if** the round trip is byte-exact, else `0` — `compression = max(0, 1 - code_bytes / original_bytes)`. Verbatim copies (no compression) and near-misses both earn zero; the only paid behavior is a code that is genuinely shorter AND reconstructs the original byte-for-byte. `similarity` (`1 - edit_distance/longer_length`) is recorded as a diagnostic, not paid. An errored trace earns zero. gzip of the original is recorded as a reference metric (`gzip_ratio`, `beats_gzip`): scoreboard, not reward. This gate pairs with an SFT warm-up: supervised on byte-exact compressed demonstrations first, then RL polishes past them.\n\nThere is no judge and no budget. Replies are parsed from the role's tag, falling back to the raw reply when the model omits it; `Trace.last_reply` itself is avoided because it strips surrounding whitespace. Ground truth is added to trace metadata only after both roles finish, never to the receiver prompt or task data.\n\nThis is a native **verifiers v1** package: exported Taskset and Env classes, built-in `null` chat harness, typed tasks, and Prime trace/reward recording. All model calls go through verifiers interception. A no-tool subprocess runtime suffices; training belongs on a Prime GPU instance.\n\n## Install and evaluate\n\nUse an environment with v1 verifiers installed. The installed Prime CLI 0.7.2 on this laptop still uses v0 environment APIs; its hosted eval path has not been verified for this package.\n\n```sh\nuv pip install -e /Users/kevin/pi/compact-language/environments/compact_language\nuv run eval @ /Users/kevin/pi/compact-language/configs/prime-eval.toml --dry-run\n```\n\nBefore a live run, replace the example model and configure a verified inference endpoint through verifiers client/agent configuration. Verify the actual provider and advertised pricing for internal inference. Run the same eval command without `--dry-run`. Upload remains enabled by default.\n\nFor the local source checkout, validation was performed with:\n\n```sh\ncd /Users/kevin/pi/verifiers\nUV_CACHE_DIR=/tmp/compact-language-uv \\\nPYTHONPATH=/Users/kevin/pi/compact-language/environments/compact_language \\\nuv run --no-sync eval @ /Users/kevin/pi/compact-language/configs/prime-eval.toml --dry-run\n```\n\nThe five-task eval config uses concurrency five. For 200 tasks use `--num-tasks 200 --max-concurrent 120`. The bundled synthetic dataset has 800 train and 200 test examples with disjoint templates; it is not representative English.\n\n## Train\n\n`configs/prime-rl.toml` validates against local prime-rl config classes. There is no budget; the shortness bonus always rewards compression relative to the original size.\n\n```sh\n# On a Prime GPU instance with prime-rl and this package installed:\nuv run rl @ /path/to/compact-language/configs/prime-rl.toml\n```\n\nThe config uses role-conditioned RAE advantage baselines and **one shared trainable model** in sender and receiver roles. It does not implement two independently optimized models or Rajan's PPO-compressor / cross-entropy-decoder alternation. `train_sender` and `train_receiver` select which role contributes RL traces. A role pinned to an external model is frozen by the framework.\n\nThe product reward means rollouts earn nothing without real compression; early averages will look small — that is headroom, not failure. First establish what prompted models actually produce. No GPU resources have been provisioned and no model evaluation or training has run. Published to the Hub as private `kevin/compact-language` 0.1.0 (verifiers v1) on 2026-10-06.\n\n## Verification performed\n\n- CLI scaffold and eval config dry-run pass.\n- prime-rl config validates; GPU training remains untested.\n- Local model-free checks pass: 800/200 split, disjoint messages, raw whitespace, similarity reward, tag parsing, receiver task isolation. `finalize` was smoke-tested against fake traces: exact-and-short > verbatim > near-miss > garbage > errored.\n- Ruff check and format pass for the environment package and tests; the retained prototype harness keeps its original compact style (two pre-existing lint exceptions).\n- Wheel builds cleanly with `uv build` (editor junk excluded) and is pushed to the Hub as public `kevin/compact-language`.\n\nGrounded against local verifiers `392d596de` and prime-rl `819993faa`. Published versions may not yet expose these interfaces. Full installation and live evaluation gates remain open.\n","encoding":"utf-8","truncated":false,"total_bytes":4722},"status":null}