{"data":{"kind":"file","path":"README.md","version_id":"k6lgtpvva7fuspv36t44hjvj","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":7635,"modified_at":"2026-09-23T08:42:28.987000","content_hash":"87e1e3b2275c17848c2d2a80529118c73b5d1b461e7335f90035c5c0c1a79088"},"entries":[],"content":"# emulatorbench\n\nEmulator Bench asks agents to implement deterministic Rust emulators for 16\nclassic computer and console platforms. This package contains the public native\nVerifiers v1 taskset, source manifests, starter workspace, feedback verifier,\nand scoring code. It is harness-agnostic and contains no private harness,\ncredentials, private artifact payloads, or private download locations.\n\n## Canonical interface\n\n- Distribution and taskset ID: `emulatorbench`\n- Python module: `emulatorbench`\n- Intended Environments Hub identity: `primeintellect/emulatorbench`\n- Task data/config classes: `EmulatorBenchData`, `EmulatorBenchTaskConfig`,\n  and `EmulatorBenchConfig`\n- Task class: `EmulatorBenchTask`\n- Taskset class: `EmulatorBenchTaskset`\n- Task/platform IDs: `emulatorbench-<platform>`\n- Sandbox label metadata: exactly `emulatorbench`\n- Standard Prime Agent evaluation: [`PRIME_EVAL.md`](PRIME_EVAL.md), using\n  `PLATFORM=<slug> ./scripts/run_prime_eval.sh`\n\nThe runtime image is configurable with `--taskset.image` or\n`EMULATORBENCH_IMAGE`. The default is the public digest-pinned\n`rust:1.85-bookworm@sha256:e51d0265072d2d9d5d320f6a44dde6b9ef13653b035098febd68cce8fa7c0bc4` image. External-controller grading rejects mutable image\nreferences. Runtime and harness selection remain caller-owned.\n\n## Platforms\n\n| Platform slug | Task ID | System | Level | Declared units |\n|---|---|---|---:|---:|\n| `chip8` | `emulatorbench-chip8` | CHIP-8 | 1 | 8 |\n| `i8080_space_invaders` | `emulatorbench-i8080-space-invaders` | Intel 8080 / Space Invaders | 2 | 5 |\n| `gameboy_dmg` | `emulatorbench-gameboy-dmg` | Nintendo Game Boy (DMG) | 3 | 2,580,686 |\n| `nes` | `emulatorbench-nes` | Nintendo Entertainment System | 4 | 293 |\n| `sms` | `emulatorbench-sms` | Sega Master System | 5 | 1,604,017 |\n| `gameboy_cgb` | `emulatorbench-gameboy-cgb` | Game Boy Color | 6 | 2,584,110 |\n| `gba` | `emulatorbench-gba` | Game Boy Advance | 7 | 57,571 |\n| `genesis` | `emulatorbench-genesis` | Sega Genesis / Mega Drive | 8 | 2,604,060 |\n| `snes` | `emulatorbench-snes` | Super Nintendo | 9 | 5,376,331 |\n| `ps1` | `emulatorbench-ps1` | Sony PlayStation | 10 | 124 |\n| `c64` | `emulatorbench-c64` | Commodore 64 | 11 | 1,544 |\n| `n64` | `emulatorbench-n64` | Nintendo 64 | 12 | 698 |\n| `psp` | `emulatorbench-psp` | PlayStation Portable | 13 | 433 |\n| `ps2` | `emulatorbench-ps2` | PlayStation 2 | 14 | 78 |\n| `zx_spectrum` | `emulatorbench-zx-spectrum` | ZX Spectrum 48K | 15 | 1,605,358 |\n| `game_gear` | `emulatorbench-game-gear` | Sega Game Gear | 16 | 1,604,027 |\n\n## Source and score integrity\n\nGit sources are pinned to full commits and fetched with an exact-commit shallow\nfetch. Downloadable HTTP artifacts are SHA-256 pinned. Documentation references\nare not treated as executable artifacts. Package load validates IDs, source\ncoverage, private descriptor boundaries, unique case metadata, and declared unit\ncounts.\n\nEvery scored declaration has an explicit expected-unit denominator. Missing\nobservations, inputs, or oracles never shrink that denominator. Runner-v1 output\nis retained only as an untrusted development diagnostic and always receives zero\ntrusted reward. Positive reward requires a deeply validated runner-v2 response,\na signed controller-owned input/oracle corpus, and a score record bound to that\ncorpus digest and version.\n\nThe public taskset does not accept private artifact paths. Private payloads and\noracles remain external-controller extensions. Public manifests contain\nnon-payload descriptors only. Candidate and grader workspaces, submissions,\nresults, lifecycle events, patches, and logs are retained under restricted host\ndirectories with credential exclusion and text redaction before teardown.\n\n## Feedback boundary\n\nSame-UID public feedback is either the explicit `full` verifier or disabled.\nAggregate and hidden modes are rejected because a candidate workspace cannot\nhide verifier inputs or enforce budgets.\n\nWith `anti_cheat_validation`, the candidate workspace receives no verifier,\nsource manifest, expected output, private payload, or signing material. When\npublic feedback is enabled, a host watcher transfers deterministic candidate\narchives to a distinct digest-pinned no-egress grader, then publishes bounded\ncontroller-signed count, diagnostic, or full feedback. Records bind task, trace,\nattempt, mode, raw archive, sanitized submission, retained grader result, and\noracle corpus. The candidate helper verifies the signature and exact raw archive\nbefore displaying feedback. Disabling public feedback keeps the same signed\nfinal-grade path without staging a helper or starting a watcher.\n\nEvery grader gets a distinct provider identity, a bounded start, staging,\nscoring, capture, and confirmed teardown lifecycle, a provider hard-lifetime\nbackstop, and restricted pre-teardown artifacts that exclude trusted `/tmp`\nverifier and oracle files. Any identity, egress, signing, artifact, lifecycle,\nresult-digest, corpus, or confirmed-teardown failure zeros trusted reward.\n\n`emulator-runner-v2` accepts input-only ordered requests and returns bounded raw\nobservations without pass or score claims. The controller keeps expected states\nand comparisons separate. The package currently includes strict schemas,\nsigned-corpus verification, fixed plans for all 13 source-separable declarations,\nand fail-closed classifications for all 65 declarations. No declaration is\nmarked ported until its complete signed corpus has been built and installed.\n\n## Safe validation\n\nImport and construction do not launch a runtime:\n\n```bash\npython - <<'PY'\nfrom verifiers.v1.loaders import taskset_class, taskset_config_type\nprint(taskset_class(\"emulatorbench\"))\nprint(taskset_config_type(\"emulatorbench\"))\nPY\n```\n\nA real evaluation requires an isolated container runtime. Do not launch one\nwithout an explicit run plan.\n\n## Provenance reconciliation\n\n`PROVENANCE_RECONCILIATION.md` records the reviewed implementation origin,\nintentional cleanup differences, and the handoff point for the CPU-run source\ncomparison. Behavioral differences not listed there require reconciliation\nbefore approval.\n\n## Changelog\n\n- 2026-07-16: Added the canonical public Emulator Bench taskset from the\n  reviewed PR #662 implementation, removed legacy harness/runtime compatibility\n  paths, pinned public sources, corrected scoring denominators/composition, and\n  added redacted terminal artifact capture.\n- 2026-07-17: Restored the external signed-controller feedback loop, added fresh\n  confirmed-teardown graders and complete restricted artifact retention, bound\n  scores to signed runner-v2 corpora, and made every unported runner-v1\n  declaration fail closed.\n- 2026-07-24: Ported `chip8:chip8_timendus_suite` as the first\n  `signed_corpus_installed` runner-v2 plan, with immutable Timendus fixtures,\n  a digest-pinned packaged verification key, bounded signed feedback, and\n  lifecycle-gated trusted scoring.\n\n\n## Deviations\n\n- **Offline-build disclosure (added 2026-09-23, v0.1.10):** the task instruction now\n  discloses that the grading build runs fully offline (`cargo build --locked`, no crate\n  fetches) and that external crates must be vendored or avoided. Rationale: the original\n  instruction invited adding dependencies without disclosing the offline grading\n  constraint, and the candidate sandbox's network access let in-session builds pass while\n  only the no-egress grader failed — an invisible trap (observed: glm psp 0.0 in run 008;\n  gpt-6-astra systematic 0.0s on chip8/sms/psp with in-session public scores up to 99.9989%).\n  The scoring contract is unchanged; the fix makes the constraint agent-visible.\n","encoding":"utf-8","truncated":false,"total_bytes":7635},"status":null}