{"data":{"kind":"file","path":"README.md","version_id":"uvcwb25qkf64fjty27p5mq4w","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":5144,"modified_at":"2026-09-30T02:51:52.419000","content_hash":"c53def19c240a245de3eaf6cbe8131f48acb8593cd0765d0e4b8675833dab7b3"},"entries":[],"content":"# runebench\n\nRuneBench on the Prime Environments Hub (verifiers v1): 40 RuneScape gameplay\ntasks ported from [MaxBittker/runebench](https://github.com/MaxBittker/runebench)\n(MIT). The agent plays an emulated RuneScape private server (rs-sdk / LostCity\nengine) running at 8x speed inside the sandbox, writing and executing\nTypeScript bot scripts through the upstream `rs-agent` MCP server\n(`execute_code` on `bot`/`sdk` globals), with the game wiki bundled in the\nimage for strategy reference.\n\n- **32 skill-XP tasks** — 16 skills x {15, 30} min. Reward: peak real-game\n  XP/min over a single 15-second sampling window (raw XP / 8 game speed / 25\n  server XP rate; windows < 12 s ignored; tracker cadence 15 s).\n- **8 gold tasks** — 4 starting conditions (vanilla, smith-alch, fish,\n  fletch-alch) x {15, 30} min. Reward: peak total coins across inventory +\n  bank (tracker cadence 5 s; max of any sample's gold and the final save read).\n\nNative verifiers v1 environment — no Harbor at runtime. The prime-agent\nharness (autonomous ACP) is the agent seat and consumes the rs-agent game\ntools as an MCP server.\n\n## Sandbox images\n\nThin layers over the upstream pre-built image\n`ghcr.io/maxbittker/rs-agent-benchmark:v71` (engine + gateway + bot client +\nwiki + MCP server + skill tracker, with upstream's anti-tamper model intact:\nroot-owned engine/saves/tracker output, agent user unprivileged):\n\n- `prime/primeintellect/runebench-skill:0.1.0` — `agent.sav` start, 15 s tracker\n  cadence (all 32 skill tasks)\n- `prime/primeintellect/runebench-gold-{vanilla,smith-alch,fish,fletch-alch}:0.1.0`\n  — per-condition starting save, 5 s tracker cadence (the 8 gold tasks)\n\nRebuild with `prime images push runebench-<kind>:<tag> --context <dir>`\n(the image Dockerfiles live with this port; the base image is pinned, never\nrebuilt).\n\n## Upstream fidelity and deviations\n\nPrompts are the upstream `generate-tasks.ts` templates byte-for-byte (verified\nagainst a fresh `bun generate-tasks.ts` run in tests). Scoring is a Python port\nof `shared/check_skill_xp.ts` / `shared/check_gold.ts` with identical math and\nconstants. Deviations, all documented here:\n\n1. **Scoring is native Python reward code**, not the upstream TypeScript\n   verifier scripts run in-sandbox. Same constants and window semantics\n   (`MIN_PEAK_WINDOW_MS = 12000`, /8 /25 normalization, rounded peak; peak-gold\n   = max(final save, any sample)). Final xp/level metadata comes from the last\n   tracker sample instead of a live gateway connection (upstream reads live\n   state for metadata only; the reward math is tracker-derived in both).\n2. **40 of 41 upstream tasks**: the 5-minute woodcutting smoke task is not\n   ported (upstream generates it as a harness sanity check, not a benchmark\n   row).\n3. **5 images instead of 40**: upstream generates one thin `FROM` layer per\n   task (baking save + `SAMPLE_INTERVAL_MS` + `BENCHMARK_DURATION_SECS`);\n   this port collapses to one image per (starting save, tracker cadence).\n   `BENCHMARK_DURATION_SECS` (display-only in the agent-facing\n   `check_xp_rate.ts` progress line) moves to the task's runtime env.\n4. **MCP wiring**: the upstream stdio MCP server (`bun run /app/mcp/server.ts`)\n   is proxied 1:1 to the harness through a colocated vf Toolset adapted from\n   verifiers' `HarborMCPToolset` — the tool catalog, schemas, and the SDK API /\n   MARKET resources are forwarded dynamically from the live upstream server.\n   Harbor itself does not run.\n5. **Game-stack lifecycle**: Prime VMs boot the image as a microVM rootfs under\n   `/sbin/init` — the Docker ENTRYPOINT does NOT run at boot (verified on the\n   platform). The task `setup` therefore starts `/entrypoint.sh` itself when\n   the stack is not already running (engine boot ~90 s on 2 vCPU), applies the\n   upstream anti-cheat (`save-generator` removal), waits for the gateway, and\n   runs the upstream `ensure-services.sh` (root-owned tracker up). The\n   running-stack probe uses bracketed pgrep patterns (`entrypoin[t].sh`) — a\n   bare `pgrep -f entrypoint.sh` inside `sh -c` matches the probe's own\n   command line and reported a phantom running stack in the first hosted\n   smoke. On stack failure, setup raises with the sandbox's process/port/log\n   state embedded in the error. `finalize` ports the upstream\n   `VERIFIER_CLEANUP` (stop ffmpeg, kill orphaned agent bun scripts) before\n   scoring.\n   The setup timeout is 600 s (upstream Harbor allowed `build_timeout = 1200`):\n   the stack boot and the harness's own in-sandbox install (node +\n   prime-agent) share one setup deadline.\n6. Harbor task metadata (author, difficulty, tags) is not carried; the\n   upstream per-model agent adapters (opencode etc.) are replaced by the\n   prime-agent harness seat.\n\n## Run\n\n```bash\nprime eval run primeintellect/runebench --hosted \\\n  -m internal/glm-5.3-fast -n 40 -r 1 \\\n  --env-args '{\"agent\": {\"harness\": {\"id\": \"prime-agent\", \"autonomous\": true}}}'\n```\n\n## Attribution\n\nUpstream: MaxBittker/runebench (MIT) and MaxBittker/rs-sdk (MIT), built on\nLostCityRS/Server. The rs-agent MCP proxy adapts PrimeIntellect-ai/verifiers'\n`HarborMCPToolset` (Apache-2.0).\n","encoding":"utf-8","truncated":false,"total_bytes":5144},"status":null}