{"data":{"kind":"file","path":"README.md","version_id":"lmdiycyd80noz3gkod6frma7","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":2947,"modified_at":"2026-09-21T22:20:58.842000","content_hash":"ffb4bd0a43f63874693908db61b75e4d532bba671834db23ee7b67b997291be2"},"entries":[],"content":"# GameDevBench — Prime Environments Hub package (eval-only)\n\nPort of https://github.com/waynchi/gamedevbench (ICML 2026): 333 Godot 4.4.1\ngame-development tasks, scored by the benchmark's official headless Godot\nvalidation (per-task scripts/test.gd + scenes/test.tscn printing\nVALIDATION_PASSED / VALIDATION_FAILED).\n\n**EVAL-ONLY: ground truths and tutorial sources are public. Do not train on\nthese tasks.**\n\n## Layout\n\n- `gamedevbench/taskset.py` — verifiers v1 Taskset + Task (setup + `@reward`\n  scoring that mirrors the benchmark's strict-confinement validation flow)\n- `gamedevbench/tasks_meta.json` — per-task instruction / display flag\n- `Dockerfile` — task image: Godot 4.4.1 (exact), xvfb/xauth, the 333 official\n  task zips at /opt/gamedevbench/tasks (pinned to gamedevbench@3a0dd3c)\n\n## Deviations from the official benchmark (documented)\n\n1. **Prompt, text-only models**: the official text-mode prompt line\n   \"You are a visual agent and can use images and videos to help you\n   understand the state of the game.\" is replaced by an explicit text-only\n   instruction (\"Do not use attach_image, generate screenshots, or create\n   movie files. Verify ... through headless Godot command output, print\n   statements, and text logs.\"). Reason: with glm-5.3-fast (no vision), the\n   official line drives agents to attach screenshots; requests carrying image\n   attachments upstream-500 and kill the episode after one retry (verified on\n   two hosted canaries, 2026-09-21). All other prompt bytes are identical to\n   `gamedevbench/src/utils/prompts.py::create_task_prompt` with\n   `use_runtime_video=False, use_mcp=False`.\n2. **Harness binding**: the env config binds the agent seat to\n   `vf.AgentConfig(harness={\"id\": \"prime-agent\", \"autonomous\": True})`.\n   glm-5.3-fast's first turn is often reasoning-only with no tool calls;\n   non-autonomous rollouts end after one call (verified on canary 1).\n   Autonomous continuation keeps rollouts alive until task completion or the\n   1800s agent timeout.\n\n## Flow per task\n\n1. `setup()`: unzip the task into /workspace under the official sandbox filter\n   (no test*, task_config.json, *.log, *.md, hidden files/dirs), write the\n   minimal task_config.json, headless `--import`.\n2. Prime Agent (ACP) solves the task in /workspace with the official\n   text-mode prompt.\n3. `godot_validation` reward: copy the workspace, restore test.gd/test.tscn\n   from the pristine zip, headless `--import`, run scenes/test.tscn\n   (xvfb-run for the 10 display-required tasks), parse VALIDATION_PASSED.\n\n## Run\n\n```bash\nprime images push gamedevbench:0.1.0 --path . --dockerfile Dockerfile --plain\nprime env push --path . --name gamedevbench --visibility PRIVATE --runtime v1 --plain\nprime eval run primeintellect/gamedevbench --hosted \\\n    -m internal/glm-5.3-fast -n 333 -r 1 --max-concurrent 32 \\\n    --max-tokens 131072 --timeout-minutes 300 \\\n    --eval-name gamedevbench-glm53fast-001 --plain\n```\n","encoding":"utf-8","truncated":false,"total_bytes":2947},"status":null}