{"data":{"kind":"file","path":"README.md","version_id":"cczae4hwggmb3uab5pp6z6kk","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":9217,"modified_at":"2026-10-01T06:51:39.373000","content_hash":"6af1217ed6230d02d41f5f713bce080bc75d773d5f42c5dbbd56ead1d225eb50"},"entries":[],"content":"# factorio-play\n\nfactorio-play is an agentic environment built on Factorio, the\nfactory-building game. The model controls a player character through tool\ncalls, one action at a time: it looks around, walks, places machines, fuels\nthem, and sees the result of every action before choosing the next. When it\nthinks the factory is ready it calls `finish`, the game runs a measurement\nwindow, and the reward is what the factory actually produced.\n\nThis makes it a test of closed-loop behaviour. The model has to explore a map\nit only partly sees, notice when an action is refused or doesn't do what it\nexpected, and recover, all within a fixed budget of decisions. The larger task\ntakes around 90 tool calls even when played well.\n\nThe maps run in [factory-sim](https://github.com/divagr18/factory-sim), a\ntick-exact simulator of Factorio's early game that is checked against the real\ngame by [FactorioGym](https://github.com/divagr18/FactorioGym). Every action a tool\ntakes goes through the simulator's own API, so the rules and the budget are\nexactly those of a trained policy.\n\nIf you want to train program synthesis instead, where the model writes one\nPython program that builds the factory and is scored on maps it has never\nseen, use the companion environment\n[factorio-build](https://app.primeintellect.ai/dashboard/environments/divagr/factorio-build).\nIt uses the same tasks, maps and scoring.\n\nIt requires `verifiers>=0.3.1` and `factory-sim>=0.3.0`, both installed from\nPyPI with prebuilt wheels. It runs on Linux and macOS; `verifiers.v1` does not\nimport on Windows.\n\n## Tasks\n\nThere are two tasks. Choose one with `sim_task`.\n\n| Task | The model has to | Reward | Decisions |\n| --- | --- | --- | --- |\n| `construct_smelting_line` (default) | walk to an iron patch, place a burner drill on ore and a stone furnace under the drill's drop point, and fuel both | 1 if at least 10 iron plates are made by machine during a verification minute | 600 |\n| `belt_smelting` | walk between an iron patch, a coal patch and a wooden output chest, 20 to 40 tiles apart, and build a line of drills, furnaces, transport belts and burner inserters that carries plates into the chest | min(1, plates / 150), counting the iron plates that reach the chest during a ten-minute window in which the model cannot act; 1 exactly when the task succeeds | 2,500 |\n\n`belt_smelting` starts with 4 burner drills, 4 stone furnaces, 40 belts, 10\nburner inserters and 20 coal. Twenty coal runs a line for about 79 plates, so\nthe line also needs coal from the coal patch, mined by hand or by a drill.\nPlates smelted from hand-mined ore do not count. Its tools are factory-sim's\n`WorldV3`: public markers for the three sites, belts, inserters and chests,\n`rotate`, hand-mining, `take_fuel` and belt lanes. The prompt carries the same\ntask text and measured mechanics that `factorio-build` shows a program writer.\n\n## Taskset\n\n- **Source:** scenes generated by `fsim.scenes.sample`, the same seed plan and\n  splits as `factorio-build`. One task is one scene.\n- **Splits:**\n  - `train`: seeds 0-99,999, the training layout families (four for\n    `construct_smelting_line`; `open`, `walled` and `split_patch` for\n    `belt_smelting`).\n  - `val`: seeds from 100,000 on, same families.\n  - `holdout`: families neither of the others contains (`obstructed_patch`;\n    `obstructed` and `far_chest`). Evaluation only: it is built only when\n    asked for by name.\n- **Tools, `construct_smelting_line`:** these cost no decision: `me`, `tile`,\n  `inventory`, `ore_tiles(kind)`, `blocked_tiles`, `entities`, `patch`,\n  `decisions_left` and `last_refused`. These spend decisions from a\n  600-decision budget: `move(direction, stride)`, `place(item, x, y, facing)`,\n  `give(row, item, amount)`, `take(row, item, amount)`, `mine(row)` and\n  `wait(count)`. `finish()` ends the build phase. Unchanged from 0.1.1.\n- **Tools, `belt_smelting`:** these cost no decision: `me`, `tile`,\n  `inventory`, `ore_tiles(kind)`, `blocked_tiles`, `entities(kind)`,\n  `marker(name)` (`iron`, `coal`, `output`), `belt_lanes(row)`, `opened`,\n  `decisions_left` and `last_refused`. These spend decisions from a\n  2,500-decision budget, exactly as `WorldV3` counts them:\n  `move(direction, stride, count)`, `place(item, x, y, facing)` for drills,\n  furnaces, belts, inserters and chests, `rotate(row, reverse)`,\n  `give(row, item, amount)`, `take(row, item, amount)`,\n  `take_fuel(row, amount)`, `mine(row)`, `mine_resource(x, y, amount, wait)`,\n  `inspect(row)`, `wait(count)` and `finish()`. `inspect` opens a chest within\n  reach and returns what it holds by item; `opened` reads it again for free\n  until the character walks out of reach.\n- **Entities** are named by their current `row` in `entities()`.\n- **Reward:** `success` for `construct_smelting_line`, 1.0 if `finish` was\n  called and the line verified. `score` for `belt_smelting`, factory-sim's own\n  score for the task: min(1, plates delivered / 150). A rollout that never\n  calls `finish` scores 0, because nothing was verified. Metrics: `finished`,\n  `verified_output`, `decisions`, `refusals`, `tool_calls`, and `success` for\n  `belt_smelting`.\n- **Runtime:** each rollout gets its own MCP server process, simulator and\n  scene. Every world action is one method call on factory-sim's `World` or\n  `WorldV3`, so the budget, refusals and verification are exactly those of a\n  program in `factorio-build`. The tests check this decision for decision.\n- **Harness:** the default agent (`FactorioPlayHarness`) calls the model through\n  the OpenAI **Responses API**, so reasoning models can think and call tools in\n  the same turn. Its reasoning is carried from one turn to the next. A refused\n  action returns the reason, and directions take `N/E/S/W` or words.\n\n## Long episodes\n\n`belt_smelting` is long. The reference builder, `fsim.belt_expert`, spends\nabout 275 to 370 decisions per scene, some 160 of them waiting for hand-mined\ncoal. So its tools fold runs of world actions into one call, each action still\ncounted by `WorldV3`:\n\n- `move(direction, stride, count)` takes up to `count` strides and stops after\n  one that does not move the character;\n- `wait(count)` waits `count` decisions;\n- `mine_resource(x, y, amount)` sends requests of 20, 5 and 1, each followed by\n  the waits for its items, so 40 coal is one call of 162 decisions;\n- `give`, `take` and `take_fuel` move any amount as transfers of 20, 5 and 1.\n\nReplies are compact JSON, and `entities()` leaves out the fields that hold\nnothing. The reference build, played through the tools on four `train` scenes,\ntakes 86 to 98 tool calls for 287 to 311 decisions. Its whole context, with\nthe system prompt and the tool schemas, is 12,000 to 13,300 tokens under\nQwen3.5's chat template, before any reasoning. A model that reads more, or\nrecovers from mistakes, uses more: set `max-turns` generously.\n\n## Run it\n\n```bash\nuv run eval factorio-play -m openai/gpt-6-luna -n 8 --env.agent.max-turns 400\nuv run eval factorio-play -m openai/gpt-6-luna -n 8 --env.taskset.sim-task belt_smelting --env.agent.max-turns 800\nuv run eval factorio-play -m openai/gpt-6-luna -n 8 --env.taskset.split holdout --env.agent.max-turns 400   # evaluation only\n```\n\n## First run\n\n`gpt-6-luna`, one `construct_smelting_line` `val` scene, reasoning on: the line\nwas built and verified (15 plates) in 12 model turns and 17 tool calls, with\nno refused action. One episode shows the environment works end to end. It is\nnot a benchmark number.\n\n## Config\n\n| Field | Default | Description |\n| --- | --- | --- |\n| `sim_task` | `construct_smelting_line` | `construct_smelting_line` or `belt_smelting` |\n| `split` | `train` | `train`, `val` or `holdout` (evaluation only) |\n| `num_examples` | 64 | Scenes, one per task |\n| `seed` | 0 | First scene index of the split |\n| `game_notes` | true | Include notes on mechanics: drop point, reach, fuel, belts and inserters |\n| `task.tools.decision_budget` | 600 | `construct_smelting_line`'s budget |\n| `task.belt_tools.decision_budget` | 2,500 | `belt_smelting`'s budget |\n\n## Changelog\n\n- 2026-10-01 (0.3.0): `belt_smelting` follows FactorioGym 1.2.0, whose\n  generator keeps scenes off the map's lake (the roughly 1% of scenes that\n  touched water are redrawn; the rest are unchanged), and gains the `inspect`\n  and `opened` tools. `construct_smelting_line` is unchanged. Requires\n  factory-sim 0.3.0.\n- 2026-09-28 (0.2.0): Adds `belt_smelting` (FactorioGym `belt_smelting`\n  1.1.0) with factory-sim's `WorldV3` as tools, its train, val and holdout\n  splits, and its score, min(1, plates / 150), as the reward. Its tools that\n  take a count or an amount run several world actions per call, and their\n  replies are compact JSON. `construct_smelting_line`'s tools, prompts and\n  reward are unchanged. Fixes its action replies, which stopped carrying the\n  refusal reason after an episode's twentieth action. Requires factory-sim 0.2.0.\n- 2026-09-24 (0.1.1): The default harness uses the Responses API. Refused\n  actions return their reason, and directions and facings accept words\n  (`east`) as well as letters (`E`). Before this, every action a model phrased\n  with a word was refused without explanation.\n- 2026-09-24 (0.1.0): Initial release.\n","encoding":"utf-8","truncated":false,"total_bytes":9217},"status":null}