{"data":{"kind":"file","path":"README.md","version_id":"oxgewitc1fs2xsknszg0n278","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":4959,"modified_at":"2026-09-23T21:07:11.422000","content_hash":"e0e22b9381487bab09f9014898ea8d3be1d702eb14d46fddd50a294b08e0df12"},"entries":[],"content":"# econarena-eval\n\nEconArena economic-games benchmark as a verifiers v1 Environments-Hub eval\nenvironment.\n\n**PAPER-FAITHFUL REIMPLEMENTATION of the alpha games described in\n\"Economics Arena for Large Language Models\" (Guo et al., arXiv:2401.01735,\n2024). No official code exists** (verified: the paper and every citing paper\nreference only the arXiv preprint; no repo exists on GitHub/Gitee/HuggingFace)\n— the games, session structure, answer formats and prompts are reconstructed\nfrom the paper's Sections 3-4 and Appendices A (games/metrics/architecture)\nand B (verbatim prompt templates). Prompt-fidelity is therefore approximate\nwhere the paper used placeholders; this package documents every reconstructed\nparameter.\n\n## What the model under evaluation plays\n\nThe model plays **every contestant seat** in a session (the paper's\n\"self-competing\" configuration): 5 contestants, one verifiers interaction\neach, the paper's per-run decision prompts, `runs` runs per session\n(default 10; the paper's history experiments use 6+ runs per session with up\nto 3 runs of history revealed).\n\n- **beauty_L / beauty_M / beauty_H** — beauty contest games (Keynes'\n  \"guess 2/3 of the average\"): pick a real number in [0, c̄]; closest to 2/3\n  of the average wins a fixed prize (ties split); unique NE = 0. The\n  paper's range groups: c̄ drawn from [10,100) / [100,1000) / [1000,10000).\n  No game history revealed (the paper's baseline environment).\n- **beauty_hist** — L-range beauty contest **with revealed history** (the\n  paper's B.4 prompt): after run 1 each prompt carries the last 3 runs of\n  everyone's choices. This is the paper's strategic-reasoning /\n  in-context-learning manipulation; convergence over runs is the signal.\n- **auction_L** — minimal second-price sealed-bid (Vickrey) auction cell,\n  the paper's group L (private values s ~ N(50,10), assets 100, entrance fee\n  10 on rule breaks, minimal-id tie-break). Truthful bidding is the unique\n  symmetric NE. **Auctions are otherwise deprioritized in this program —\n  auctions/bargaining are covered by the GTBench port; this one cell is kept\n  only because the paper's second game family is cheap to include.**\n\n## Metrics (paper Section 3.2 / Appendix A.3, as documented components)\n\nComponents in [0,1], higher = more rational; composite **EAS (EconArena\nScore)** = unweighted mean, recorded at reward weight 1 with components at\nweight 0 and all raw quantities + per-run series in `trace.info[\"econarena\"]`:\n\n- `rationality` — the paper's Eq. 1: mean payoff / NE payoff (beauty: prize\n  share vs prize/5 under the all-tie NE; auction: assets vs the run's NE asset).\n- `ne_proximity` — 1 − mean normalized NE-deviation distance. NOTE: the\n  paper's ratio formula `d = π(a)/π(NE) − 1` is **undefined at the\n  beauty-contest NE (π(0) = 0)**, so the deviation is implemented as the\n  paper's deviation *distance*: |action − NE action| / scale (c̄ for beauty,\n  assets for auctions).\n- `win_rate` — beauty: mean prize share won; auction: fraction of runs won.\n- `rule_following` — fraction of runs with a valid in-range parsable action\n  (the paper's rule-breaking frequency, inverted).\n- `convergence` — NE proximity over the LAST THIRD of the runs (the paper's\n  \"convergence rate to the optimal NE strategies\"; the in-context learning\n  signal on history cells).\n\n## Reconstructed parameters (not fixed by the paper)\n\nFixed prize x = 100 (beauty); entrance fee = assets/10 (auction); 5\ncontestants; 10 runs per session default; history window = 3 runs; sampling\ntemperature 0.7 (the paper does not state one). Everything is exposed in\ntask info so episodes are reproducible.\n\n## Running\n\nLocal (model-free stub tests first):\n\n    pytest tests/\n\nLocal live smoke (2-run session, 1 task, chat program runs locally):\n\n    .venv/bin/eval econarena_eval -m internal/glm-5.3-fast -n 1 -r 1 \\\n        --max-concurrent 1 --max-tokens 131072 --no-serve --no-rich \\\n        --env.agent.runtime.type subprocess \\\n        --env-args '{\"taskset\": {\"cells\": [\"beauty_L\"], \"seeds\": [8], \"runs\": 2}}'\n\nHosted (canary probes only; 1 episode, 2-run and 12-run variants):\n\n    prime eval run primeintellect/econarena-eval --hosted \\\n        -m internal/glm-5.3-fast -n 1 -r 1 --max-concurrent 1 \\\n        --max-tokens 131072 --timeout-minutes 60 \\\n        --eval-name econarena-t2-direct --plain \\\n        --env-args '{\"taskset\": {\"cells\": [\"beauty_L\"], \"seeds\": [8], \"runs\": 2}}'\n\nPrimeAgentHarness mode: add `\"agent\": {\"harness\": {\"id\": \"prime-agent\"}}` to\n`--env-args` (seats then run as Prime Agent daemons in a container runtime).\n\n## Provenance and licensing\n\nThis package contains **no vendored upstream code** (none exists). The engine\n(`econarena_eval/engine.py`) is an original reimplementation of the paper's\nrules and prompts, MIT-licensed here. Cite the benchmark as: Guo, Bu, Wang,\nRen, Sui, Shang, Lu. \"Economics Arena for Large Language Models.\"\narXiv:2401.01735 (2024).\n","encoding":"utf-8","truncated":false,"total_bytes":4959},"status":null}