{"data":{"kind":"file","path":"README.md","version_id":"nc7rfp0kjgscw2jmsqa8ekjo","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":4670,"modified_at":"2026-09-23T20:57:11.878000","content_hash":"d485145d96c9bcd3b79d38a8cd455589aa80f7f1141496f61aeeba966103bfbe"},"entries":[],"content":"# gtbench-eval\n\nGTBench (\"GTBench: Uncovering the Strategic Reasoning Limitations of LLMs\nvia Game-Theoretic Evaluations\", NeurIPS 2024,\nhttps://github.com/jinhaoduan/GTBench) — economic game subset — as a\nverifiers v1 Environments-Hub eval environment. Vendored from the\nbenchmark (MIT) with its engine re-implemented on `open-spiel==1.4`\n(the benchmark's own pin) as a single-seat direct-model environment.\n\n## What the model under evaluation plays\n\nOne task = one OpenSpiel match driven with GTBench's own language\nprotocol (system prompt, per-turn observation + step prompt, regex\naction parsing, last match wins). The model plays ONE seat; the other\nseat is a scripted opponent (no model calls). One verifiers interaction\nper episode carries the whole match transcript.\n\n| game | opponent | what it measures |\n|---|---|---|\n| `first_sealed_auction` | equilibrium bidder `b(v)=v/2` | bid shading vs. the FPSBA equilibrium |\n| `negotiation` | greedy half-pie bargainer | multi-issue bargaining: deals, own-value share |\n| `kuhn_poker` | Kuhn equilibrium mixture (α=1/3) | imperfect-information betting |\n| `liars_dice` | minimum-raise bidder | bluffing / probability under hidden dice |\n| `python_iterated_prisoners_dilemma` | tit-for-tat | repeated-game cooperation vs exploitation |\n\nDefault suite: 5 games x 2 model seats x 3 seeds = 30 tasks\n(`first_sealed_auction`, `negotiation`, `kuhn_poker`, `liars_dice`,\n`python_iterated_prisoners_dilemma`). GTBench has **no public-goods\ngame**; the non-economic board games (tictactoe, connect4, breakthrough,\nnim, pig) are out of scope per the port brief.\n\n## Scoring\n\nPer-game components in [0,1] (composite = unweighted mean, recorded as\nreward `gtbench` weight 1.0 on the model seat trace; components at\nweight 0; all raw quantities + the full step log in\n`trace.info[\"gtbench\"]`):\n\n* auction: `valid` (normal completion), `won`, `surplus` = (v − price)/v\n* negotiation: `valid`, `deal`, `util_share` = own deal utility / own pie\n* kuhn_poker: `valid`, `win` (1/0.5/0), `payoff_norm` = (r+2)/4\n* liars_dice: `valid`, `win`, `payoff_norm` = (r+1)/2\n* PD: `valid`, `payoff_share` = r/(r+o), `abs_norm` = r/(5·rounds)\n\nAn invalid/unparseable action ends the match Abnormal (GTBench\nsemantics) and scores `valid` = 0.\n\n## Faithfulness notes (upstream quirks kept on purpose)\n\nPorted verbatim from `gamingbench/games/*` and `gamingbench/prompts/*`:\n\n* kuhn_poker renders past-move roles with GTBench's modulo formula\n  (attribution is wrong for seat 1 — upstream behavior);\n* the PD past-round rendering mirrors GTBench's (opponent history is\n  rendered for both players);\n* negotiation accepts only the strict `[a, b, c]` bracket spacing of\n  upstream's inner regex (`[1,2,3]` fails the move — benchmark-faithful);\n* liars_dice parses both dice from the debug state string (upstream\n  does too) but the prompt only shows the model its own die;\n* negotiation runs with upstream's default (fixed) valuation vectors.\n\nDeviations (documented): chance nodes sample from a dedicated\n`np.random.default_rng(seed)` (upstream uses the process-global numpy\nRNG); GTBench calls the model stateless per move, this port keeps one\ngrowing seat transcript per match (the per-turn prompts already render\nthe benchmark's visible history); the LLM-vs-LLM win-rate/Elo scoring of\nthe paper is replaced by the single-seat scripted-opponent components\nabove.\n\n## Running\n\nLocal smoke (subprocess runtime, no Prime tunnel):\n\n    .venv/bin/eval gtbench_eval -m internal/glm-5.3-fast -n 2 -r 1 \\\n        --max-concurrent 1 --max-tokens 131072 --no-serve --no-rich \\\n        --env.agent.runtime.type subprocess \\\n        --env-args '{\"taskset\": {\"games\": [\"first_sealed_auction\"], \"seeds\": [8], \"seats\": [0]}}'\n\nHosted canary / probes:\n\n    prime eval run primeintellect/gtbench-eval --hosted \\\n        -m internal/glm-5.3-fast -n 1 -r 1 --max-concurrent 1 \\\n        --max-tokens 131072 --timeout-minutes 60 --eval-name gtbench-canary --plain\n\nPrimeAgentHarness mode: pin the seat harness via `--env-args\n'{\"agent\": {\"harness\": {\"id\": \"prime-agent\"}}}'` (verified pattern from\nagent-bazaar-eval).\n\n## Model-free tests\n\n    .venv/bin/python -m pytest tests/ -q\n\nStub seats drive every game to terminal without model calls (canned\nlegal actions), covering: match wiring, forced auto-moves, invalid-move\nAbnormal ends, the empty-reply guard, metric bounds, and the\ndeterministic negotiation deal vs. the greedy opponent.\n\n## Traces + audit\n\n`scripts/audit_eval.py <eval-id> [--json out.json]` folds traces into\nepisodes per task and aggregates per game (means of composite +\ncomponents), reporting harness id, stop conditions, and usage.\n","encoding":"utf-8","truncated":false,"total_bytes":4670},"status":null}