{"data":{"kind":"file","path":"README.md","version_id":"ssexso99qj50aze1ytbc581d","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":4697,"modified_at":"2026-09-23T20:56:48.150000","content_hash":"19bfccba9c41eb8f6decfd32bf42bbebc30c5cea8f90a00a593b55f87e73260a"},"entries":[],"content":"# llm-economist-eval\n\nLLM Economist (Stackelberg tax-policy world, https://github.com/sethkarten/LLM-Economist,\narXiv:2507.15815) as a verifiers v1 Environments-Hub eval environment.\nVendored from the MIT-licensed upstream repo.\n\n## What the model under evaluation plays\n\nEach task = one full tax-policy episode for one (scenario cell, seed);\n6 cells x 3 seeds = 18 tasks. One verifiers interaction (seat) per LLM sim\nagent; each benchmark decision prompt is one seat turn.\n\n- **stackelberg_llm**: the TaxPlanner seat (bracket-delta tax schedules every\n  tax year from worker (z, utility) histograms) + all 5 LLM worker seats\n  (weekly labor hours in [0,100]).\n- **planner_fixedpop**: only the planner seat; procedural FixedWorkers.\n- **workers_saez / workers_usfed**: only worker seats, governed by the\n  analytic Saez schedule (three brackets, elasticity 3.0) or the statutory\n  2024 US federal schedule (seven brackets).\n- **bounded_personas**: planner + persona worker seats; workers answer a\n  per-step satisfaction reflection (utility scaled by 1.0 YES / 0.5 NO).\n- **democratic_platforms**: worker seats only; every tax year workers\n  campaign with a platform (proposed deltas), vote, and the elected leader\n  sets the tax deltas.\n\nRewards recorded per seat trace: `econ` (composite, weight 1) +\n`econ_welfare` / `econ_productivity` / `econ_equity` (weight 0), with raw\nquantities and per-step series in `trace.info[\"llm_econ\"]`.\n\n## Metrics\n\nAll components in [0,1], higher = better, evaluated over the terminal tax\nyear (see `metrics.py` docstring for the definitions and caveats):\n\n- **welfare**: mean over agents of clip01(u / max(z, 1)) — the benchmark's\n  equity-weighted SWF per capita, per-agent clipped at 1 (unclipped SWF in\n  trace.info raw).\n- **productivity**: mean labor / 100.\n- **equity**: 1 - Gini(consumption) at episode end (consumption = post-tax\n  income + rebate).\n\nComposite = unweighted mean of the three.\n\n## Scenario cells (experiment fidelity)\n\nCells mirror `experiments/run_experiments.py` base configs (rational\nscenario, `us_income` GB2 skills, history-len 50, io prompts, temperature\n0.7). Scaling deviation, documented per the paper: the paper's central\nconfig is N=100 / T=3000 / K=128 tax years; hosted full-effort episodes must\nstay eval-sized, so cells default to N=5 / T=50 / K=25 (two tax years = one\nin-context planner revision per episode). Vendored patches: wandb/torch\nstripped; provider model clients lazily imported; the fixed-persona path\nreturns a proper dict (upstream returned a bare list, which crashed\nGEN_ROLE_MESSAGES.update) and skips the census CSV that the upstream repo\ndoes not ship; `--timeout` (JSON-retry budget per decision, vendored default\n10) is actually forwarded to agents (upstream parses it but drops it) and\npinned to 3 in all cells.\n\n## Reproducibility\n\nSeeds are applied exactly as `llm_economist.main` does (numpy + python\nglobals before construction). The benchmark draws from process-global RNGs\nduring stepping and shares the module-global `GEN_ROLE_MESSAGES` persona map,\nso run episodes one-per-process for exact reproducibility:\n`--serve.max-concurrent 1`, or `--max-concurrent 1` in-process. Throughput\nthen scales with pool workers (processes).\n\n## Running\n\nLocal:\n\n    .venv/bin/eval llm_economist_eval --env-dir-path . \\\n        -m internal/glm-5.3-fast --env.agent.runtime.type subprocess \\\n        -n 1 -r 1 --max-concurrent 1 --max-tokens 131072 --no-serve --no-rich\n\nHosted (canary first; probe overrides via --env-args):\n\n    prime eval run primeintellect/llm-economist-eval --hosted \\\n        -m internal/glm-5.3-fast -n 1 -r 1 --max-concurrent 1 \\\n        --max-tokens 131072 --timeout-minutes 60 \\\n        --eval-name econ-canary --plain \\\n        --env-args '{\"taskset\": {\"cells\": [\"stackelberg_llm\"], \"seeds\": [8],\n                                  \"timesteps\": 2, \"two_timescale\": 1}}'\n\nPrimeAgentHarness mode: add `\"agent\": {\"harness\": {\"id\": \"prime-agent\"}}`\nto --env-args (and pin `\"max_concurrent_agents\": 4`, `\"runtime\": {\"cpu\": 8,\n\"memory\": 8}` for seat-heavy episodes).\n\n## Vendoring notes (llm_economist)\n\n- `llm_economist/` is vendored top-level (absolute `llm_economist.*` imports).\n- Patched: `main.py` strips wandb + the torch bootstrap (kept only for\n  `create_argument_parser()`); `models/__init__.py` ships only `BaseLLMModel`;\n  `agents/llm_agent.py` + `agents/worker.py` import provider clients lazily;\n  `agents/worker.py::distribute_personas` fixed-persona path returns\n  `{persona_i: description}` without the missing census CSV.\n- The eval env drives the episode loop itself (`env.py::_drive` mirrors\n  `run_simulation` step by step); `main.run_simulation` is never called.\n","encoding":"utf-8","truncated":false,"total_bytes":4697},"status":null}