{"data":{"kind":"file","path":"README.md","version_id":"owjw2vsd2kpnd99lgi217yql","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":2735,"modified_at":"2026-09-21T21:43:38.066000","content_hash":"8570994f7f5007e10e7a8d9d6f92cb00316d4277aab74fd849211733836719d6"},"entries":[],"content":"# agent-bazaar-eval\n\nAgent-Bazaar (THE_CRASH + LEMON_MARKET) as a verifiers v1 Environments-Hub eval\nenvironment. Vendored from https://github.com/CameronCrow/Agent-Bazaar (MIT).\n\n## What the model under evaluation plays\n\n- **crash_baseline / crash_stabilizing**: all 5 B2C firms — one seat per firm,\n  one verifiers interaction per seat, the benchmark's own per-step pricing\n  prompts (supply / production / price in one JSON decision per step).\n- **lemon_main / lemon_norep**: all 12 C2C buyers — one seat per buyer; the\n  benchmark's bid prompts (bid/pass on visible listings) and post-purchase\n  review prompts (upvote/downvote). Sellers and the Sybil cluster are fed from\n  the bundled listing corpus, so only buyers make model calls.\n\nEach task = one full market episode for one (scenario cell, seed); 4 cells x\n3 seeds = 12 tasks. Rewards recorded per seat trace: `eas` (composite,\nweight 1) + `eas_<component>` (weight 0), with raw quantities and per-step\nseries in `trace.info[\"bazaar\"]`.\n\n## Scenario cells (benchmark-experiment fidelity)\n\n- crash cells replicate `scripts/exp1.py` `_BASE_FIXED` (`THE_CRASH`, 5 LLM\n  firms, 50 CES consumers, dlc=3, io prompts, no diaries), with\n  `--num-stabilizing-firms 0|3`.\n- lemon cells replicate `scripts/exp2.py` `_BASE_FIXED` (`LEMON_MARKET`,\n  12 buyers, K=3 sybils, rho_min 0.3, rep 0.8, dlc=3), plus the bundled\n  listing corpus; `lemon_norep` adds `--no-buyer-rep`.\n- 50 timesteps per episode by default (crash dynamics emerge within that\n  horizon; lemon keeps the benchmark default).\n\n## Reproducibility\n\nSeeds are applied exactly as `agent_bazaar.main` does. The benchmark draws\nfrom process-global RNGs during stepping, so run episodes one-per-process\nfor exact reproducibility: `--serve.max-concurrent 1`, or `--max-concurrent 1`\nin-process. Throughput then scales with pool workers (processes).\n\n## Running\n\nLocal:\n\n    prime eval run agent_bazaar_eval --env-dir-path . -m internal/glm-5.3-fast \\\n        -n 12 -r 1 --max-concurrent 2 --max-tokens 131072\n\nHosted (canary first, then the full 12):\n\n    prime eval run primeintellect/agent-bazaar-eval --hosted \\\n        -m internal/glm-5.3-fast -n 1 -r 1 --max-concurrent 1 \\\n        --max-tokens 131072 --timeout-minutes 60 --eval-name bazaar-canary --plain\n\n## Vendoring notes (agent_bazaar)\n\n- `agent_bazaar/` is vendored top-level (absolute `agent_bazaar.*` imports).\n- Patched: `agents/llm_agent.py` imports the provider model clients lazily;\n  `models/__init__.py` ships only `BaseLLMModel`; `main.py` strips wandb and\n  is used only for `create_argument_parser()`.\n- `agent_bazaar_eval/corpus/listing_corpus.json` is the repo's pre-compiled\n  listing corpus (15.5 MB) so lemon episodes need no seller LLM calls.\n","encoding":"utf-8","truncated":false,"total_bytes":2735},"status":null}