{"data":{"kind":"file","path":"README.md","version_id":"zs7fluaz83gzxc470o05uxhy","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":7521,"modified_at":"2026-09-24T00:06:47.093000","content_hash":"10f02d08fc20fb38fb8878f76856ccc9bf089a0c03d6b1876cfaf036b14d8bca"},"entries":[],"content":"# magentic-marketplace-eval\n\nMagentic Marketplace (multi-agent agentic-markets benchmark) as a verifiers v1\nEnvironments-Hub eval environment. Vendored from\nhttps://github.com/microsoft/multi-agent-marketplace (arXiv 2510.25779, MIT).\n\n## What the model under evaluation plays\n\n- **Every LLM marketplace seat**: all customers AND all businesses — one seat\n  per agent, one verifiers interaction per seat. Customers run the\n  benchmark's autonomous-shopping loop (search / message / check / pay /\n  end), businesses respond to inquiries with the benchmark's\n  inquiry-response prompt (text or structured order proposals).\n- Each upstream `generate()` decision call becomes one turn on that seat's\n  interaction; structured-output parse retries become nudge turns on the\n  same interaction (mirroring upstream's append-the-error-and-retry loop).\n- An episode = one full marketplace simulation: the in-process FastAPI\n  marketplace server + SQLite store + all agent loops run on a worker\n  thread; seat turns are bridged onto the episode's event loop with\n  `asyncio.run_coroutine_threadsafe` and served concurrently under the\n  episode's agent gate.\n\n## Scenario cells (benchmark-experiment fidelity)\n\nEach cell = bundled population x marketplace search algorithm (the\nbenchmark's own market-design axis) x bounded customer step budget:\n\n- `mexican3_simple` / `mexican3_lexical` / `mexican3_optimal` — the paper's\n  quickstart population (mexican_3_9: 3 customers, 9 businesses) under\n  rating-ranked / lexical-relevance / fully-fulfilling search. 12 seats.\n- `mexican5_lexical` — first 5 customers / 15 businesses of mexican_10_30,\n  lexical search; the market-size axis (3 -> 5 customers). 20 seats.\n- `contractors5_simple` / `contractors5_lexical` — first 5 customers / 15\n  businesses of contractors_10_30 (home-services domain); cross-domain\n  generalization cells. 20 seats.\n- `mexican3_simple_tight` / `mexican5_lexical_tight` /\n  `contractors5_lexical_tight` — the same three populations with a 6-decision\n  customer budget instead of 10. The standard budget saturates at these market\n  sizes (per-cell means 0.91-1.00), so the tight budget is the discrimination\n  axis; nothing else changes between the pairs.\n\n9 cells x 2 seeds = 18 tasks (the rollouts per task supply the sampling\nvariance; the seed axis labels tasks, since the simulation has no RNG upstream\nof LLM decisions).\n\nUpstream's 40-seat populations are STRUCTURALLY UNRUNNABLE here, not merely\ncontended: one 40-seat episode can consume the org-wide ~32-request model\nceiling by itself (measured: 39% of seats lost to provider 429s), and in\nPrimeAgentHarness mode it also exhausts the team's 2000-tunnel quota. They\nstay bundled but are not a cell.\n\nCell size is deliberately bounded at 12-20 seats. Upstream's full bundled\npopulations (10 customers / 30 businesses = 40 seats) were measured to be\nunviable: each seat's agent run consumes a concurrency slot (the platform's\norg-wide ~32-request ceiling is shared by every running eval) and, in\nPrimeAgentHarness mode, a team tunnel slot — 40-seat cells died with\n`Maximum number of team tunnels (2000) reached` while 12/20-seat cells ran\ncleanly. Medium cells deterministically subset a bundled population (first N\nin file order); the full populations remain bundled and reachable through the\ntaskset knobs. The env's default seat gate is `max_concurrent_agents: 8` for\nthe same reason.\n\nDirect-model seats default to a LOCAL SUBPROCESS runtime\n(`SubprocessConfig`): the null-harness chat program is tool-less, so it runs on\nthe runner VM with no container and no interception tunnel. Prime-runtime seats\nmint one tunnel per rollout and the team's tunnel quota (2000) is shared across\nevery running eval, so staying local removes that failure class. Pin\n`\"agent\": {\"runtime\": {\"type\": \"prime\", \"cpu\": 8, \"memory\": 8}}` for\nPrimeAgentHarness-mode seat runs (those need a container).\n\nCells x seeds = 12 tasks. Rewards recorded per seat trace: `mm` (composite,\nweight 1) + `mm_<component>` (weight 0) with raw quantities in\n`trace.info[\"magentic\"]`.\n\n## Metrics (documented components, all in [0, 1], higher = better)\n\n- `purchase_completion` — fraction of customers who paid for at least one\n  proposal (upstream `purchase_completion_rate`).\n- `needs_met` — fraction of customers whose paid orders matched the full\n  requested menu AND amenities (upstream needs-met flag).\n- `welfare` — total customer utility (upstream utility: 2x requested menu\n  value - payments for needs-met purchases) normalized by the population's\n  theoretical optimum (cheapest fully-matching business per customer).\n- `proposal_validity` — fraction of order proposals without integrity errors\n  (off-menu items, wrong menu prices, inconsistent totals, nonexistent\n  agents; upstream proposal-error classes).\n\nThe composite is the unweighted mean. All raw quantities + per-customer /\nper-business records land in `trace.info[\"magentic\"][\"raw\"]`.\n\n## Vendoring + deviation notes (magentic_marketplace)\n\n- `magentic_marketplace/` is vendored top-level from\n  `packages/magentic-marketplace/src/magentic_marketplace` (MIT,\n  `LICENSE.magentic-marketplace` in this repo).\n- **DB**: upstream runs on PostgreSQL (docker-first). This port runs on the\n  upstream SQLite controller (`platform/database/sqlite/`) — episodes need\n  no docker/postgres and state stays inside the episode sandbox. Analytics\n  run over the live SQLite controller with the upstream\n  `MarketplaceAnalytics` engine.\n- **LLM layer**: upstream dispatches to OpenAI/Anthropic/Gemini structured\n  outputs. This port installs a seat bridge at the vendored\n  `marketplace/llm` entrypoint (`functional.install_seat_bridge`): calls\n  route to the verifiers seat interaction. Structured outputs are emulated\n  on the plain-chat seat by appending the JSON schema to the decision\n  prompt (upstream's prompts rely on provider-side structured outputs, so\n  the schema suffix is an eval-specific addition) and parsing + bounded\n  retry (3 attempts, mirroring upstream).\n- Pruned (unused on the eval path): the FastAPI UI static assets, the\n  postgres-only CLI/experiment tooling (`cli.py`,\n  `experiments/list_experiments.py`, `experiments/export_experiment.py`,\n  `experiments/run_audit.py`). The postgres connector remains but imports\n  lazily (asyncpg not required).\n- **Determinism**: no RNG upstream of LLM decisions (search ranking,\n  proposal validation, analytics are deterministic given actions); the seed\n  axis labels tasks; the stochastic element is model sampling.\n- **Empty-reply guard** (copied from the agent-bazaar-eval template): a\n  bounded self-heal nudge turns empty model replies into a retry, otherwise\n  the episode fails loudly (`SeatPoisonedError`); counters recorded into\n  `trace.info[\"magentic_seat\"]`.\n\n## Running\n\nLocal:\n\n    .venv/bin/vf-eval magentic_marketplace_eval --env-dir-path . \\\n        -m internal/glm-5.3-fast -n 12 -r 1 --max-concurrent 1 \\\n        --max-tokens 131072 --env.agent.runtime.type subprocess \\\n        --no-serve --no-rich\n\nHosted (canary first, then the full 12):\n\n    prime eval run primeintellect/magentic-marketplace-eval --hosted \\\n        -m internal/glm-5.3-fast -n 1 -r 1 --max-concurrent 1 \\\n        --max-tokens 131072 --timeout-minutes 60 \\\n        --eval-name mm-canary --plain\n\nPrimeAgentHarness seat mode (probes): add\n`--env-args '{\"agent\": {\"harness\": {\"id\": \"prime-agent\"}}}'` and for\nseat-heavy cells pin `\"runtime\": {\"cpu\": 8, \"memory\": 8}` +\n`\"max_concurrent_agents\": 4`.\n","encoding":"utf-8","truncated":false,"total_bytes":7521},"status":null}