{"data":{"kind":"file","path":"README.md","version_id":"dp14hgyrg3aghovn81km33f0","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":23925,"modified_at":"2026-09-10T00:45:16.005000","content_hash":"c11a8ab31d3adb5dd526ca7820acc2aa4f65b5b05d5738a77d85a96742523c94"},"entries":[],"content":"**English** · [한국어](README.ko.md) · [中文](README.zh.md)\n\n# nego-eval\n\nA negotiation environment where an agent **chooses** its counterparty and **meets it again**.\n\nEvery negotiation benchmark that lets an agent pick whom to deal with throws the\nrecord away when the session ends, and every one that keeps a record assigns the\ncounterparty. So the thing a relationship is actually made of — what happened\nlast time, with someone you chose — has never been the variable.\n\nHere it is the variable, and it is one bit.\n\n|  | Picks its partner | Memory across games |\n|---|---|---|\n| [Cattle Trade](https://arxiv.org/html/2605.14537v1) | yes | no |\n| [M3-Bench](https://arxiv.org/pdf/2601.08462) | fixed dyad | within game |\n| [RLVR negotiation](https://arxiv.org/abs/2604.09855) | one seller | no |\n| [TERMS-Bench](https://arxiv.org/abs/2605.13909) | assigned | within one negotiation |\n| **nego-eval** | **yes** | **yes** |\n\nTERMS-Bench is worth reading next to the numbers below. It diagnoses thirteen\nfrontier agents inside a single bilateral negotiation and finds that they\n**saturate deal rate while diverging in surplus extraction** — an aggregate\noutcome hiding the differences that matter. The same thing happens here one level\nup: total profit puts four models within noise of each other, and the share of\nthe relational surplus they capture spreads them from 25% to below zero. Its\ncounterpart is assigned and its horizon is one negotiation; the divergence it\nfinds inside a deal is the divergence this board looks for between them.\n\n---\n\n## The game\n\nThree sellers quote openly. The buyer takes one unit a round. Delivery sometimes\nfails, and a failure creates a loss `L` that settles only when two integers add\nup to it:\n\n```\nL_buyer + L_seller = L\n```\n\nThe two sides bargain over that split across up to three exchanges. If they never\nagree, the loss stays where it landed — on the buyer — and the pair stops trading\nfor three rounds. Both are told this before they start, which is what gives the\nseller something to concede against. Every amount is on a grid of ten, so at\n`L = 110` there are exactly twelve legal splits.\n\nA **match** is four games of twelve rounds with the same three sellers. The\nmanipulation is whether each game begins holding the match ledger or among\nstrangers. Seeds, cast, rules and length are identical between the two.\n\n### What is published and what is not\n\n| | Where it lives | Priced? |\n|---|---|---|\n| delivery rate for a buyer with no history | on the quote sheet, exact | yes |\n| share of a loss it absorbs at arm's length | learned by failing with it | no |\n| how much further it goes for a regular | learned by staying | no |\n\nA supplier's on-time rate is in the catalogue and you pay for it. What it does\nwhen a shipment is ruined is not in the catalogue, is not contractible, and is\nfound out by having it happen. All of it is countable in whole units, so the\nreward is arithmetic — **no judge model anywhere in the loop.**\n\n---\n\n## What it is an analogy for\n\nEvery mechanic answers to something a buyer does, and it is worth being explicit\nabout which ones, because a board game that resembles nothing is a puzzle rather\nthan an environment.\n\n| In the game | In procurement |\n|---|---|\n| a published delivery rate, and a higher price for a better one | on-time-delivery figures, quality certifications, SLA tiers — disclosed, and you pay for them |\n| the loss when a delivery fails | the consequential loss: a stopped line, expedited freight, a sale that did not happen, rework |\n| two integers that must sum to the loss | somebody bears it. Contracts are usually silent on consequential damages, and silence puts it on the buyer |\n| three exchanges over the split | the call after the incident — a credit note, a free replacement, splitting the freight |\n| impasse: the buyer bears it all, and the pair stops trading | no agreement, so you eat it, and you stop putting that supplier on the next RFQ |\n| how much of a loss it absorbs at arm's length | not in the catalogue. Found out by having a shipment ruined |\n| how much further it goes for a regular | the unwritten difference between a new account and a ten-year one |\n| the ledger carried across games | a purchase history surviving a fiscal year, a contract renewal, or the buyer changing desks |\n| prices that jitter each round | spot movement. Staying with your supplier costs you the difference some weeks, which is what makes staying a decision |\n\nThe asymmetry in the middle of the table is the point, and it is not invented\nhere. What a supplier discloses is priced into what it charges; what it does once\nsomething has gone wrong is neither disclosed nor contractible, and is learned\nonly by having it happen. That gap is what\n[Brown, Falk & Fehr](https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1468-0262.2004.00511.x)\nput in a laboratory and what Macchiavello and Morjaria measure in a real supply\nchain: relationships form there because nothing else enforces the part of the\nbargain a contract cannot reach.\n\n### For an agent, \"no carry-over\" is not hypothetical\n\nThe manipulation has a second reading that needs no supply chain at all. A fresh\ncontext window is a buyer with no history. So is a new session, a summarised\ntranscript that dropped the ledger, a tool call that returns only the current\nquote. The condition this environment toggles is one that agent systems are in by\ndefault, and the measurement is what it costs.\n\n### Where the analogy breaks\n\nStated because a reader will find these anyway, and should find them here first.\n\n- **Three sellers, not a market.** No entry, no exit, no competition between\n  buyers for a good supplier's attention.\n- **One product, one unit a round.** No baskets, no lead times, no quality\n  grades, no minimum order quantities.\n- **Failure is exogenous and its rate is published.** In reality neither holds:\n  reliability responds to how a supplier is treated, and nobody publishes it\n  honestly.\n- **Solvency is switched off.** Real suppliers squeezed below cost cut corners\n  and then disappear. That mechanism was in an earlier board and is off here for\n  a measurement reason given below; it is a real omission.\n- **No courts, no insurance, no Incoterms.** Consequential loss often *is*\n  allocated in advance in real contracts. This board is the case where it was\n  not.\n- **The record is not a sufficient statistic.** How much a seller concedes\n  depends on the buyer's share of the *last eight* rounds; the record shows\n  lifetime counts. Two histories with identical ledgers can differ by 33 in what\n  a seller will pay on a loss of 110. Every agent measured here faced the same\n  gap, so the comparison holds — but an agent is being asked to build a\n  relationship whose mechanism it can only partly see. One line stating the\n  recent window would close it, at the cost of re-measuring everything.\n- **The horizon is known.** The agent is told how many rounds remain, which\n  sharpens endgame behaviour in a way an open-ended relationship does not.\n\n---\n\n## Quick start\n\n```bash\npip install -e \".[dev]\"\npytest                                   # 34 invariant tests, no API key needed\n\npython scripts/run_carry4.py             # the manipulation, scripted policies\npython scripts/run_dp4.py                # the exact optimum, by dynamic programming\npython scripts/train_rl4_shaped.py       # tabular Q-learning, two reward shapings\n```\n\nLanguage-model buyers need a key; nothing else does.\n\n```bash\ncp .env.example .env && $EDITOR .env\npython scripts/run_probe4.py             # one move per frozen position\npython scripts/run_fullplay4.py \"[('deepseek/deepseek-chat', 12)]\" on\npython scripts/run_fullplay4.py \"[('deepseek/deepseek-chat', 12)]\" off 910012\npython scripts/analyse_paired4.py        # models against references, same seeds\n```\n\nThe third argument is the first seed, so an arm continues where an earlier one\nstopped rather than replaying matches already paid for.\n\n---\n\n## Reference points\n\nAt `L = 110`, four games of twelve rounds, premium 120:\n\n| Policy | Score | Share of the relational surplus |\n|---|---|---|\n| optimum, by dynamic programming | **1099** | 100% |\n| optimum without carry-over | 1031 | 59% |\n| hand-written, reads the record | 996 | 38% |\n| hand-written, quote sheet only | 934 | 0% |\n| tabular Q-learning, shaped reward | 898 | — |\n\nAnd the models, playing whole matches, paired against those references on the\nsame seeds. Zero is the hand-written rule; every model sits to the left of it.\n\n<picture>\n  <source media=\"(prefers-color-scheme: dark)\" srcset=\"docs/models.dark.svg\">\n  <img alt=\"Each model against the hand-written rule, one standard error\" src=\"docs/models.light.svg\">\n</picture>\n\n| Model | n | Score | vs the rule | vs quote sheet only |\n|---|---|---|---|---|\n| gpt-5.6-terra | 12 | 958 | −127 ±68 | +6 ±75 |\n| gpt-5.6-sol | 10 | 897 | −153 ±75 | −27 ±96 |\n| claude-fable-5 | 10 | 876 | −174 ±63 | −48 ±82 |\n| gemini-3.7-flash | 24 | 762 | −118 ±35 | −66 ±34 |\n| qwen3.8-max | 8 | 786 | −251 ±101 | −190 ±105 |\n| claude-haiku-4.5 | 12 | 663 | −421 ±72 | −288 ±79 |\n| deepseek-chat | 12 | 610 | −474 ±87 | −342 ±91 |\n\nThose rows were measured with a system prompt that misstated the rules twice:\nit told models a failed delivery costs them the price, which `World` does not\ncharge, and that sellers do not publish their failure rates, one message before\nprinting all three. Both are corrected in `sim/agents.py`, so a rerun today will\nnot reproduce these numbers exactly. Tested on one model, the contradiction did\nnot change the opening pick — 0.88 with the old text against 0.75 with the new,\npaired, −1.3 sigma — but the misstatement about the price flips which seller\nmaximises expected value on 7 of 48 boards, so the table should be re-measured\nbefore it is leaned on.\n\n**A deficit here does not say what it is made of.** A score is one number and\nthree things move it: how reliably the buyer picks the expected-value best\nseller, how it settles a failure, and what the relationship earns it. Two very\ndifferent accounts fit gemini's −66 equally well.\n\n| | picking | settling | relationship | total |\n|---|---|---|---|---|\n| one story | 0.91 accuracy, −71 | like the null, 0 | 0 | −66 |\n| another | perfect, 0 | takes what is offered, −194 | **+128** | −66 |\n\nUnder the second, the model is *gaining* from the relationship and losing more\nthan that elsewhere. The score cannot tell them apart, and it does not even fix\nthe sign of the relational term. The three figures come from\n`scripts/accuracy_to_profit.py`, `scripts/policy_map.py` and\n`scripts/noise_taxes_relationship.py`, all replayed on the same board.\n\nSo these rows should not be read as measuring how well a model uses the record.\nThey measure a sum. Splitting it needs the pick and the settlement logged every\nround — the model's choice, the expected-value best at that moment, and whether\nthe offer was taken or countered — which leaves the relationship as a residual\nwith a sign. `run_fullplay4.py` keeps only per-seed profit and discards the\nrounds, so it cannot be recovered from `data/`; it has to be re-measured.\n\n**What is forced, though, is that the information was not the limit.** Every\nchoose prompt prints the delivery rate beside each price, and the 934 reference\nconsults nothing else — no record, no memory, three expected values and a\nmaximum. Six of seven models see that same sheet *plus* a ledger and score below\nit. However the split between picking and settling turns out, the shortfall is\nin executing on visible numbers rather than in access to them, so arithmetic\nreliability across the fifty-odd comparisons a match asks for is an axis this\nboard is sensitive to, and one worth reporting apart from any claim about the\nrecord.\n\nFor two rows that much is already pinned. Picking at chance costs about 276\n(`scripts/accuracy_to_profit.py`); haiku-4.5 at −288 and deepseek-chat at −342\nare further down than picking at chance would put them, so the settlement has to\nbe leaking as well. Nothing similar is pinned for the rest.\n\nPairing matters more than it looks. Match-level spread is two to six hundred, and\none early seed was worth +530 to a policy that consults nothing at all — so a\nmodel's mean against a reference mean computed on other seeds would mostly\nmeasure which boards were drawn.\n\nThe surplus denominator is `1099 − 934 = 165`: what reading the record is worth\nwhen read perfectly. Total profit is a poor scale for this environment because\nroughly 85% of it comes from buying at a sensible price, which every competent\npolicy does about as well.\n\n**The optimum is computable.** A seller's concession depends only on the buyer's\nshare of the last eight rounds, delivery rates are published, and prices are a\nfixed sequence per seed — so the state is `(round, last eight picks)` and value\niteration solves it exactly. Very few environments can report a true ceiling\nrather than a best-known score.\n\n---\n\n## The contract\n\nThree boards failed before one worked, and each failed differently. The\nconditions are therefore a module, `nego_eval.game.contract`, run against any\ncandidate board — every one is a measurement, not a guideline.\n\n| Condition | Threshold | Why |\n|---|---|---|\n| answer recoverable from the prompt | ≥ 0.95 | otherwise the task is estimation under noise and no score can be read against a ceiling |\n| quote sheet alone does not settle it | ≤ 0.75 | otherwise the record is decoration |\n| \"fewest failures\" is not a trap | ≥ 0.45 | a board where an ordinary heuristic scores chance measures compliance with its author, not capability |\n| no single heuristic suffices | ≤ 0.80 | |\n| no answer-by-name | ≤ 0.42 | |\n| countering can gain something | = 1.00 | if the opening offer is already the ceiling, there is no negotiation to fail at |\n\nThree earlier boards are **not** in this repository. Each was withdrawn for a\nreason the corresponding condition now encodes, and shipping a package with four\nboards in it — three of them wrong — only invites someone to build on the wrong\none. `table4.py` is the board. What the discarded ones were, and what each got\nwrong, is in `notes/`.\n\n### The condition a checklist cannot hold\n\nBoard three satisfied every measurable condition and still measured the wrong\nthing: making the published rate exact had switched off the channel by which\nloyalty improved delivery, restoring room to bargain left the loyalty term in the\nconcession ceiling non-binding, and between them nothing remained by which\nreturning to a seller made that seller *better*. What was left was a fixed hidden\nattribute and a buyer discovering it. Carrying the ledger then saved time and\nalmost nothing else — which is exactly what it measured.\n\nNo number computed on frozen positions catches that. It took reading a\ntranscript.\n\n---\n\n## What is known so far\n\nMeasured on this board unless noted. See `notes/note.html` for the full\nwrite-up, including five figures that earlier revisions got wrong.\n\n**Carrying the record is worth the gap between how long a relationship takes to\nbuild and how soon it is cut off.** Holding total rounds at 48 and moving only\nwhere the boundaries fall, its value rises monotonically from +61 at four games\nto +181 at twenty-four. With memory the score is flat; without it, each game\nrestarts a relationship that never gets far enough to pay.\n\n<picture>\n  <source media=\"(prefers-color-scheme: dark)\" srcset=\"docs/shape.dark.svg\">\n  <img alt=\"What carrying the ledger is worth as boundaries multiply\" src=\"docs/shape.light.svg\">\n</picture>\n\n**Handing a record to an agent that was not built around one is worse than\nwithholding it.** A tabular policy trained without carry-over, given a carried\nledger at test time, picks the right opening 0.14 of the time against a chance\nrate of 0.33 — a third of its lookups land in states training never visited and\nthe rest carry a median of 178 visits against 55,161. The same shape shows up\ntwice more: appending a pre-computed ratio to the ledger cost one model 0.15 and\ncollapsed its use of the column it had been reading, and six of the seven models\nbelow choose worse once a ledger exists than before there was one.\n\n**Subtracting the non-relational baseline from the reward is worth +49** to an\notherwise identical learner — the control variate removes the 85% of the signal\nthat is already solved.\n\n**Seven language models play whole matches, and six of them choose worse once\nthe ledger arrives than before it.** The opening pick of games 2+ against the\nopening pick of game 1: qwen3.8-max goes 0.88 → 0.29, below the 0.33 of chance,\nhaving spent an hour a match to do it; claude-fable-5 0.70 → 0.47; three others\nfall a little. Only gemini-3.7-flash improves, 0.29 → 0.61, and it is the worst\nof the seven when there is no history to read. The hand-written rule goes\n0.52 → 0.71.\n\n<picture>\n  <source media=\"(prefers-color-scheme: dark)\" srcset=\"docs/ledger.dark.svg\">\n  <img alt=\"Opening pick before and after the ledger exists\" src=\"docs/ledger.light.svg\">\n</picture>\n\n**Every model loses to that rule**, by 118 to 474 paired on the same seeds, at\ntwo standard errors or better. Against the weaker reference — a policy that\nignores the record entirely — only the bottom two are clearly behind, so \"frontier\nmodels score below the null\" is *not* what this shows.\n\n**Scale does not predict any of it.** A flash-tier model is the only one the\nrecord helps; the two top-tier models measured sit mid-table and last. Three\nmodels from one family spanning generations and sizes scored 0.48 / 0.47 / 0.46\non the same probe.\n\n**Whether a model earns anything from carrying the ledger is still open.** One\nmodel was run in both conditions on 35 shared seeds: +40 ±35, against the +64 the\nenvironment offers a policy that uses it perfectly. Consistent with both,\nseparated from neither. At sixteen pairs the same measurement read +88 ±43 and\nlooked settled; it was not.\n\n---\n\n## Training on it\n\n    prime env install hanjinkim/nego-eval@latest    # or: pip install \"nego-eval[rl]\"\n\n`verifiers` is an extra rather than a dependency: it wants Python 3.11 and\nbrings about two dozen packages, and the board, the settlement rules and the\ndynamic-programming optimum need none of them.\n\nThe board is packaged as a `verifiers` environment in `nego_eval/rl/`, and\n`deploy/` holds what it takes to run GRPO against it on one rented H100. Three\nruns happened. None of them improved anything, and the useful part is why.\n\n**A row is a turn, not a match.** The environment hands the model `[system] +\nthe current position` and nothing else. Flattening a match into one training\nsequence would let round 2 attend to round 1's text, which at inference it\ncannot see — training a model to lean on information it will not have is the\nfailure this board exists to measure. So the rollout freezes a prefix and\nbranches `num_generations` ways from one cut point. The prefix is shared inside\na group, so GRPO's group mean removes its contribution exactly, and the\nwhole-match total is the return-to-go up to that constant.\n\n**Run 2 moved nothing, and the shape of the nothing was predicted.** Eight\nevaluation points on fixed seeds fit a trend of −29.8 ± 24.4. The verdict in\n`scripts/verdict.py`, written before any of it ran, failed two of three: `g1`\nfell 0.375 to 0.323 while `gk − g1` rose 0.080 to 0.128 — the opening move with\nno history behind it got worse while the ledger appeared to start helping. That\nis the terminal-reward signature the tabular study recorded, and whole-match\nreward is terminal reward.\n\n**The reason is a ratio.** `scripts/reward_snr.py` freezes a prefix, takes each\nlegal seller, and plays the tail out. Best minus worst, in tail standard\ndeviations: 0.44 on the whole match, 0.85 over eight rounds, 4.75 over two. A\ngroup of eight has a mean standard error of 0.35 SD, so the whole-match reward\noffers about 1.2 sigma per group. Eight rounds is the shortest window that\nstill spans a game boundary, which is where carry-over lives.\n\n**The model was not playing the board.** Thinking was turned off to fit a\n64-token budget. On the opening pick — no history, the answer a pure function\nof the quote sheet — Qwen3-8B scores 0.44 with thinking off against a bar of\n0.53, because naming one seller every time already scores that. With reasoning\nit scores 0.75–0.88, and its spread of picks matches the answer key's rather\nthan a name preference. So the runs were optimising a model with no way to do\nthe board's arithmetic, which is where `scripts/policy_map.py` had placed it\nindependently: on top of random-pick-and-accept-everything.\n\nTurning it back on is not free. One decision with reasoning costs upward of\nfour thousand tokens against a rollout of thirty-five decisions, which puts\nGRPO with thinking on outside a rented-hour budget rather than merely dearer.\nThat trade is not resolved here.\n\n`deploy/README.md` carries all three runs, what changed between them, and the\nprediction each was measured against.\n\n## Layout\n\n```\nnego_eval/\n  sim/       world, bargaining, agents, the learned baseline, the LLM client\n  game/      the board, the denomination grid, the match runner, the contract\n  rl/        the board as a verifiers environment, for training on\nscripts/     every measurement in the write-up, one file each\ndata/        the numbers those scripts produced\ntests/       66 invariants, including the prompt against the rules it describes\nnotes/       the write-up\ndeploy/      running GRPO against it on a rented card\n```\n\nEach module's docstring says what went wrong before it looked like that. That is\ndeliberate: most of the design here is scar tissue, and the scars are the part\nworth reading.\n\n## What would move this forward\n\nTwo questions are open, and neither is open for want of effort.\n\n**Splitting a score into its channels.** The model table reports a sum — picking\narithmetic, settlement, and whatever the relationship earned — and one number\ncannot separate them, so those rows cannot be read as measuring how well a model\nuses the record. Separating them needs the pick and the settlement logged every\nround: the model's choice, the expected-value best seller at that moment, and\nwhether the offer was taken or countered. That is a re-run of the seven models,\nabout 850 calls each for twelve matches, and the cost is dominated by reasoning\ntokens rather than by call count. `run_fullplay4.py` keeps only per-seed profit,\nso it cannot be recovered from `data/`.\n\n**Training something that can play it.** The board is packaged for RLVR in\n`nego_eval/rl/` and three GRPO runs are recorded in `deploy/`. None\nimproved anything, for a reason no reward shape reaches: the models with the\npick accuracy — all above 0.9 implied — are the ones whose weights are closed,\nand the ones that can be fine-tuned score at or below what naming one seller\nevery time would score. Reasoning is what buys the arithmetic, and leaving it on\ncosts about seventy times the tokens.\n\nIf you have inference credits, access to open weights that clear roughly 0.80\npick accuracy on `scripts/thinking_probe.py`, or you want to use the board for\nsomething else, open an issue. The commit log has an address if that is easier.\n\n---\n\n## License\n\n[Business Source License 1.1](LICENSE) — source-available, not OSI open source.\n\n**Evaluating any model or agent with it is free**, including your own models,\nincluding internally at a company, including publishing what you find. So is\nresearch and teaching. What the licence withholds is using it as a training\nenvironment in a pipeline that ships commercial weights, or reselling it as a\nproduct.\n\nIt converts to Apache 2.0 on 2030-09-01, or four years after any given version\nis published, whichever comes first. For anything the grant does not cover,\ncontact the licensor.\n","encoding":"utf-8","truncated":false,"total_bytes":23925},"status":null}