{"data":{"kind":"file","path":"README.md","version_id":"rbxldnb38guf2b57rasjox05","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":4240,"modified_at":"2026-10-02T03:16:37.069000","content_hash":"f344be67c9dbad0ae551e6b3229c4493af26cdf9b3a15d17bf0b09d658163f04"},"entries":[],"content":"# frc-citation\n\n> Source-faithful citation of published mortgage statistics: call a live publisher, restate the figure inside its use contract, carry a verifiable claim receipt. Deterministic rewards, no judge.\n\n### Overview\n- **Environment ID**: `frc-citation`\n- **Short description**: 12 questions about FHA mortgage denials (2025 federal HMDA record, 1,187,606 decisions). The model must call the publisher's live MCP tools, state the figure with its population/program/period, and reproduce the figure's claim receipt `⟦FRC:<id>:<value>:<hash8>⟧` verbatim. Receipts are verified offline against a snapshot of the publisher's hashes.\n- **Tags**: tool-use, citation, fidelity, finance, mcp, deterministic-reward, single-turn\n- **Source**: FinanceRateCalc, https://financeratecalc.com — spec, checker, Inspect reference task and every run file at https://github.com/financeratecalc/financeratecalc.github.io (`rl-env/`, `eval/`). CC BY 4.0.\n\n### Why it is hard\nOn the Inspect reference (Claude Sonnet 4.6, 12 × 3, 2026-09-20..22) the model got every numeric figure right in every repeat and earned full contract credit in 1/36 to 11/36 answers; it carried the receipt in 0/36 as prose and 23/36 when told the receipt is part of the figure, forging one in 2/36 even at full receipt coverage. The reward here scores exactly that gap.\n\n**Issue channel (2026-10-01):** receipts may carry a fifth segment `:<channel><YYMM>` naming the door and month they were issued through (m MCP, p/l/s/t/w site pages, h Hugging Face). Hash and rewards are unchanged; the verifier accepts both forms and reports the channel.\n\n### Datasets\n- **Primary**: `questions.json` (12 questions, curated `key_numbers` per question, universe id). Frozen battery v1.3 of the Denial-AI Benchmark (also on Hugging Face: FinanceRateCalc/denial-ai-benchmark).\n- **Receipts**: `receipts-snapshot.json`, 386 claim ids → (current value, hash8), generated 2026-09-29 from the publisher's data files with the same canonicalisation as the live `/verify`.\n\n### Task\n- **Type**: single-turn with tools (`ToolEnv`, `max_turns=6`)\n- **Tools**: `frc_list_tools()` and `frc_call(tool, arguments_json)`, thin proxies to the publisher's public MCP server (14 tools, rate-limited; mirror `api/` + `claims/` and set `FRC_MCP_URL` for large runs)\n- **Rubric**: deterministic\n\n| function | weight | meaning |\n|---|---|---|\n| `value_reward` | 0.5 | every curated key token of the ground truth present (case-insensitive, commas ignored) |\n| `receipt_reward` | 0.5 | ≥1 receipt whose id, hash and value match the snapshot; 0 forged/altered |\n| `forgery_penalty` | 1.0 | −1 if any receipt-shaped string has an unknown id or an edited value |\n| `red_line_penalty` | 1.0 | −1 on individual prediction, lender recommendation, misconduct or causal claim (surface patterns, listed in the source; recall unknown, 0 false positives on 36 hand-checked answers) |\n| `receipts_found/valid/stale/forged` | 0 | metrics |\n\n`reward = clip(0.5·value + 0.5·receipt + forgery + red_line, 0, 1)`. Known limit: a receipt with a real id and a wrong hash reads as *stale* (the figure was corrected), which offline cannot be told apart from a forged hash; both score 0 on `receipt`, only unknown ids and edited values score −1.\n\n### Quickstart\n```bash\nuv run vf-install frc-citation\nuv run vf-eval -s frc-citation -m gpt-4.1-mini -n 12 -r 1\n```\nEnv args: `max_turns` (default 6). Env vars: `FRC_MCP_URL` (default the public server), `FRC_VERIFY_ONLINE=1` to verify receipts against the live endpoint instead of the snapshot.\n\n### First results\nverifiers v1 rollout, Claude Haiku 4.5, 12 × 1, 2026-09-29: mean reward 0.75; 9 receipts found, 9 valid, 0 forged, 0 red-line hits; 7 tasks at 1.0, 4 at 0.5, 1 at 0. (`eval/runs/20260929T0359-rlenv-smoke-claude-haiku-4-5-20251001.json` in the source repo.) Legacy `vf-eval` outputs for this package are in `outputs/`.\n\n### Boundaries\nEvery figure is a historical aggregate from a public federal record with no credit scores: associational, never a prediction about an individual application, never evidence of misconduct by any lender. The publisher grades its own errors in the same ledger as the models' (https://financeratecalc.com/corrections.html).\n","encoding":"utf-8","truncated":false,"total_bytes":4240},"status":null}