{"data":{"kind":"file","path":"README.md","version_id":"q61jxu8h6hk151mzyt28vs2s","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":11627,"modified_at":"2026-10-05T19:36:19.382000","content_hash":"1a5dc4ddbd56fc3dcdfbde8f2d60771aa46968726d1017e6986078b2f6c17258"},"entries":[],"content":"# examqm-v1\n\nSingle-turn reconstruction of one curated merged exam-question record from its\nsource page images and, optionally, extractable native text. The record schema\nis defined in [schema.py](examqm_v1/schema.py) and embedded in the model prompt\nfrom that same validator.\n\nThe [repair record](../../docs/examqm-repair.md) describes corrections to the\nexisting 79 rows. The [scorer 3.1 follow-up](../../docs/examqm-scoring-repairs.md)\nrepairs the later checklist findings. The [historical deep audit](../../docs/examqm-deep-audit.md)\nrecords the original defects. Seven benchmark rows still await review; the eval\nsplit remains a narrow challenge set, mostly Fuvest.\n\nThe [source-checked scoring checklist](../../docs/examqm-scoring-checklist.md)\nreplays historical mistakes and equivalent representations through the native\nparser/reward. All 14 stated expectations currently pass. The review page also\nshows the source-checked printed markers used by the grader and checks the\nrepaired poem and Q4 italics against literal source transcriptions.\n\n## Run and inspect\n\nFrom the repository root, `uv sync --extra eval` installs this package and the\nnative v1 runtime from the Git revision pinned in `pyproject.toml`/`uv.lock`.\nThe root application's former verifiers 0.2 dependency is incompatible with\nthis environment. The installed Prime CLI 0.7 uses its legacy eval interface;\nuse the pinned native `eval` entrypoint below.\n\nFirst run the model-free checks:\n\n```bash\nuv run --extra eval python scripts/audit_examqm.py\nuv run --extra eval validate examqm_v1 --runtime.type subprocess --only-gold --no-rich\nuv run --extra eval eval @ configs/eval/examqm-gemini.toml --dry-run --no-serve --no-rich\n```\n\n`validate` checks schema validity and the gold's oracle score. It does not approve\nunreviewed source annotations or establish representativeness. The reproducible\n[coverage/scoring audit](../../docs/examqm-eval-audit.md) documents those limits.\n\nThe live config retains the earlier Gemini model ID and output budget; verify\nthe model is available on your account before a full run. It reads\n`GEMINI_API_KEY`, uses the native `examqm_v1` harness (one reply with no tools), local\nsubprocesses, concurrency 1, bounded startup/provider retries, and no result upload.\nThe harness runs the upstream null program with the already installed, locked\nPython dependencies. It avoids the upstream startup step that downloads uv even\nwhen uv is installed. It supports the local subprocess runtime only.\nRun a small check, inspect it, then omit `--num-tasks` for the full benchmark:\n\n```bash\nuv run --extra eval eval @ configs/eval/examqm-gemini.toml --num-tasks 3 --no-serve --no-rich\nuv run --extra eval python scripts/report_examqm.py outputs/evals/examqm/<run>\n```\n\nThe runner persists the resolved config, full native episodes in `traces.jsonl`,\nand logs. Resume with that exact saved configuration:\n\n```bash\nuv run --extra eval eval @ outputs/evals/examqm/<run>/configs/resolved/eval.json --resume\n```\n\nThe report writes `summary.json`: valid outputs, invalid outputs, truncated\nreplies, and provider/run errors are separate categories. It includes component\nmetrics, literal transcription agreement, review-status groups, and image/table/\nformula/choice/subquestion groups. Input modes remain distinct. Error rates\naccompany scores; a high mean over a few successful replies is not a full-run\nresult. Coverage includes planned/missing rollouts, and the report rejects mixed\ndataset, prompt or scoring versions. The seven non-ready benchmark annotations\nremain visible separately.\n\nThe existing direct-provider probe scripts are exploratory; their caches and\nsummaries are separate from native eval results.\n\n## Colab vision evaluation\n\nInstall `ob1/examqm-v1@0.1.0` and the `colab` extra. The environment's\n`colab_models.json` freezes model revisions, the native verifiers commit and runtime\nversions. The dataset is `oliveirabruno01/vestgraph-examqm-v1`; use a pinned HF\ncommit and verify its manifest. The package also includes the same frozen gold and\nsource images for portable loading without a dataset download.\n\nThe packaged command runs the native taskset, harness, parser and scoring rubric:\n\n```bash\npython -m examqm_v1.colab --model gemini-3.5-flash --root /path/to/results --full-eval\npython -m examqm_v1.colab --model qwen35-4b --root /path/to/results\npython -m examqm_v1.review /path/to/results/qwen35-4b/calibration-eval \\\n  /path/to/results/qwen35-4b/calibration-train --out /path/to/review\n```\n\nGemini uses three rotating key slots. Local Qwen/Gemma models use their real vision\nencoders on a single T4, with frozen quantization settings. Run one model at a time;\nlocal CUDA loading must pass on the real GPU before scaling. Provider/runtime\nfailures stop that model and remain separate from invalid JSON or OCR errors.\nThe bridge preserves raw model replies and native scoring; it never repairs them.\nThe default local comparison uses 15 development questions, including three train\nquestions, so those results are not held-out performance.\n\nPrime's installer has its own dependencies. Keep the CLI in a separate tool\nenvironment (`uv tool run --from prime==0.7.0 prime env install ...`) and run\nevaluation with the pinned native verifiers version.\n\n## Contract\n\nThe model returns one raw JSON object with `schema_version: \"3.0\"` and exactly\none question. Every field is required; extra keys, duplicate JSON keys, aliases,\nmissing values, markdown fences, surrounding prose, invalid types, and multiple\nquestions are rejected. Blocks always contain `kind`, `role`, `text`,\n`table_rows`, and `boxes`; the kind determines which payload is populated.\nThe question number is a locator supplied in the task input, not an output\nfield. Alternative choices and lettered subquestions are ordered lists with no\nprinted IDs; list order carries their sequence. The frozen data row keeps the\nexact labels in a separate `source_metadata` sidecar, aligned by list position.\nThat sidecar is stored with task data and is not sent in the model prompt. The\nprinted IDs remain visible\nin the source page, but the model must not copy them into the content object or\ninclude them in an alternative image crop. Alternatives use `blocks` and\nsubquestions use `body`. All containers use the same fixed role enum, with no\nkind-to-role rule in the schema. All answer-choice blocks use `prompt`, including\ngraphics. Empty table rows, wholly blank tables, and hollow choice/part objects\nare rejected. There are no implicit defaults or parser repairs.\n\nThe task prompt includes the generated JSON Schema and a compact example.\nPreserve source reading order, rows, cells, alternatives and written parts.\nAdjacent text/formula blocks of the same role can be split or joined. Equivalent\npartitions of a visual region are accepted. Preserve punctuation, legend glyphs,\nemphasis and actual printed color. Text values remain strings because\ntranscription is the content being reconstructed; the record structure and\nfield names are fixed by the schema.\n\n## Data and splits\n\nRun `scripts/export_examqm_gold.py` to refresh the frozen JSONL. It includes\nthe union of reviewed `ready` records and benchmark-flagged records, preserving\neach actual review status. A benchmark-flagged `needs_fix` record is therefore\ngold for eval without being relabeled `ready`. The loader rejects stale or\nunversioned rows; `eval` (default) selects every benchmark flag, while `train`\nselects ready rows outside the benchmark. `trusted` reflects `ready`, never\nbenchmark membership. Counts follow current source files.\n\nShared materials, directions and sources are assembled into each standalone\ntarget. Required context pages join the question pages; original question pages\nremain in `source_question_pages`. The exporter freezes 31 deduplicated PNGs in\n`gold/images`, records their SHA-256 hashes, and bundles them in the wheel.\nLoading checks the hashes and does not require the original run outputs.\n\n## Scoring\n\nScorer **3.1** has one reward, `reconstruction_fidelity`: the **lowest required\ncomponent score**, across active body/choice/part sections, visual recovery,\nmath and meaningful symbols. This prevents correct choices from masking an\nabsent passage, or correct prose from masking absent required figures.\n\nText blocks carry weight proportional to their transcription length; missing or\ninvented text reduces matched-token coverage. Choices and written parts stay\nordered. Roles remain exact. Text and formula representations are equivalent\nfor scoring when they preserve the same content and role. Table rows/cells stay\nordered. Prose case, whitespace and the underscore count in one continuous\nanswer blank are tolerated. Standalone runs of at least three underscores encode\none blank; missing or additional blanks lose credit. Math subscripts and Markdown\nemphasis keep their meaning. Accents, punctuation, Greek\nletters, legend glyphs and formatting markup are preserved.\n\nVisual geometry uses page-aware union IoU within each content container; correct\nadjacent partitions are equivalent. Duplicate/overlapping crops and reordered\nboxes lose credit. Geometry is not multiplied into the content score again.\nMissing all required images scores zero.\n\nSource-checked printed choice-marker regions are frozen separately from the\nanswer schema. All 15 graphical choices in the current set have annotations\nverified against the page images; export/loading rejects missing coverage.\nAny content image intersecting one of these markers sets\n`alternative_id_in_image=1` and makes the reconstruction reward zero. Gold crops\nhave no marker overlap. This flag is a violation metric (1 is bad); the report\naverages it only for complete, valid outputs on annotated tasks.\n\nMath allows optional single-atom braces, common equivalent LaTeX commands and\nmatching Unicode glyphs. Changed operation/comparison/grouping signs score zero\nfor that expression. Other math differences lose token credit. No algebraic\nrewriting or symbolic equivalence is inferred. Missing meaningful legend symbols\nscores zero for the symbol component. These conservative rules are deterministic,\nnot a calibrated estimate of downstream question usefulness.\n\nInvalid or truncated output scores zero. Empty extraction scores zero. Supplied\npage numbers earn no OCR reward; their accuracy is a separate diagnostic.\nA perfect reward does not imply literal text equality, pixel-perfect layout,\nor correct page metadata. The report also measures literal transcription agreement.\n\nThe native runner persists `scoring_version` in task data, so changing the scorer\ninvalidates old resume matches. Dataset and prompt fingerprints are also recorded.\nOld scores are not comparable to scorer 3.1 results. The output schema remains 3.0.\n\n## Verification\n\n`tests/test_examqm_proof.py` checks that every ready-or-benchmark source record\npreserves its content, choices, written parts, linked context and required pages,\nand that images retain recoverable boxes. This establishes conversion and input\ncoverage, not a proof that every annotation matches the PDF.\n\n`uv run --extra eval pytest tests/test_examqm_proof.py tests/test_examqm_scoring.py tests/test_examqm_runner.py -q`\nalso checks the actual native CLI against a local oracle endpoint: the outbound\nrequest includes the schema and screenshots, private task metadata stays out\nof the prompt, scores persist, torn writes can resume, and completed tasks are\nnot requested again. Scoring checks cover all 79 oracles, empty outputs,\nmissing required images, equivalent partitions, symbols and math corruptions.\nThese tests make no paid\nmodel calls.\n","encoding":"utf-8","truncated":false,"total_bytes":11627},"status":null}