{"data":{"kind":"file","path":"README.md","version_id":"ocmgx4hpob0xm324hwflm42z","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":674,"modified_at":"2026-09-18T06:33:27.342000","content_hash":"18f3d2144b3e13792ed33daa140be7fbdf5922842d8bc94a51b8acf1c7a681fc"},"entries":[],"content":"# OfficeQA Pro V2\n\nA re-implementation of [OfficeQA Pro V2](https://github.com/databricks/officeqa) using Verifiers v1.\n\n## Changes compared to upstream\n\n- Basic blocklist of domains where the dataset is hosted to avoid simple lookups (Hugging Face, GitHub, and common mirrors).\n- Changed the regex-based grader to an LLM-as-judge to catch semantic nuances better. We found this to be less brittle. The judge is **GPT-5.6 Luna @ medium**.\n\nSet `env.taskset.network_block = []` to allow all domains for an evaluation. Omitting\nthis setting preserves the default dataset-host blocklist. For hosted evaluations,\npass it through `--env-args '{\"taskset\":{\"network_block\":[]}}'`.\n","encoding":"utf-8","truncated":false,"total_bytes":674},"status":null}