{"data":{"kind":"file","path":"README.md","version_id":"ou7hmnecsy0o2yihbvlrt5d7","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":1079,"modified_at":"2026-09-17T17:52:29.615000","content_hash":"b3b631351558dc4c41396232a9f80bcd4eaf1495bdbff756953d79489886d4a0"},"entries":[],"content":"# AutomationBench-Verified\n\nA benchmark for evaluating AI agents on realistic business workflows.\n\n- **White Paper:** https://arxiv.org/abs/2604.18934\n- **GitHub:** https://github.com/zapier/AutomationBench\n- **Artificial Analysis:** https://artificialanalysis.ai/evaluations/automationbench-aa\n\nLearn more at [zapier.com/benchmarks](https://zapier.com/benchmarks) or run it on the [Prime Intellect Environments Hub](https://app.primeintellect.ai/dashboard/environments/zapier/AutomationBench).\n\n## Changes compared to AutomationBench\n\n- The tools are exposed as native tools to the models with their full schema, not as stringified JSON or prose\n  - This heavily relies on top-level `allOf`, `anyOf` and `oneOf`. Some open model tool parsers do not parse those values correctly. [vLLM PR](https://github.com/vllm-project/vllm/pull/53729), [SGLang PR](https://github.com/sgl-project/sglang/pull/36626)\n- Fixes some of the APIs to be called; some of the loaded tools for some of the tasks as well as the graders.\n- Builds on top of verifiers v1, removes all the legacy (v0) code.\n","encoding":"utf-8","truncated":false,"total_bytes":1079},"status":null}