{"data":{"kind":"file","path":"README.md","version_id":"quvoy95ngekqvve0e4zxiksk","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":9296,"modified_at":"2026-10-06T06:36:28","content_hash":"fd6194688d2545a540514ee0bddd2e8b0580a7f06b35790c04b76ec460ce9c7f"},"entries":[],"content":"# hf-tasksmith-v1\n\nTwelve coding tasks taken from merged pull requests to Hugging Face's machine-learning\nlibraries. In each one the agent gets the repository as it was just before the pull request,\nplus a description of the change, and has to make it. A separate grader then runs the pull\nrequest's own tests against the agent's files.\n\nThe tasks come from the Hugging Face dataset\n[FineEnvs/HF_ML_Tasksmith](https://huggingface.co/datasets/FineEnvs/HF_ML_Tasksmith) at\nrevision `3c63c8b059d734fe74f932101f44dd0b16e7ee26`. Regents Labs packaged them for\n[Techtree](https://techtree.sh), where they make up the Tasksmith Climb.\n\n| | |\n|---|---|\n| **Taskset** | Harbor tasks, each with its own agent image and its own grader image |\n| **Grading** | Separate: the agent's box is torn down, the files each task lists are copied into a fresh box from the grader image, and the tests run there. The agent never sees the tests |\n| **Reward** | 1.0 when every listed test holds, 0.0 otherwise |\n| **Network** | None, for the agent and for the grader |\n| **Needs** | Docker (or the Prime runtime) and the `images` setting below |\n\n## The tasks\n\n| Task | Pull request | Tests that must start holding | Tests that must keep holding |\n|---|---|---|---|\n| `tasksmith-0d2d1e298e86` | [huggingface/peft#3350](https://github.com/huggingface/peft/pull/3350) | 4 | 1 |\n| `tasksmith-35c1a487454c` | [huggingface/peft#3212](https://github.com/huggingface/peft/pull/3212) | 2 | 14 |\n| `tasksmith-7bc616ce49d7` | [huggingface/diffusers#13921](https://github.com/huggingface/diffusers/pull/13921) | 38 | 2 |\n| `tasksmith-9f3fb07e5766` | [huggingface/trl#6206](https://github.com/huggingface/trl/pull/6206) | 12 | 11 |\n| `tasksmith-c488fc138ba1` | [huggingface/peft#3098](https://github.com/huggingface/peft/pull/3098) | 12 | 1 |\n| `tasksmith-ecb1931929df` | [huggingface/peft#2939](https://github.com/huggingface/peft/pull/2939) | 20 | 1 |\n| `tasksmith-0df9b3cd6191` | [huggingface/trl#6001](https://github.com/huggingface/trl/pull/6001) | 3 | 7 |\n| `tasksmith-0ea2241cceeb` | [huggingface/trl#5501](https://github.com/huggingface/trl/pull/5501) | 7 | 5 |\n| `tasksmith-1c5704b1f07f` | [huggingface/diffusers#12703](https://github.com/huggingface/diffusers/pull/12703) | 12 | 21 |\n| `tasksmith-5dd11b34cce1` | [huggingface/transformers#35348](https://github.com/huggingface/transformers/pull/35348) | 12 | 1 |\n| `tasksmith-9aea936e30e4` | [huggingface/peft#2952](https://github.com/huggingface/peft/pull/2952) | 14 | 1 |\n| `tasksmith-c98f3a4a299d` | [huggingface/accelerate#3529](https://github.com/huggingface/accelerate/pull/3529) | 12 | 1 |\n\nTechtree's Climb runs the first six every round and keeps the last six for its held-out check.\nEach task folder holds `instruction.md` (what the agent reads), `task.toml` (limits, the files\nhanded to the grader, and the pull request and commits it came from), `tests/` and\n`solution/` (the reference answer). `hf_tasksmith_v1/provenance.json` records the one line\nchanged in each published `task.toml` and the digest of every packaged folder.\n\n## Images\n\nEvery task runs in two images, pinned by digest and public on GitHub's container registry. The\npackage carries none of them: you name them in `[env.taskset.images]`, one entry per task to\nrun, and only the tasks named there are loaded. To run fewer tasks, leave entries out. These\nare the images Techtree uses:\n\n```toml\n[env.taskset]\nid = \"hf-tasksmith-v1\"\n\n[env.taskset.images.tasksmith-0d2d1e298e86]\nagent = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:b6425ee9d5512122d1fcdc4f465849fb81c52709a48d3b266c269d2b68e59b49\"\ngrader = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:e76502673d91bf2261f3e9f4892f23e9e22902c89632e8c6e31300d84c8827aa\"\n\n[env.taskset.images.tasksmith-0df9b3cd6191]\nagent = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:694c4525e372bdd314ca36028977058eae28ba63965db66e78feddb405900139\"\ngrader = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:ed1340461c0e5e24fa6e5708615c135ced64aeedc8308c4b51d4733ce1be967e\"\n\n[env.taskset.images.tasksmith-0ea2241cceeb]\nagent = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:0ef23048795c6a35d155470cfd453c10fe3c5386856798966bf4bddb860f7317\"\ngrader = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:4ba112600db8712280d2eedd57b81eadb2b676310151af9cd4a4d52aa77c68f5\"\n\n[env.taskset.images.tasksmith-1c5704b1f07f]\nagent = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:3a93126842c8d9a529c4bb47e674f2b9509296fc76f46db1cde266052996c19e\"\ngrader = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:c59e680ddfb5b9c5a32d1926678cf39009422afb4e685ad4e4cb8380ea2ebd4b\"\n\n[env.taskset.images.tasksmith-35c1a487454c]\nagent = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:5dea2dad3530951b996c36b6d9083fcfe3538b5202ad8f70b03842a36b9b1a96\"\ngrader = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:43820bc79efb0e182fea2f0dea84433e3ee35790d201c8c8d02cc3aa75f0bebc\"\n\n[env.taskset.images.tasksmith-5dd11b34cce1]\nagent = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:5f285375a3a307cbc19468b63f7c9facef0872b84bdea3cb7064fd7f5d20f33b\"\ngrader = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:cae1da55c43030f41df9bee874bc1e2be3781f0fa4b646b06f5f9ce0ea002adb\"\n\n[env.taskset.images.tasksmith-7bc616ce49d7]\nagent = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:9151a207fe78b63c3e67150dbee499af9a1dd1e98ffaf70b930d86fb3f1ddec9\"\ngrader = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:4a62fdba3463a277856fbe036975c2a559fb0e6200c8b65f161f74174f1c0968\"\n\n[env.taskset.images.tasksmith-9aea936e30e4]\nagent = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:85740c29c5fa66c6fec405f136722beeeef629eaf89a86dd1d2c5f70a00e15c5\"\ngrader = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:fc815cf3b02722bfd1cdd27636a8a48197acd7ee8dbf1ac3ea7fdd3aa82bc13d\"\n\n[env.taskset.images.tasksmith-9f3fb07e5766]\nagent = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:4f3c9b94179ae550f7a6c60297bdd7db5d76f2bacff9cd0fb7c307a513cf67e4\"\ngrader = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:b4de5c2e0c1b2a83b8c0a9927b53f6d586bf5ea4fa006d84487de2bd0985699c\"\n\n[env.taskset.images.tasksmith-c488fc138ba1]\nagent = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:b499356ba08485bd456839d032399ededf770368c571c1001b5656b78f0ef5ad\"\ngrader = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:f2c81997be59d589d615fba778a101eca79d8c814cba80326d3b3040fb8386f6\"\n\n[env.taskset.images.tasksmith-c98f3a4a299d]\nagent = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:a09990596df5a270ac4a4bc54535b4087613020f8ce15742b1c8cad96859aa22\"\ngrader = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:535678fe51a24c55553eea2a95075b3431e45118c895c0e9a0a8bc8ee2636eb9\"\n\n[env.taskset.images.tasksmith-ecb1931929df]\nagent = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:626658b8e65a1f8ea61a69183180a14def3a24527d05718c2bc1e20b99c8415b\"\ngrader = \"ghcr.io/regents-ai/techtree-tasksmith@sha256:45afad838d6da0fcfb6a7206b1fac0d82cdf77de42fc99b74648e2caa7180f5a\"\n\n[env.agent.runtime]\ntype = \"docker\"\n```\n\nSave that as `tasksmith.toml`. The agent images are built from each task's `environment/`\nfolder in the dataset and the grader images from its `tests/Dockerfile`; both hold the\nrepository at the pull request's head commit with the pull request's source changes taken back\nout.\n\n## Run\n\n`vf-validate` reads the same settings without the `env.` and `env.agent.` prefixes:\n\n```bash\nsed -e 's/^\\[env\\.taskset/[taskset/' -e 's/^\\[env\\.agent\\.runtime\\]/[runtime]/' \\\n  tasksmith.toml > validate.toml\n\n# Model-free: for each task, grade the untouched repository (must score 0), then the\n# reference answer (must score 1), each in a fresh grader box\nuv run vf-validate @ validate.toml\n\n# Check the configuration without starting anything\nuv run vf-eval @ tasksmith.toml --model <model-id> --dry-run\n\n# Evaluate\nuv run vf-eval @ tasksmith.toml --model <model-id>\n```\n\nWhile loading, Verifiers warns for each task that `[verifier.environment]` names no\n`docker_image` and that it will grade in the agent's image. The taskset then gives every task\nits grader image from `images`, and grading starts a fresh box from that image; the warning is\nabout the published `task.toml`, not about what runs.\n\nWithout other flags the agent is Verifiers' `bash` harness; choose another with\n`--env.agent.harness.id`. Each task's own time limits (600 seconds for the agent, 150 for the\ngrader) are ignored unless you pass `--no-env.taskset.ignore-timeouts`, as for every Harbor\ntaskset. `tasksmith-5dd11b34cce1` hands the grader all of `src/transformers`, so\n`artifact_max_bytes` defaults to 128 MiB here instead of Verifiers' 32 MiB.\n\n## Status\n\n| Check | Result |\n|---|---|\n| `vf-validate`, Docker runtime, Verifiers commit `fc73e02` | 12 of 12 valid: every task scores 0 untouched and 1 with its reference answer |\n\n## Dependencies\n\n`verifiers[harbor]==0.3.2.dev147` (Verifiers commit `fc73e02`). No API keys or environment\nvariables of its own.\n\n## Licences\n\nThis package's own code is MIT (`LICENSES/MIT.txt`). The task folders are republished from the\ndataset under whatever terms their authors set: the dataset's statement\n(`hf_tasksmith_v1/LICENSES.md`) says it does not relicense its sources, and Regents Labs grants\nno licence to them either. The reference solutions are copies of Hugging Face source files\nunder the Apache License 2.0 (`LICENSES/Apache-2.0.txt`). `LICENSES/NOTICE.md` has the details.\n","encoding":"utf-8","truncated":false,"total_bytes":9296},"status":null}