{"data":{"kind":"file","path":"README.md","version_id":"pw6fnkviz1gu4u6vqtyltw9c","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":6444,"modified_at":"2026-08-07T06:30:03.599000","content_hash":"ebf48c9e9dedbf1b720f68ab5ece235c69d0472bc9e1de8ccc6bbe1577999f2b"},"entries":[],"content":"# pytorch-env\r\n\r\n### Overview\r\n- **Environment ID**: `pytorch-env`\r\n- **Short description**: RL environment for PyTorch tensor/autograd/nn tasks using expected_output comparison\r\n- **Tags**: pytorch, deep-learning, autograd, nn\r\n\r\n### Datasets\r\n- **Primary dataset(s)**: `eltociear/pytorch-tasks-v1`\r\n- **Source links**: [HuggingFace Dataset](https://huggingface.co/datasets/eltociear/pytorch-tasks-v1)\r\n- **Split sizes**: train (45 examples)\r\n\r\n### Task Categories\r\n\r\n| Category | Tasks | Description |\r\n|----------|-------|-------------|\r\n| tensor_ops | 16 | reshape, transpose, reductions, clamp, masking, sort/argsort, cat/stack, matmul, normalise, cumsum, dtype cast |\r\n| nn_modules | 12 | Linear and Conv2d with explicit weights, relu/gelu/sigmoid, softmax/log_softmax, LayerNorm, pooling, flatten, Embedding |\r\n| autograd | 6 | backward on scalar losses, `torch.autograd.grad`, `detach` semantics, second derivatives via `create_graph` |\r\n| losses | 6 | MSE, cross-entropy (mean and per-sample), BCE, L1, smooth L1 |\r\n| optim | 5 | one SGD step, SGD+momentum, one Adam step, three chained steps, `clip_grad_norm_` |\r\n\r\n### Task\r\n- **Type**: Multi-turn tool use (default `max_turns=5`)\r\n- **Rubric overview**: Binary pass/fail using `torch.allclose` (rtol 1e-5, atol 1e-7) against the reference tensor\r\n\r\n### Quickstart\r\n\r\n```bash\r\nuv run vf-eval pytorch-env -p prime -m openai/gpt-5.4-nano -s\r\n```\r\n\r\nFull sweep:\r\n\r\n```bash\r\nuv run vf-eval pytorch-env -p prime -m openai/gpt-5.4-mini -n 45 -r 3 -s\r\n```\r\n\r\n### Environment Arguments\r\n\r\n| Arg | Type | Default | Description |\r\n|-----|------|---------|-------------|\r\n| `split` | str | `\"train\"` | Dataset split to use |\r\n| `dataset_name` | str | `\"eltociear/pytorch-tasks-v1\"` | HuggingFace dataset name |\r\n| `max_turns` | int | `5` | Maximum interaction turns per task |\r\n\r\n### Tools Available\r\n- `execute_code(code: str)` — run Python in the sandbox; the task's input tensors are preloaded by name and `result` persists across turns\r\n- `bash(command: str)` — run shell commands in the sandbox\r\n\r\nThe sandbox installs the **CPU** torch wheel on purpose: every task is CPU-deterministic and\r\nthe CUDA build is a multi-gigabyte download it does not need.\r\n\r\n### Grading\r\n\r\nThe reference answer never enters the sandbox — it is held host-side and compared after the\r\nrollout, so the model cannot read the answer key.\r\n\r\nComparison is `torch.allclose` with a tolerance, **not** exact equality (the sibling\r\n`pandas_env` uses exact). matmul, convolution and pooling dispatch to BLAS/oneDNN kernels whose\r\nlast bits legitimately differ across builds and CPU architectures; bit-exact grading would fail\r\ncorrect solutions for reasons unrelated to the model.\r\n\r\nThe tolerance cannot launder a wrong answer: **shape is compared exactly**, and an\r\ninteger-valued reference **requires** an integer answer.\r\n\r\nA wrong answer scores `0.0` silently. A broken scorer scores `0.0` and additionally sets\r\n`state[\"scoring_error\"]`, so eval runs can tell \"the model was wrong\" from \"the harness failed\".\r\n\r\n### Building the dataset\r\n\r\n```bash\r\npython build_tasks.py --verify          # check every task\r\npython build_tasks.py --out train.jsonl # regenerate\r\n```\r\n\r\nTasks are (deterministic input tensors, instruction, reference solution); expected outputs are\r\n**computed** from the reference, never hand-written. `--verify` independently checks each task\r\nfor clean execution, real-valued tensor output, determinism across two runs, finiteness (no\r\nNaN/inf in the answer key), non-emptiness, non-identity, and an exact serialisation round-trip\r\nincluding dtype and shape. All 45 pass on torch 2.13.0+cpu.\r\n\r\n### Why the nn layers use explicit weights\r\n\r\n`Linear`, `Conv2d` and `Embedding` are filled with `arange`/`linspace`/constant values rather\r\nthan `manual_seed` plus default initialisation. Default initialisers and RNG stream details are\r\nimplementation choices that have changed between PyTorch releases; `arange` cannot. A dataset\r\nwhose labels silently change under `pip install -U torch` would be worse than no dataset.\r\n\r\n## Environment arguments\r\n\r\n`load_environment()` deliberately exposes very little: the program's guidance is that an\r\nenvironment should have one correct way to be run, so nothing about the prompts, the parsing or\r\nthe grading is configurable.\r\n\r\n| Argument | Default | Meaning |\r\n|---|---|---|\r\n| `split` | `'train'` | Dataset split to load. |\r\n| `dataset_name` | `'eltociear/pytorch-tasks-v1'` | Hugging Face dataset of tasks. Change only to point at a fork. |\r\n| `max_turns` | `5` | Tool-use turns the model gets before the rollout ends. |\r\n| `**kwargs` | — | Passed through to the underlying `SandboxEnv`. |\r\n\r\n## Reward rubric\r\n\r\n| Reward function | Weight | What it returns |\r\n|---|---|---|\r\n| `correctness` | 1.0 | Binary: 1.0 when the answer matches the reference, else 0.0. |\r\n\r\nThere is no LLM judge and no partial credit. The score is computed on the host in\r\n`post_rollout` and read back by the rubric, so the reward is a deterministic function of the\r\nvalues the model left in `result`. A **harness** failure (dead sandbox, unreadable read-back)\r\nalso scores 0.0 but additionally sets `state[\"scoring_error\"]`, so an eval run can tell a\r\nbroken harness from a wrong answer instead of blaming the model.\r\n\r\n## Dependencies\r\n\r\n`datasets>=4.1.0`, `torch>=2.0.0`, `verifiers>=0.1.8`\r\n\r\n## Sample `vf-eval` usage\r\n\r\n```bash\r\nuv run vf-install pytorch-env\r\nuv run vf-eval -s pytorch-env -m gpt-4.1 -n 5 -r 3\r\nuv run vf-tui                      # inspect the outputs/ folder it writes\r\n```\r\n\r\n### Known limitation\r\n\r\nThe Docker **sandbox transport** has not been executed — `docker run`, `pip install`, and the\r\nmodel's code running inside the container — because that needs a Docker runtime and an\r\ninference provider key, neither available where this was authored.\r\n\r\nEverything else is exercised. `environments/test_scoring_path.py` runs this environment's\r\n`post_rollout` and `Rubric` with the sandbox mocked out, asserting that a correct answer scores\r\n1.0, a well-formed wrong answer scores 0.0, and a dead sandbox scores 0.0 *and* sets\r\n`scoring_error` so a harness failure is never mistaken for a bad model. It also checks that\r\nwhat `build_tasks.py` emits is exactly what the scorer expects. The environment mirrors the\r\nstructure of `polars_env` (already accepted into the Environments Program) and imports cleanly\r\nagainst `verifiers` 0.2.1.\r\n","encoding":"utf-8","truncated":false,"total_bytes":6444},"status":null}