{"data":{"kind":"file","path":"README.md","version_id":"g3efvc4c81afr796bswmgcto","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":3940,"modified_at":"2026-02-12T04:39:02.359000","content_hash":"7b9b957133142bf629629e1d5914071f8d06b3eeb6a6f08f1c7b3f3623b2f8b9"},"entries":[],"content":"# python-context-env\n\nPython tool-use environment based on `vf.PythonEnv` with key features:\n\n1. The Python sandbox is pre-seeded before turn 1 with:\n   - `task` = task description string\n   - `data` = full task payload object\n   - `tool_call_log` = string log of prior python tool calls (`Turn <n>` + code)\n   - `context_list` = list with one pre-populated item (task_summary, schema, data_keys, cursor)\n2. Each turn prompt is rebuilt to show only:\n   - one memory-reset instruction (static prefix for KV cache optimization),\n   - `context_list` items rendered as `[i]: <json>`,\n   - turn progress and reminders.\n3. The model is instructed to keep static reference info at the front of `context_list` and dynamic state at the back, maximizing KV cache prefix hits across turns.\n4. A rollout log file is written with per-turn:\n   - assistant tool call(s) from the previous turn,\n   - the newly rendered user prompt for the current turn.\n5. Rollout termination + scoring:\n   - if any item in `context_list` contains `\"final_answer\"`, rollout stops early,\n   - reward combines judge model evaluation and a prefix match metric (gated by judge score).\n\n### Overview\n- **Environment ID**: `python_context_env`\n- **Type**: multi-turn tool-use\n- **Base**: `vf.PythonEnv`\n\n### Quickstart\nFirst generate a dataset (writes to the default path):\n```bash\npython /Users/eligottlieb/Documents/research-environments/environments/python_context_env/synthetic_task_generator.py\n```\n\nThen run eval:\n```bash\nuv run vf-eval python_context_env\n```\n\nExample with custom truncation limit:\n```bash\nuv run vf-eval python_context_env -a '{\"max_context_chars\": 2000, \"max_turns\": 10}'\n```\n\nGenerate synthetic tasks to the default dataset path:\n```bash\npython /Users/eligottlieb/Documents/research-environments/environments/python_context_env/synthetic_task_generator.py\n```\n\nGenerate tasks to a custom JSONL path:\n```bash\npython /Users/eligottlieb/Documents/research-environments/environments/python_context_env/synthetic_task_generator.py \\\n  --output /tmp/context_tasks.jsonl \\\n  --num-tasks 100 \\\n  --families all \\\n  --difficulty 4 \\\n  --seed 123\n```\n\nRun env against a generated JSONL file:\n```bash\nuv run vf-eval python_context_env \\\n  -a '{\"dataset_path\":\"/tmp/context_tasks.jsonl\",\"max_turns\":25}'\n```\n\nSpecify a custom rollout log directory (or disable file logging with `null`):\n```bash\nuv run vf-eval python_context_env \\\n  -a '{\"turn_log_dir\":\"/tmp/python_context_turn_logs\"}'\n```\n\nSupported synthetic families:\n- `streaming_clue`\n- `entity_state`\n- `checkpoint_reasoning`\n- `error_recovery`\n- `compression_pressure`\n- `min_max_tracking`\n- `graph_degree`\n- `fsm_execution`\n- `unique_counting`\n- `conditional_accumulation`\n\nSynthetic JSONL row format (minimum):\n- `task`: short instruction string (used for REPL `task`)\n- `data`: structured payload object (used for REPL `data`)\n- `answer`: target output string (typically minified JSON)\n- optional: `info` (may contain `task_summary` — injected into `context_list[0]` so the model sees operation definitions and a worked example every turn)\n\n`question` is optional; if missing, the loader falls back to `task`.\n\n### Environment Arguments\n| Arg | Type | Default | Description |\n| --- | ---- | ------- | ----------- |\n| `dataset_path` | str \\| Path \\| None | `environments/python_context_env/my_data/train.jsonl` | JSONL dataset path. If `None`, uses the default path. |\n| `max_turns` | int | `20` | Maximum interaction turns |\n| `max_context_chars` | int | `4000` | Character cap for rendered context_list items in each prompt |\n| `turn_instruction` | str | `DEFAULT_TURN_INSTRUCTION` | Instruction shown each turn |\n| `turn_log_dir` | str \\| Path \\| None | `None` | Directory for per-rollout turn logs (`None` disables file logging) |\n| `pip_install_packages` | str | `\"numpy sympy scipy\"` | Packages installed in sandbox startup |\n| `max_startup_wait_seconds` | int | `30` | Worker readiness timeout |\n","encoding":"utf-8","truncated":false,"total_bytes":3940},"status":null}