{"data":{"kind":"file","path":"README.md","version_id":"asbqjoc1orv82q67vwmomhl7","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":16837,"modified_at":"2026-10-02T06:15:58.490000","content_hash":"d6a99265f64cdd0ee31a052a11db78da9518c8d329ec6917e93206ae2edd2835"},"entries":[],"content":"# SMDD-Bench on Prime\n\n[Paper](https://arxiv.org/abs/2605.21740) | [SMDD-Bench website](https://smddbench.com/) | [Blog](https://r-suresh07.github.io/writing/is-human-taste-overrated-in-harness-engineering/)\n\nSMDD-Bench tests whether language-model agents can carry out long-horizon small-molecule drug design. This package is the Prime-native release: it bundles the task data, original agent harness, scoring code, and Docker build sources needed to run SMDD-Bench with Prime's verifiers.v1 evaluation runner.\n\nThere are five task families:\n\n| Type | Task | Submission |\n| --- | --- | --- |\n| 1 | 2D Pharmacophore Identification | `solution.py` |\n| 2 | Interaction Point Discovery | `solution.csv` |\n| 3 | Scaffold Hopping | `solution.smi` |\n| 4 | Lead Optimization | `solution.smi` |\n| 5 | Fragment Assembly | `solution.smi` |\n\nThe full benchmark contains 502 tasks. Some validation builds contain only a subset; run smdd-bench list to see exactly what is installed in your release.\n\n## How the setup fits together\n\nThe deployment has two moving parts:\n\nThe evaluation host runs Prime, calls the language model, and launches the task and verifier containers.\n\nThe oracle service runs Boltz and ADMET on GPU. It can live on the same machine or on a separate GPU machine / Ray cluster.\n\nPrime handles the short-lived task and verifier containers for you. The oracle is the only long-running service you start yourself.\n\nThe task and verifier are intentionally separated. Public task files are staged into the agent runtime; only declared submission artifacts are copied into the verifier runtime, where the private grader inputs are staged. The operator still has access to the release files, but the agent does not receive the grader, oracle token, or host model credentials.\n\n## Before you start\n\nUse a Linux host with:\n\n- Docker\n- Python 3.12\n- `uv`\n\nEvery GPU host also needs the NVIDIA driver and NVIDIA Container Toolkit. Before building anything, confirm that nvidia-smi works and that Docker can access the GPU.\n\nFor Boltz, we recommend at least 48 GB of VRAM per GPU to reduce out-of-memory failures. Actual memory use depends on the protein and inference settings, so 48 GB is not a guarantee for every task. For larger workloads, an 8 x A100 80 GB machine is one option.\n\nMultiple GPUs increase throughput by serving independent requests. Their memory is not automatically pooled into one larger device for a single prediction. Start with one trial, watch GPU and host memory, and only then increase concurrency.\n\nYou will also need enough disk space for Docker images and model checkpoints. The oracle requires outbound network access to its MSA service. The evaluation host needs network access to both your model provider and the oracle endpoint.\n\n## 1. Install the Prime environment\n\nRun the following in Bash on the Linux evaluation host.\n\nThe package is published as looni-lab/smdd-bench. Version 0.1.4.post0 is currently private, so log in with a Prime account that has access.\n\n```bash\nuv venv --python 3.12 \"$HOME/.venvs/smdd-bench\"\nsource \"$HOME/.venvs/smdd-bench/bin/activate\"\ncommand -v prime >/dev/null || uv tool install prime\nprime login\nprime env install looni-lab/smdd-bench@0.1.4.post0\n```\n\nPrime may install the downloaded wheel into its own tool Python even when a virtual environment is active. Check the install output. To make sure the release is available inside the evaluation environment, install the downloaded wheel explicitly:\n\n```bash\nWHEEL=\"$HOME/.prime/wheel_cache/looni-lab/smdd-bench/0.1.4.post0/dist/smdd_bench-0.1.4.post0-py3-none-any.whl\"\nuv pip install --python \"$VIRTUAL_ENV/bin/python\" \"$WHEEL\"\n```\n\nIf Prime printed a different wheel path, use that path instead.\n\nNow check the installed task set and create a writable run directory:\n\n```bash\nsmdd-bench list\nsmdd-bench setup \"$HOME/smdd-bench-run\"\ncd \"$HOME/smdd-bench-run\"\n```\n\nNo repository checkout is required; the package already includes its task data and runtime sources.\n\nWhen migrating from an older SMDD distribution, use the new virtual environment above and generate a fresh run directory with `smdd-bench setup`; old TOML configs still reference the previous loader IDs.\n\nKeep this Python environment activated while running SMDD. Avoid installing another SMDD distribution into the same environment, since different releases may contain incompatible versions of the shared harness module.\n\nIf smdd-bench: command not found appears, first check that the virtual environment is still active and that the explicit wheel installation succeeded.\n\nSome Prime CLI versions print a generic from verifiers import load_environment example. That uses the older API. This release exports native classes, so use the smdd-bench eval commands in this guide instead.\n\nFor these VM-based evaluations, you do not need Prime Hub Secrets or Variables. Model and oracle credentials are configured in the VM shell below; Hub settings are for hosted services and do not configure this local runner.\n\n### What was installed\n\nThe Python package contains the canonical task definitions, harness, runtime sources, and task data:\n\n```text\nsmdd_bench/\n|-- taskset.py              # Canonical tasks and native Prime loading\n|-- harness.py              # Reference runs and model-backed agent runs\n|-- worker.py, relay.py     # Host agent loop and isolated task operations\n|-- release.json            # Task IDs, image tags and source hashes\n|-- assets/smdd-runtime.zip # Docker and oracle build sources\n+-- tasks/<task-id>/\n    |-- task.toml           # Resources, images and separate verifier policy\n    |-- instruction.md      # Public task instructions\n    |-- environment/        # Public files placed in the agent workspace\n    |-- tests/              # Grader and private scoring inputs\n    +-- solution/           # Reference solution, when available\n```\n\nsmdd-bench setup creates a separate writable run directory:\n\n```text\nsmdd-bench-run/\n|-- release.json\n|-- lead.toml               # One real lead-optimization trial\n|-- lead-smoke.toml         # Lead reference solution and verifier\n|-- smoke.toml              # One reference task from each family\n+-- smdd-runtime/           # Extracted sources, including build_images.sh\n```\n\n## 2. Build the Docker images\n\nThe release ships Docker build sources rather than prebuilding everything on your machine. build_images.sh uses the exact image tags recorded in release.json.\n\n| Image role | Build on | Used for |\n| --- | --- | --- |\n| Agent science and evaluator images | Evaluation host | Parent images for task and verifier containers |\n| Task science and task lead images | Evaluation host | Agent workspaces |\n| Five verifier images | Evaluation host | Separate scoring environments for the five task families |\n| Oracle model image | GPU build host | Boltz, ADMET, and Ray dependencies |\n| HTTP oracle image | Every GPU VM | Long-running oracle service |\n\nChoose the build that matches your deployment:\n\n```bash\n# Evaluation and oracle on this same GPU host:\nbash smdd-runtime/build_images.sh all\n```\n\n```bash\n# Alternatively, evaluation host only, or reusing an existing GPU oracle:\nbash smdd-runtime/build_images.sh tasks\n```\n\n```bash\n# Separate GPU host, after extracting the same package's runtime sources:\nbash smdd-runtime/build_images.sh oracle\n```\n\nPrime launches task and verifier containers automatically. You start the oracle service yourself.\n\nRay workers only need the HTTP oracle image.\n\n## 3. Start the oracle\n\nThe oracle exposes the same HTTP API whether it runs on one GPU host or across several GPU VMs with Ray.\n\nUse private or VPN addresses that are mutually reachable. In particular, the oracle must be reachable from both the evaluation host and the verifier containers. 127.0.0.1 inside a verifier points back to that verifier container, not to the evaluation host.\n\nOn the GPU host — or on the Ray head — start from the extracted run directory and initialize the persistent state and cache:\n\n```bash\nexport ORACLE_IMAGE=\"$(python -c 'import json; print(json.load(open(\"release.json\"))[\"images\"][\"smdd-oracle-http:latest\"])')\"\nexport STATE=\"$HOME/smdd-oracle/state\"\nexport CACHE=\"$HOME/smdd-oracle/cache\"\ninstall -d -m 700 \"$STATE\" \"$CACHE\"\ntest -s \"$STATE/token\" || openssl rand -hex 32 > \"$STATE/token\"\nchmod 600 \"$STATE/token\"\n```\n\n### Option A: one GPU host, no Ray\n\n```bash\ndocker run -d --name smdd-oracle-http --restart unless-stopped \\\n  --gpus all --shm-size=16g -p 8090:8090 \\\n  -e ORACLE_ROLE=standalone \\\n  -v \"$STATE:/srv/oracle_state\" -v \"$CACHE:/srv/oracle_cache\" \\\n  -v \"$STATE/token:/run/secrets/oracle_token:ro\" \\\n  \"$ORACLE_IMAGE\" --max-samples 10 --max-parallel-samples 2 --microbatch-size 1\n```\n\n```bash\ndocker logs --tail 100 -f smdd-oracle-http\n```\n\nThis works with one or several GPUs on the same host. docker logs -f only follows the logs: pressing Ctrl+C closes the log viewer and leaves the oracle running.\n\n### Option B: two or more GPU VMs with Ray\n\nAll nodes need the same oracle image and bidirectional private/VPN connectivity.\n\nRay uses auxiliary ports in addition to 6379, 6380, 6381, and 20000-20063. Keep these ports on a trusted network. Public NAT addresses are not local bind addresses.\n\nCopy the built oracle image to each worker, replacing the SSH address:\n\n```bash\ndocker save \"$ORACLE_IMAGE\" | ssh ubuntu@WORKER_SSH_ADDRESS docker load\n```\n\nOn the head, use its real private/VPN address and set RAY_EXPECTED_GPUS to the total number of GPUs across the cluster, including the head:\n\n```bash\nexport HEAD_IP=\"10.77.0.1\"\ndocker run -d --name smdd-head --restart unless-stopped \\\n  --gpus all --network host --shm-size=16g \\\n  -e ORACLE_ROLE=head -e RAY_NODE_IP=\"$HEAD_IP\" \\\n  -e RAY_EXPECTED_GPUS=2 -e ORACLE_HTTP_HOST=0.0.0.0 \\\n  -v \"$STATE:/srv/oracle_state\" -v \"$CACHE:/srv/oracle_cache\" \\\n  -v \"$STATE/token:/run/secrets/oracle_token:ro\" \\\n  \"$ORACLE_IMAGE\" --max-samples 10 --max-parallel-samples 2 --microbatch-size 1\n```\n\nOn each worker, use the same image tag as the head and replace the example address with that worker's own private/VPN address:\n\n```bash\nexport ORACLE_IMAGE=\"smdd-prime/oracle-http:0.1.4\"\nexport HEAD_IP=\"10.77.0.1\"\nexport WORKER_IP=\"10.77.0.2\"\ndocker run -d --name smdd-worker --restart unless-stopped \\\n  --gpus all --network host --shm-size=16g \\\n  -e ORACLE_ROLE=worker -e RAY_NODE_IP=\"$WORKER_IP\" \\\n  -e RAY_ADDRESS=\"$HEAD_IP:6379\" \"$ORACLE_IMAGE\"\n```\n\nStart the workers before waiting for HTTP readiness. The head waits until the expected GPU actors have initialized.\n\nWorkers do not need the oracle bearer token or a shared filesystem. If a node does not appear, inspect Ray and the container logs:\n\n```bash\ndocker exec smdd-head /opt/oracle/bin/ray status\n```\n\n## 4. Connect Prime to the oracle\n\nThere are two pieces of configuration:\n\n- the URL of the running HTTP oracle;\n- the local token file containing its authentication token.\n\nThese are deployment settings, not fixed defaults in the package.\n\nIf the oracle is running on the evaluation VM, inspect its network binding and token mount. For a Ray deployment, use smdd-head in place of smdd-oracle-http.\n\n```bash\ndocker ps --format 'table {{.Names}}\\t{{.Status}}\\t{{.Ports}}'\nexport ORACLE_CONTAINER=smdd-oracle-http\ndocker inspect \"$ORACLE_CONTAINER\" --format 'Network={{.HostConfig.NetworkMode}} Ports={{json .NetworkSettings.Ports}}'\nexport SMDD_AGENT_ORACLE_HTTP_TOKEN_FILE=\"$(docker inspect \"$ORACLE_CONTAINER\" --format '{{range .Mounts}}{{if eq .Destination \"/run/secrets/oracle_token\"}}{{.Source}}{{end}}{{end}}')\"\nprintf 'Local token file: %s\\n' \"$SMDD_AGENT_ORACLE_HTTP_TOKEN_FILE\"\n```\n\nThe command above prints the path to the token file, not the token itself. If it prints an empty path, check the container name and mounts before continuing.\n\nIf the oracle lives on another VM, copy the token securely to the evaluation host and set SMDD_AGENT_ORACLE_HTTP_TOKEN_FILE to that local absolute path.\n\nChoose the oracle URL from the way the service is actually bound:\n\n| Oracle setup | URL to use |\n| --- | --- |\n| Same-VM mapping `172.17.0.1:8090->8090/tcp` | `http://172.17.0.1:8090` |\n| Mapping `0.0.0.0:8090->8090/tcp` | The oracle VM's private/VPN IP on port 8090, reachable from the evaluation host and verifier containers |\n| Ray head with host networking | The head's reachable private/VPN IP on port 8090; no Docker port mapping is expected |\n\nUse the address and port reported by your deployment. 172.17.0.1 is the Docker bridge address in the tested same-VM setup; it is not a universal oracle address and should not be used from another VM.\n\nFor the same-VM bridge setup above, configure the agent and verifier clients like this. Change only the URL if your deployment differs:\n\n```bash\nexport SMDD_AGENT_ORACLE_HTTP_URL=\"http://172.17.0.1:8090\"\nexport SMDD_AGENT_ORACLE_CACHE_DIR=\"$HOME/.cache/smdd-prime\"\nexport SMDD_EVALUATOR_ORACLE_HTTP_URL=\"$SMDD_AGENT_ORACLE_HTTP_URL\"\nexport SMDD_EVALUATOR_ORACLE_HTTP_TOKEN=\"$(cat \"$SMDD_AGENT_ORACLE_HTTP_TOKEN_FILE\")\"\ncurl --fail --silent --show-error \\\n  -H \"Authorization: Bearer $SMDD_EVALUATOR_ORACLE_HTTP_TOKEN\" \\\n  \"$SMDD_AGENT_ORACLE_HTTP_URL/v1/health\" | python -m json.tool\n```\n\nInitial model loading can take several minutes. Wait until the health endpoint reports healthy model status before starting evaluations.\n\nKeep --max-samples 10 on the oracle. SMDD verifiers request ten diffusion samples regardless of how many GPUs are serving requests.\n\n## 5. Smoke-test the installation\n\nBefore paying for a model rollout, run the supplied reference solutions through the real verifier.\n\nThe oracle harness mode does not call a language model. It stages the reference answer and scores it with the same verifier used for a real run. Science tasks still use the oracle.\n\nStart with Lead Optimization, then test one task from every family:\n\n```bash\nsmdd-bench eval @ lead-smoke.toml\nsmdd-bench eval @ smoke.toml\n```\n\nA healthy installation should finish without runtime exceptions and report:\n\n```text\nverifier_error = 0\n```\n\nDo not treat reward = 0 by itself as an infrastructure failure. The scientific reward is task-dependent: a valid molecule can be scored successfully and still fail one of the benchmark gates. If a smoke task returns zero, inspect the individual verifier metrics first.\n\n## 6. Run a real model trial\n\nThe generated lead.toml runs one real Lead Optimization task. For example, to use Claude Sonnet 4.6 through Anthropic:\n\n```bash\nread -rsp 'Anthropic API key: ' ANTHROPIC_API_KEY\nprintf '\\n'\nexport ANTHROPIC_API_KEY\nsmdd-bench eval @ lead.toml \\\n  --model claude-sonnet-4-6 \\\n  --client.base-url https://api.anthropic.com/v1 \\\n  --client.api-key-var ANTHROPIC_API_KEY\n```\n\nPrime routes model calls through its interception endpoint so the full trajectory can be captured.\n\nThe original SMDD harness exposes Python, Boltz, ADMET, and submission tools. Its default per-trial budgets are:\n\n- 100 turns\n- 8 Boltz calls\n- 15 ADMET calls\n\nChange these under [env.agent.harness] if you want a different budget.\n\n### Run another task or a larger set\n\nCopy lead.toml and change [env.taskset].tasks to the canonical IDs printed by:\n\n```bash\nsmdd-bench list\n```\n\nSet the top-level num_tasks to the number of task IDs in that list.\n\nTwo settings control test-time scaling:\n\n- `num_rollouts`: independent repeats per task;\n- `max_concurrent`: simultaneous trials.\n\nStart with max_concurrent = 1. Once one end-to-end trial is stable, increase concurrency while watching GPU memory, host memory, and oracle throughput.\n\n## 7. Follow a run and inspect the results\n\nPrime writes its native traces and results under:\n\n```text\noutputs/\n```\n\nThe SMDD harness also keeps a per-trial directory under:\n\n```text\nsmdd-prime-logs/\n```\n\nA completed real trial contains the trajectory, a readable transcript, and — once the agent has terminated — agent_result.json.\n\nThe native Prime trace points back to this directory through info.smdd_logs.\n\nTo see active trajectories:\n\n```bash\nfind smdd-prime-logs -name trajectory.jsonl\n```\n\nThen follow one of the returned paths:\n\n```bash\ntail -f smdd-prime-logs/TASK-AND-RUN-ID/trajectory.jsonl\n```\n\nUse a real path returned by find rather than the placeholder above.\n\nA quiet terminal, or a container that is still running, does not necessarily mean model calls are progressing. Check trajectory timestamps and tool responses, then inspect the final verifier metrics.\n\nCompleted workers also preserve:\n\n```text\nworker.stdout\nworker.stderr\n```\n\nThe example configs use push = false, so evaluation results are not uploaded automatically.\n\n## Deployment boundary\n\nThis guide covers Linux + Docker evaluation with an operator-managed GPU oracle.\n\nHosted sandbox execution is a different deployment mode: task/verifier images must be accessible to the hosted runtime, and the oracle must be reachable from that infrastructure.\n","encoding":"utf-8","truncated":false,"total_bytes":16837},"status":null}