{"data":{"kind":"file","path":"README.md","version_id":"herq6msw4iuud004dxbg3qny","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":84354,"modified_at":"2026-09-11T12:54:15.765000","content_hash":"b3fb5a844cb3191f98ac55ed7380a6d09bb74e29c55ce3423690a4d18da1f6dd"},"entries":[],"content":"# DemoCart / Amazon Cart 001\n\n## Weighted stage result (2026-09-10)\n\nEach completed search stage contributes 0.5 and each completed cart stage contributes 0.8. The reported result is `(0.5 * passed_search_stages + 0.8 * passed_cart_stages) / 3`. It is stored as `stage_result` plus numeric `final_result`. All six stages therefore produce 1.3; this is intentionally not normalized to 1.0. The terminal verifier reward remains separate (PASS=1, FAIL=-1, INVALID=0), and historical artifacts are unchanged.\n\n## 2026-09-10 — Startup deadline repair before attempt 3\n\nThe first unscored 0.5.4 Prime check built the app and booted Android but started no task or model call. Its original manifest proves System UI restarted (PID 898 to 2524) and a clean observation arrived at 171.597 seconds. The 180-second deadline could not fit the required fifteen-second clean interval. Diagnostics were exported and sandbox tr56omtlvo1dw3hh5dpza8pa was confirmed terminated.\n\nThis startup failure does not consume another model-attempt slot: two completed failures remain counted. The check's observed cost was USD 0.0618, cumulative USD 0.8918. The unpublished 0.5.4 candidate now allows 240 seconds for startup UI and 1200 seconds for transport exchange. VM/episode lifetimes, task, prompt, model, action limit and grading stay unchanged. A new bounded live check must pass before attempt 3.\n\n| File / receipt | Exact update | Validation / boundary |\n|---|---|---|\n| harness/emulator_config.py | Software startup UI allowance 180 → 240 seconds; exchange 1140 → 1200 seconds. | Retain changed-PID and fifteen-second clean-UI requirements. No extra recovery tap. VM stays 30 minutes; episode stays 1680 seconds. |\n| tests/test_startup_budget_regression.py | Virtual-clock regression reproduces late recovery: old deadline fails, new deadline permits the full clean interval; permanently absent UI still fails. | Two new tests passed. This is a synthetic timing test, not an evaluation. |\n| Scoped harness054-prime-check/summary.json; session/.../evidence.zip | Preserve original startup XML, logs, diagnostic PNG, cost and cleanup receipts. | Zero task actions and zero model calls. |\n| Scoped next-check and launch helpers | Require fresh source/installed tests and a successful live check with matching emulator settings. | One-shot creation, prior spending and original USD 9 cap stay enforced. |\n\n## 2026-09-10 — Input-readback safeguard before requested eval 4\n\nThe next harness revision, 0.5.4, waits for the app's existing read-only SQLite probe to confirm the complete requested text before dismissing the keyboard, and checks it again afterward. It does not retype, append missing characters, submit search, or change the task, app UI, model prompt, sampling or scoring rules. Every probe retains its original response and ADB hash binding.\n\nThe source suite passed 242 offline tests, including 12 new input/audit regressions and the actual Java/D8 check. A real Prime harness check is still required before the next model run. Only two completed attempts are verified; the user's requested label “eval 4” must not manufacture a third completed result or reset the twelve-attempt series.\n\n| File / component | Exact change | Verification / boundary |\n|---|---|---|\n| amazon_cart_001/harness/peach.py | Add two bounded read-back stages through the fixed read-only provider URI; up to five reads / twenty seconds per stage. Preserve the original input command sequence. | Truncated/unreadable input raises TextInputUnconfirmed; no unknown input is replayed. |\n| amazon_cart_001/harness/peach_episode.py | Save input_readback/<step>-<index>.txt and trace-linked metadata; keep known-dispatched input separate from acceptance. Preserve pipeline-origin errors and stop the episode. | Malformed actions cannot reuse prior input receipts. Existing INVALID handling applies; no task success is fabricated. |\n| amazon_cart_001/verification/strict.py | Validate the fixed query URI, exact trace position, original response hash/length, episode, requested/actual text, order and final confirmation of both probe groups. | Only those bound read-only entries are removed before applying the unchanged exact input-command audit. Extra/mutating commands remain rejected. |\n| tests/test_text_readback.py | Twelve tests cover complete/delayed/truncated/empty input, probe errors, non-text actions, malformed output, forged probes and controlled pipeline-invalid outcomes. | Fake-device regressions only, not scored benchmark attempts. |\n| pyproject.toml; amazon_cart_001/__init__.py; specs/campaign_12.json | Record 0.5.4 and verified completed_attempts=2 with the next slot 3; preserve the requested run label separately pending clarification. | No new evaluation or publication has occurred at this checkpoint. |\n\nThe text-loss observation is confirmed by attempt 2's original ADB/UI/state evidence; keyboard-dismissal timing as its root cause remains a hypothesis. The live check must establish actual delivery before launch, not merely that these new unit tests pass.\n\n\n## Attempt 2 of 12 — harness repair and continuation\n\nAttempt 2 completed on Prime as [ct9asc4nna0d9zq1y48kgs4a](https://app.primeintellect.ai/dashboard/evaluations/ct9asc4nna0d9zq1y48kgs4a): **FAIL (-1), 18 GPT-4.1 actions, 5/6 workflow stages**. The original first FAIL is retained. The [series ledger](artifacts/prime_hosted_20260910/campaign-20260910-gpt41-attempt02-timeout60/series-ledger.json) now records **2 of 12 completed, 10 remaining**; no third attempt was launched. Version 0.5.3 was loaded by Prime; published version ID cvvl1jfyhpiqa9460ltdpp4m and source SHA acb5faac69da86e8c255db5909c9578e4daf52190b2b1f415cdafdfb0b895d50 are unchanged.\n\nThe first attempt-2 submission was rejected before creation with HTTP 422 because timeout_minutes=40 is below Prime's 60-minute minimum. Its launch intent, complete error and payload are preserved in campaign-20260910-gpt41-attempt02/. A separately recorded timeout-only correction uses 60 minutes for the managed controller; Android's independent 30-minute lifetime and the eighteen-request limit are unchanged. The accepted receipt directory is campaign-20260910-gpt41-attempt02-timeout60/. No duplicate model attempt was created or discarded.\n\nThe user explicitly retained the first completed Prime result as attempt 1 (FAIL, -1) and approved repairing the harness before attempt 2. Version 0.5.3 separates the bottom Cart tab (`cart_button`) from the detail-page Go to cart control (`detail_cart_button`). Both remain subject to visible/enabled/exactly-one targeting and both require observed cart state for acceptance. The Java/XML app, task goal, system prompt, required separate searches, reward formula and success criteria are unchanged.\n\nThis launch used exactly one hosted GPT-4.1 rollout, numbered 2 of twelve: completed_attempts=1, attempts_this_run=1. Previous spending remained charged against the original USD 9 approval; each model request and VM was reserved before execution. The abandoned pre-model startup remains in the infrastructure/cost history, not counted as a completed model rollout. No original grade was rescored or discarded, and no model/task/scoring change was made during this attempt. The repaired implementation has a new source identity; these two attempts must not be presented as an unchanged-harness pass@3 sample.\n\n| File / component | Exact change | Validation / boundary |\n|---|---|---|\n| harness/peach.py | Define two cart IDs, allow them in the peach profile, and map each UI description to its own ID. | The exact-target guard is unchanged; truly duplicated, disabled and invisible controls remain rejected. |\n| verification/peach.py | Recognize either declared cart ID as a navigation operation, still requiring page == cart. | This repairs an action binding, not the task-success definition or reward weights. |\n| tests/fixtures/peach_detail_cart_controls.xml; tests/test_cart_selector_regression.py | Focused fixture reconstructed from the two native nodes in the saved hosted-1 observation; test unique resolution, dispatch consistency, forbidden controls and accepted/rejected cart transitions. | The fixture is explicitly labeled reconstructed, not an original XML artifact or a live rerun. |\n| integrations/campaign_budget.py; capped_env.py | Add validated completed-attempt offset, per-run attempt limit, prior-spend reservation and campaign attempt number in result/model-request receipts. | One new attempt has at most eighteen requests; loading remains side-effect-free. No change to GPT-4.1, temperature 0.6, token limits or zero retries. |\n| tests/test_campaign_continuation.py | Test attempt 2 numbering, no nineteenth call, no thirteenth slot, prior-cost enforcement and safe loading. | 230 offline tests passed with zero skips, including Java/D8; these are not evaluation results. |\n| pyproject.toml; __init__.py; specs/campaign_12.json | Version 0.5.3 and explicit continuation metadata. | Harness and acceptance-source hashes change. Keep per-version records; an overall twelve-attempt count is not proof of an unchanged-harness statistical sample. |\n\n### Completed attempt-2 evidence\n\n| Item | Observed result | Evidence / limit |\n|---|---|---|\n| Cart-selector repair | Action 9 `cart_button` executed and was accepted; the model reached the cart. | Original action receipts in [worker observations](artifacts/prime_hosted_20260910/campaign-20260910-gpt41-attempt02-timeout60/worker-observations.jsonl); final cart is visibly present in [frame 018](artifacts/prime_hosted_20260910/campaign-20260910-gpt41-attempt02-timeout60/hosted-1/episodes/20260910T045435Z_2441e501/frames/018/screen.png). This run exercised the bottom tab, not the separate detail-page CTA. |\n| Why the task failed | Headphones were added directly without their required separate search. Bottle and backpack searches succeeded; actions 11–18 repeated rejected finish. Action 10 invented the unsupported `review_demo_cart_button` ID. | [Native verdict](artifacts/prime_hosted_20260910/campaign-20260910-gpt41-attempt02-timeout60/hosted-1/hosted_verdict.json) retains FAIL for action_audit, cart_lineage and exact_outcome. Visible cart summary passes; workflow progress is 5/6, not a partial episode reward. |\n| Additional input defect | At action 3, the full requested `Trail Steel Water Bottle` was dispatched with ADB exit code 0, but trusted app state and UI both held only `Trail Steel Water `. | [Read-only audit](artifacts/prime_hosted_20260910/campaign-20260910-gpt41-attempt02-timeout60/diagnostics/context-snapshot-20260910T050107Z-audit.json). The receipt currently attributes this to the agent; that attribution is not established. Text loss is confirmed; its low-level cause is unresolved. No live patch or rescore. |\n| Screenshots / policy reports | 19 original PNGs and 19 checkpoints (000–018), including initial and every after-action state, downloaded and hash-matched. | [Download manifest](artifacts/prime_hosted_20260910/campaign-20260910-gpt41-attempt02-timeout60/hosted-1/download-manifest.json). Checkpoints retain per-policy statuses, verifier reasons and separate progress. |\n| Recording | Native recording receipt has four segments, no capture errors and a 929,405-byte final MP4; Prime stores MP4 and ZIP artifact references. | [Recording receipt](artifacts/prime_hosted_20260910/campaign-20260910-gpt41-attempt02-timeout60/hosted-1/episodes/20260910T045435Z_2441e501/video/recording.json). [Supplement manifest](artifacts/prime_hosted_20260910/campaign-20260910-gpt41-attempt02-timeout60/supplement-manifest.json) verifies only three locally backed-up segments. Final MP4/ZIP local retrieval and dashboard playback remain unverified; partial backups are not called the complete recording. |\n| Live UI / cleanup | Public unauthenticated read-only PNG was verified during this model run. Viewer terminated at 05:06:51 UTC; Android worker terminated at 05:07:02 UTC. | [Final audit](artifacts/prime_hosted_20260910/campaign-20260910-gpt41-attempt02-timeout60/final-audit.json): zero active account sandboxes. Managed controller inspection returns 403, so direct controller-deletion verification is not claimed. Old live URL is no longer usable. |\n| Cost / counter | Observed attempt-2 wallet decrease USD 0.3819; cumulative series decrease USD 0.8228. Two completed FAIL results, ten slots remain. | Includes prior startup costs; wallet deltas are not final itemized invoices. No third attempt, reward tuning or successful pass@3 claim. |\n\n## Previous native hosted evaluation — attempt 1\n\n[GPT-4.1 evaluation dashboard](https://app.primeintellect.ai/dashboard/evaluations/mdgvdtqkzoum9s51dols4b58). Prime confirmed CANCELLED at 03:59:26 UTC on 2026-09-10. One of twelve planned attempts completed: FAIL, reward -1, eighteen model actions, and 3/6 workflow stages. The agent added all three products without the required separate searches. It then repeatedly attempted cart_button, but both the detail-page Go to cart button and bottom Cart tab mapped to that same ID, so the harness rejected the ambiguous target. This is a confirmed app/harness selector defect in harness/peach.py, not evidence that Android or GPT-4.1 never ran.\n\nMonitoring stopped the remaining paid attempts after inspecting the first completed sample. The following worker was cancelled during Android startup, before a second model trajectory; it did not consume the completed attempt-2 slot. Remaining attempts in that original submission did not start. The original failure remains unchanged. The twelve-attempt campaign is incomplete, so its primary pass@3 is not reportable. No rubric, model, published source, or task action was changed during monitoring; no replacement evaluation was launched.\n\n| Evidence / control | Verified result | Boundary / next action |\n|---|---|---|\n| Original screenshots and checkpoints | [Downloaded evidence](artifacts/prime_hosted_20260910/campaign-20260910-gpt41-12-serial/hosted-1/download-manifest.json): 19/19 PNG hashes match the checkpoint receipts; 19 checkpoint JSON files and the transcript are saved. | Includes the initial frame plus all eighteen actions; per-policy results remain preserved. |\n| Video and complete ZIP | Prime's sample includes artifact references and SHA-256 digests for the original recording and ZIP, and reports evidence_export_verified=true. | The samples API replaces embedded ZIP/MP4 bytes with bucket/key references. Local video/ZIP retrieval and dashboard playback were not verified in this monitoring turn. Do not claim the local evidence directory contains an MP4. |\n| Resource cleanup | [Final read-only audit](artifacts/prime_hosted_20260910/campaign-20260910-gpt41-12-serial/final-audit.json): both exact Android VMs TERMINATED; viewer tunnel terminated; zero active account sandboxes. | Managed controller inspection returns 403, so its deletion is not independently asserted. The evaluation itself is CANCELLED. |\n| Spending | Wallet observed at 47.1755 before the approved submissions and 46.7346 at 04:09:46 UTC: USD 0.4409 difference, below the USD 9 cap. | Balance observation, not final itemized billing; includes the earlier scheduling submission. |\n| Required fix | Separate the two cart-control identities and regression-test a real product-detail UI containing both controls. | Keep exact-target validation and the strict search/cart rubric. Validate offline before publishing or requesting another paid run. |\n\nThe earlier submission bem8qa7h3bm7yfjqsq7j08ec was cancelled after a pre-model grouping error. The corrected submission above used independent_scoring=true so max_concurrent=1 applies per rollout. All launch/error/monitoring receipts are retained under artifacts/prime_hosted_20260910/. Historical preview and launch sections below describe their recorded times, not current run readiness.\n\n## Approved hosted campaign — 2026-09-10\n\nThe user approved a USD 9 maximum and requested immediate launch of twelve GPT-4.1 attempts on Prime. Version 0.5.2 uses the already-tested Ubuntu software-emulator provisioning recipe in a fresh Prime VM per attempt. Both the model controller and Android worker run on Prime; no model trajectory is executed on the local server.\n\nThe strict task/rubric remain unchanged. The requested pass@3 band is a target, not an outcome guarantee. The campaign retains all attempt records and does not retry failures, discard invalid attempts or change rewards to reach the band. Submission and live URLs must be read from current run receipts, not the expired preview below.\n\n| File / control | Exact change and purpose | Limit / evidence |\n|---|---|---|\n| integrations/runtime_bootstrap.py; pyproject.toml | Package canonical Java/XML app source and provisioning assets in the wheel, scan a bounded source bundle, then provision Ubuntu 24.04 in the selected Prime VM. | Explicit software CPU emulation; no custom-image rebuild, local emulator or account credential sent to Android. |\n| integrations/hosted_session.py; hosted_env.py | Opt-in bootstrap runtime and remote process environment; preserve existing versioned-image path. | One 4-CPU/8-GB/32-GB VM, 30-minute hard lifetime per attempt. Provisioning logs and archive export retained. |\n| integrations/campaign_budget.py; capped_env.py | Reserve compute and every attempted model call before sending; freeze GPT-4.1, temperature 0.6, output 2048, zero transport retries and at most 18 calls per attempt. | Twelve VM reservations plus 216 requests at at most 8000 input tokens: listed-price upper bound USD 8.866944. No refund of ambiguous call reservations. |\n| Model context and transcript | Send the latest original 1080x2400 PNG, compact visible UI and action-only history. Do not repeatedly resend old images. Save all actual requests and replies for the dashboard. | Input token check includes conservative image/framing allowance; oversized prompts stop before inference, not silently truncate. |\n| tests/test_campaign_budget.py; test_pass_at_k.py | Add budget, no-duplicate, context, wheel-source and mocked real-client adapter checks; update the old null-budget expectation to the explicit approval. | Offline tests are not evaluation outcomes. First packaging test exposed a scanner matching its own private-key regex; narrowed to actual key headers without weakening the key checks. |\n\n\n## Current Prime readiness — 2026-09-10\n\nThe public read-only DemoCart UI was verified inside Prime at 02:45:56 UTC on 2026-09-10. Android, the screenshot HTTP server and the tunnel client all run in Prime sandbox ad78x8509c5zu48npiheabai. A real Chromium browser rendered the app correctly, fresh frame IDs advanced and the original browser screenshot was visually inspected. This is an unscored preview: **0 of 12 model attempts have run**.\n\n[Expired preview URL](https://t-2-b26ceb9ffdc42ae2.tunnel.pinfra.io). The viewing hold ended at **02:53:53 UTC / 04:53:53 CEST on 2026-09-10**; the tunnel and VM are now terminated, and the public endpoint returns 404. It was not a permanent deployment. Public visitors need no token and cannot control the app. The controller exported evidence and deleted this exact tunnel and VM after the ten-minute hold. [Independent cleanup readback](artifacts/prime_public_live_20260910_pid_wait/cleanup-readback.json) confirms termination; the original summary's false tunnel-cleanup flag was a checker limitation because Prime returns a terminated object with HTTP 200. The user approved USD 0.19 for this one preview; a separate model-evaluation budget remains required.\n\n| Component | Directly verified evidence | Remaining boundary / next check |\n|---|---|---|\n| Android without KVM | Explicit software CPU emulation and SwiftShader. Android boot and APK installation passed on the current Prime VM. | Software emulation is slower; this does not establish twelve-attempt reliability. |\n| Patched startup | READY after 149.477 seconds: one exact System UI Close, one absent replacement PID, four empty UI roots, then changed PID 872 to 2602 and clean UI. | Missing-process observations remain bounded; no task action or unknown input is replayed. |\n| Public live UI | HTTPS and real PNG passed. Control POST returned 403; private-path request returned 404. Browser rendered 1080x2400 Android content; two fresh frame IDs; no JavaScript errors. | [Browser receipt](artifacts/prime_public_live_20260910_pid_wait/browser-check.json), [browser screenshot](artifacts/prime_public_live_20260910_pid_wait/public-browser.png), [app screenshot](artifacts/prime_public_live_20260910_pid_wait/public-browser-android.png). These are live-preview proof, not model actions. |\n| Source and tests | 211 offline tests passed, zero skips, including actual Java/D8. Source loader verified with verifiers 0.3.1, allow_eval=false, one dataset row, nineteen turn slots and twelve-attempt maximum. | Strict rubric SHA 68b54be405b06a14c5dee20b32bffee6e7589b8c603d140fa5d37ea9e9bf3d00 unchanged. |\n| Prior recording | Previous successful Prime readiness exported real Android screenrecord video; it decoded successfully. | [MP4](artifacts/prime_software_readiness_20260910_restart_wait/episodes/20260909T232635Z_de316e6f/video/screen-recording.mp4). The completed [preview recording](artifacts/prime_public_live_20260910_pid_wait/episodes/20260910T024311Z_4f7f63bd/video/screen-recording.mp4) also exported and decoded successfully; no model recording exists yet. |\n| Hosted model evaluation | Read-only account check confirms openai/gpt-4.1 with image input. Both earlier custom image tags still FAILED and the Hub source is the older PRIVATE version. | [Campaign preflight](artifacts/prime_public_live_20260910_pid_wait/campaign-preflight.json). Publish a tested usable runtime/source, enforce context/retry/spend limits and obtain a separate total evaluation budget before launching. |\n\nThe first browser-check attempt failed because its own string-evaluation helper was blocked by the viewer's Content Security Policy. The check was changed to normal locator/request assertions; the service's security policy was not weakened. The successful browser run is retained separately from that checker error.\n\nOfficial background: without VM acceleration, Android Emulator translates guest machine code in software; SwiftShader separately renders graphics on the CPU. See [Android acceleration](https://developer.android.com/studio/run/emulator-acceleration). [Prime Tunnel](https://docs.primeintellect.ai/sandboxes/tunnel) supplies the public HTTPS address. Live PNG refresh is not a video-call stream or interactive remote control.\n\nDirectly verified: current public browser rendering, frame freshness, denied controls, startup recovery and safe loader. Not yet verified: a Ready custom runtime-image publication, the new source on the Hub, an enforced full-campaign spending cap, or any GPT-4.1 task result. The [implementation log](IMPLEMENTATION_LOG.md) preserves failures, exact patches and run receipts. The dated sections below are historical.\n\n---\n\n## Previous Prime readiness — 2026-09-09\n\nThe Prime-hosted Android runtime builds DemoCart with Java 21, boots the API-33 AOSP emulator in explicit software mode and installs the APK, but it is **not evaluation-ready**. The three additional readiness runs in this continuation all stopped before creating a task episode. Their original screenshots and Android logs identify an obstructing System UI ANR and UIAutomation startup failures; this is not a GPT-4.1 task failure.\n\nThe infrastructure patches passed **192 offline tests**, including the actual Java/D8 regression, and the final installed wheel loads correctly. No model requests, scored attempts, recordings, new public live URL, Hub publication or custom-image build occurred in this continuation. The app/business logic, strict grading and twelve-attempt maximum are unchanged. All three temporary Prime VMs were exported and confirmed TERMINATED; reusable images were retained.\n\n| Component / file | Exact change or directly observed evidence | Readiness boundary / next check |\n|---|---|---|\n| APK build and installation | All three latest Prime tests built the APK and returned Success from the installation command. Android boot times were 316322, 412957 and 316427 ms. | The startup gate stopped before the episode's installed-APK identity check; do not equate the saved build hash with that later check. |\n| amazon_cart_001/harness/startup_ui.py | Internal temporary XML path is now live-verified. Record System UI Wait attempts before dispatch; use a 30s command allowance within the shared 120s deadline. Unknown outcomes are explicit and never replayed. | Maximum two positively identified System UI Wait attempts; only later clean UI observations can establish readiness. |\n| amazon_cart_001/harness/startup_ui.py | Retry only completed, empty UI dumps: maximum three empty results and a three-second pause within the existing deadline. Unexpected nonempty output, command errors and unknown app dialogs still fail closed. | This newest patch passed offline tests. Its live retest returned exit 137 with a UIAutomation connection timeout, not an empty-root result; recovery is not proven. |\n| amazon_cart_001/harness/runtime_diagnostics.py | Diagnostic PNG capture allowance increased from 10 to 30 seconds. Other six commands unchanged; total ceilings 85 seconds. | Real AOSP5/AOSP6 PNGs captured. They are failure diagnostics, not scored task frames. |\n| tests/test_startup_ui.py; tests/test_runtime_diagnostics.py | Cover unknown tap outcome without replay, bounded empty-root retries and precise diagnostic deadlines. | Full suite: 192 passed in 18.668s, zero skips. See the [test log](artifacts/prime_eval_continuation_20260909/offline-tests.log). |\n| Final installed package | Wheel SHA f116e9af7dbfae20635031e8e72040d9c1a44c7323152459c34e2bee70d88b93; one dataset row, allow_eval=false, 19 turn slots, twelve-attempt guard. | [Package receipt](artifacts/prime_eval_continuation_20260909/package-validation.json). Loading is not deployment or inference. |\n| Prime publication and model | Authenticated devjangid-wootzapp metadata lists openai/gpt-4.1 with image input. The earlier Hub source remains PRIVATE and the two earlier runtime-image builds remain FAILED. | [Metadata receipt](artifacts/prime_eval_continuation_20260909/prime-metadata.json). No newer source or Ready runtime image was published. |\n| Twelve-attempt campaign | Zero attempts started; no pass@3 can be computed. Readiness probes are excluded from the benchmark denominator because no benchmark episode started. | Require stable Prime app interaction, fresh PNGs, decodable MP4, a Ready versioned image, updated Hub source, verified public viewer and an enforced campaign spending cap. |\n\nLatest original receipts: [AOSP4: unknown Wait outcome](artifacts/prime_software_readiness_20260909_aosp4/summary.json), [AOSP5: empty first UI root](artifacts/prime_software_readiness_20260909_aosp5/summary.json), and [AOSP6: UIAutomation connection timeout](artifacts/prime_software_readiness_20260909_aosp6/summary.json). The [latest screenshot](artifacts/prime_software_readiness_20260909_aosp6/diagnostics/screen.png) visibly shows the System UI dialog. The [continuation report](artifacts/prime_eval_continuation_20260909/readiness-status.json) binds exact sandbox IDs, archive hashes, tests and cleanup. The previous six readiness receipts remain unchanged in their own dated directories.\n\nDirectly verified: builds, boots, installation command results, original failure evidence, package tests and exact-ID cleanup. First-boot resource contention is an inference from slow dispatch/lock-contention logs, not a proven sole cause; exit 137 alone does not establish an out-of-memory kill. The current software recipe remains unreliable. There is no successful task action, recording, public live browser rendering or hosted model outcome from these runs.\n\nThe observed account balance decreased from USD 47.6282 to 47.4960 during this continuation (USD 0.1322). Settlement may lag; these are wallet observations, not final itemized charges. No local Docker or emulator operation was performed. The strict rubric identity remains 68b54be405b06a14c5dee20b32bffee6e7589b8c603d140fa5d37ea9e9bf3d00.\n\nNext: establish stable Android UI execution on Prime through a KVM-capable placement or a separately validated emulator configuration before starting the actor campaign. Do not bypass the readiness gate or treat a successful build as a successful evaluation. The [implementation log](IMPLEMENTATION_LOG.md) records each patch and failure. Sections below are historical and do not supersede this status.\n\n---\n\n## Current target and Prime capability checks — 2026-09-09\n\nThe target remains a fully Prime-hosted Android evaluation, not a local evaluation followed by upload and not a Prime controller operating our local emulator. The successful local KVM run below is staging evidence only. At the user's request, the current runtime source was uploaded privately to Prime and a bounded set of sandbox configurations was tried without model calls.\n\nThe upload was accepted, but the image did not reach Ready: Prime's VM artifact build failed and its container artifact was automatically rolled back. Four sandbox configurations were requested; two started and lacked KVM, while two were rejected before creation. Both created sandboxes were deleted and independently read back as TERMINATED. These are capability checks, not benchmark attempts.\n\n| Operation / configuration | Exact ID | Observed result |\n|---|---|---|\n| Image prime/devjangid-wootzapp/democart-peach:0.5.1-kvm2 | Build ml2kj3fy4dsa0r7b6ev23i57; VM artifact qlrso3vtuhz3dl5vttfjd966 | Source upload accepted; VM artifact failed with Build failed with exit code 1. Container artifact rolled back. No usable image. |\n| VM, region us | khi437br5wgcjn51jkm0vxw6 | Ran; /dev/kvm absent and VMX/SVM CPU flags absent. Cleanup verified. |\n| VM, region eu-west | No sandbox ID | HTTP 400: region is not supported for VM sandboxes. Not a negative KVM measurement because no VM started. |\n| Container, region us | fkqyl3w5dr8n59nkner2c2o7 | Ran with vm=false; /dev/kvm absent and VMX/SVM CPU flags absent. Cleanup verified. |\n| Container, region eu-west | No sandbox ID | HTTP 400: Invalid sandbox region. No container started. |\n\n[Final Prime readback](artifacts/prime_kvm_matrix_20260909/final-readback.json), [timestamped event log](artifacts/prime_kvm_matrix_20260909/events.jsonl), [probe summary](artifacts/prime_kvm_matrix_20260909/summary.json) and individual probe JSON files are retained under artifacts/prime_kvm_matrix_20260909. The final readback reports usable_runtime_image=false and image_deletion_requested=false. The initial summary's image_retained flag meant that our helper did not request image deletion; it is not a claim that Prime retained usable image bytes after rollback.\n\nThe image context contained 110 files / 130,051 bytes, SHA-256 b69c703d6cfc07e302901098576c24a85acf9cd00b4331ef8e10fc6f18cb9321. The SDK command-line-tools download's SHA-256 matched the Dockerfile, so a bad pin was not established as the build failure. The returned build metadata includes no detailed build log; its underlying failing command remains unknown. The wallet read back as USD 47.9174 before and after this batch; zero observed delta is not a guarantee of no later-settling charges.\n\nNext: obtain a Prime-supported KVM-capable placement and the failed image build's detailed logs. Do not repeat ordinary sandbox launches indefinitely, invent undocumented privileged settings, weaken the KVM gate or relabel local Android as Prime-hosted. The existing Hub source publication is unchanged; this operation was a runtime image-context upload, not a new source-version publication.\n\nCleanup requirement: retain reusable task images while work is unfinished. Once the full work and verified export of screenshots, recordings, logs and JSON reports are complete, stop/remove the task-owned runtimes and remove only their explicitly identified Docker images. Temporary diagnostic sandboxes are deleted after each probe to avoid idle charges. No local Docker image or unrelated container was removed in this batch.\n\n---\n\n## Staging readiness — local KVM, not Prime-hosted (2026-09-09)\n\nThe local server now boots the peach DemoCart app using real KVM acceleration and an API-33 x86_64 emulator. The new launcher rebuilds the current APK, installs and verifies its hash, captures the clean home screen and SQLite/provider evidence, starts a read-only browser viewer, records Android video, and stops only the processes it created. It performs no shopping actions, model requests, scoring or evaluation upload.\n\nThis is the current readiness result; the dated hosted sections below retain their earlier history. Prime source version 0.5.1 is published, but its Android image build failed and the tested Prime VM exposed neither /dev/kvm nor VMX/SVM CPU flags. Local KVM does not enable nested virtualization on Prime. The hosted integration and strict rubric are preserved; a Prime-controller-to-local-Android bridge and the twelve GPT-4.1 attempts have not run.\n\n| Component / file | Role and exact behavior | Evidence / current boundary |\n|---|---|---|\n| amazon_cart_001/harness/local_runtime.py | Uses a private AVD, ADB server and per-run Android/XDG/temp directories. Requires KVM and free dedicated ports; never replaces another service or falls back to software emulation. | Real API-33 boot passed. ADB 5048; emulator 5584/5585; JWT-protected emulator gRPC 8578. Existing shared ADB 5037 is not stopped. |\n| scripts/prepare_local_kvm.sh | Launches only the readiness module with the existing Python environment. A bounded hold keeps the viewer available; cleanup exports video and removes only its own temporary runtime directory. | Default hold five seconds; maximum 900 seconds. This command is not an evaluation command. |\n| scripts/local_kvm.env.example | Points at existing SDK, Java and ffmpeg executables without changing global configuration. | The SDK at /data/Balram/android-sdk and JDK under /usr/lib/jvm are read-only inputs. New build/evidence/cache data stays under /data/Tirtha. |\n| tests/test_local_runtime.py | Sixteen offline tests cover output boundaries, occupied ports, required tools/KVM, owned-process cleanup, environment restoration and conflicting Android preference paths. | Full task suite: 159 offline tests passed. Offline tests are separate from the real boot/browser checks. |\n| integrations/live_viewer.py; viewer/index.html; viewer/viewer.js | Adds explicit local_readiness context to the status and browser labels, including the SSH-only access explanation. Hosted defaults and input-denial behavior are unchanged. | Three new regression tests and a real Chromium check confirm this distinction. |\n| Real readiness receipt | Run 20260909T144945Z_8uc8_x6j: clean app state, matching installed APK hash, original PNG, browser rendering with no JavaScript errors, and decodable Android screen recording. | No model or task action. Viewer control POST returned 403; private-file request returned 404. Temporary emulator, private ADB and runtime data were removed after evidence export. |\n\n### Run the local readiness check\n\n~~~bash\ncd /data/Tirtha/browser-is-all-you-need/android-apk-rl-adk/environments/amazon_cart_001\nsource scripts/local_kvm.env.example\nbash scripts/prepare_local_kvm.sh\n~~~\n\nFor up to fifteen minutes of read-only UI inspection, use `bash scripts/prepare_local_kvm.sh --hold-seconds 900`. After Android boots, the command prints the actual `http://127.0.0.1:<port>` URL. The port is allocated per run; an old printed URL is not a permanent endpoint. To view it from your computer, keep the launcher running and forward that printed port through SSH:\n\n~~~bash\nssh -N -L 8765:127.0.0.1:<printed-port> ubuntu@static.235.31.55.162.clients.your-server.de\n~~~\n\nThen open `http://127.0.0.1:8765` locally. Replace `<printed-port>` with the numeric port in the running launcher's output. This uses SSH access to our server; it is neither a public Prime link nor proof of Prime-hosted Android.\n\nSaved evidence: [readiness report](artifacts/local_kvm/20260909T144945Z_8uc8_x6j/readiness.json), [initial app screenshot](artifacts/local_kvm/20260909T144945Z_8uc8_x6j/frames/000/screen.png), [browser screenshot](artifacts/local_kvm/20260909T144945Z_8uc8_x6j/live-viewer-browser.png), [browser checks](artifacts/local_kvm/20260909T144945Z_8uc8_x6j/viewer-check.json), and [screen recording](artifacts/local_kvm/20260909T144945Z_8uc8_x6j/video/screen-recording.mp4). These are local generated artifacts, not newly published Hub files. The recorded session has ended. The viewer explicitly labels this LOCAL KVM RUNTIME and NOT AN EVALUATION; it does not present the local test as a Prime model run.\n\nDirectly verified: local KVM, actual app boot, installed APK identity, readable initial evidence, browser display, blocked controls, recording decode, and cleanup. The strict rubric identity remains `68b54be405b06a14c5dee20b32bffee6e7589b8c603d140fa5d37ea9e9bf3d00`. No inference is drawn from an empty home screen about task success. Hosted KVM, a public Prime live URL, GPT-4.1 execution and any pass@3 result remain unverified. Existing legacy local model commands retain their older free-provider guard; this launcher does not silently change them.\n\nNext: establish the Prime-side KVM runtime before launching the frozen twelve-attempt campaign. Local or hybrid execution is not an authorized substitute for the fully hosted target. Local KVM readiness alone is not hosted evaluation readiness. Android's official guidance explains [emulator acceleration](https://developer.android.com/studio/run/emulator-acceleration) and the separate [SDK, user and AVD paths](https://developer.android.com/tools/variables).\n\n---\n\n## New account and strict hosted campaign — version 0.5.1 (2026-09-09)\n\nThis version prepares the same peach-header Java/XML/SQLite shopping task for **devjangid-wootzapp**, using the newly supplied account rather than terrano09. The actor must still search separately for the three specified products, add exactly one of each, and open the INR 4,997 cart. The app, task values, screenshots and historical results have not been edited to improve a score.\n\nThe strict scoring profile is implemented, and **143 offline tests passed**. At the user's request, the active actor is now **GPT-4.1**, using Prime Inference directly as **openai/gpt-4.1**. Official OpenAI documentation confirms image input and structured outputs; GPT-4.1 has no reasoning-effort setting. An authenticated lookup of this account's 116-model Prime catalog lists the exact full GPT-4.1 model with image input, USD 2 per million input tokens and USD 8 per million output tokens; cached pricing was unspecified. These were documentation/metadata checks, not inference. The runtime image, actual Prime Android/KVM boot, live tunnel, secret delivery and 12 hosted attempts remain unverified. Source publication is separate from image building. The user has requested a PUBLIC Hub page and a public read-only live UI: public_viewer=true removes the login requirement for live frames only. Protected inspection remains available separately. A total spending cap is still required before billed work.\n\n| File / component | What changed in the code | What it establishes and what it does not |\n|---|---|---|\n| verification/strict.py, specs/strict_rubric.json | Adds six mandatory checks: task/APK identity, frozen evidence integrity, timeline, action audit, current search-to-cart lineage, and exact final outcome. Reads original files again instead of trusting cached derived state. | All required checks PASS gives 1; trusted FAIL gives -1; otherwise missing/unreliable required evidence gives INVALID=0. No result depends on the desired pass rate. |\n| harness/peach_episode.py | Adds explicit rubric_profile; saves dispatch/NNN.xml immediately before targeted actions, with its hash. Checkpoints expose strict checks and retain PENDING/null terminal reward until actor stop. | Separates six-stage progress, old policy votes and final reward. Local default remains legacy_v1; no historical run is rescored or relabeled. |\n| verification/strict.py::action_audit | Matches raw action, schema, permitted target, pre-action/dispatch XML, ADB commands/stdout hash, and committed app mutation. Preserves rejection and extra-product constraints. | A successful tap alone cannot prove acceptance. Missing receipts cannot silently pass. |\n| verification/strict.py::cart_lineage | Binds each present SKU to the matching search when its quantity last changed from zero to positive. Removal/zero clears the association. Requires three distinct searches. | Old add/remove evidence cannot justify an unsearched re-add. Identical units have no invented FIFO/LIFO identity. |\n| specs/campaign_12.json | Records owner, proposed image/version, 12 sequential attempts, zero replacements/retries, sampling proposal, estimator and reporting-only target. Budget remains null pending approval. | A preparation manifest, not a claim that 12 attempts ran. Freeze actual model/settings, rubric, task and source digests before launch. |\n| verification/pass_at_k.py | Validates unique attempt/episode IDs, matching rubric/task/model/sampling/campaign hashes and signed verdicts. Retains setup INVALID without invented Android episodes. | Does not pool different settings, drop INVALID runs, stop at a desired result, or fabricate successes. |\n| integrations/hosted_env.py, hosted_session.py, harness/hosted_worker.py | Hosted default is peach_strict_v1; forwards it into the VM, records the Prime trajectory ID, and guards 12 starts per environment instance. Uses prime/devjangid-wootzapp/democart-peach:0.5.1. | CLI rollout count and zero retries set the campaign total; an in-process counter is not a distributed budget. Importing/loading creates no resources. |\n| tests/test_strict_verification.py, test_pass_at_k.py, test_hosted_runtime.py | Adds 30 scoring/hosted tests and ten public-viewer tests over the previous 103, using synthetic SQLite/XML/PNG trajectories, corruption, malformed actions, receipt mismatch, lineage, arithmetic and hosted guards. | 143 offline tests passed. Fake devices and synthetic results are not an emulator run or model evaluation. |\n| pyproject.toml, __init__.py | Prepares package version 0.5.1 with the new implementation/specifications. Dependencies and Android source are unchanged. | The terrano09 0.4.0 publication remains historical and untouched remotely. |\n\n### Strict rubric and policy receipts\n\nThe original fourteen policies still have five verifier entries each, and their original .5 + .1 × PASS-count results remain visible as diagnostics. They no longer decide the new profile's final reward by themselves: two correlated votes cannot override a failed mandatory check. The strict JSON maps six required checks to all fourteen policies. These are six explicit additional checks, not seventy new independent evidence sources.\n\nAll required frozen files must have matching hashes and episode identity. SQLite, provider and journal share an app/database; agreement detects inconsistency but is not independent ground truth. PNG integrity is not visual-semantic grading. When a cart subtotal is visible, it is cross-checked; being off-screen is an optional-channel INVALID, not a task failure. A visible contradiction prevents a clean pass.\n\nA readable empty cart is FAIL, not INVALID. Corrupt essential state or missing dispatch evidence is INVALID unless another required check already establishes a trusted failure. Malformed/forbidden actions, app rejections, premature finish and extra-product adds retain their original action-contract failure semantics even if the final cart is later corrected. Verifiers are inspection-only.\n\n### Twelve-attempt reporting\n\nUse **pass@3 = 1 - C(n-c,3) / C(n,3)**. With twelve valid attempts, four successes gives **0.74545**, five gives **0.84091**, and six gives **0.90909**. Only four or five successes fall in the requested 0.70–0.90 band. This is an observed performance target, not a grading input or quota.\n\nRun all twelve; never stop at the target or replace INVALID attempts silently. If any attempt is INVALID, the primary twelve-valid-attempt estimate and target verdict stay unavailable. A conditional valid-attempt estimate and a separately labeled operational estimate counting invalid as non-success are still reported. The confidence summary explicitly assumes IID attempts and does not establish competence across different tasks.\n\n~~~bash\npython -m amazon_cart_001.verification.pass_at_k path/to/all_attempts.json\n~~~\n\nThe JSON list must be exported from actual hosted receipts. Each record needs attempt_id, episode_id (null only for documented setup INVALID), actual model, rubric_sha256, task_sha256, actor_config_sha256, campaign_sha256, status and reward. Hash actual launch/request settings; do not infer them from a model name.\n\n### Future hosted launch — approval and preflight required\n\nConfirm Prime identity is devjangid-wootzapp, publish version 0.5.1 with runtime v0, confirm the Hub visibility is PUBLIC, then build and verify the private 0.5.1 runtime image. The native Prime model route uses the funded Prime account; it does not require a separate OpenRouter or OpenAI key. Keep the historical local runner's free-only guard unchanged. The public campaign does not read or require DEMOCART_VIEWER_TOKEN. A protected-viewer secret was created under the earlier approved private-view plan; it stays secret and is unused by public viewing. No secret is placed in source, public HTML, environment arguments or screenshots.\n\nThis command has **not** been run:\n\n~~~bash\nprime --plain eval run devjangid-wootzapp/amazon-cart-001 \\\n  --hosted --follow --allow-sandbox-access --allow-tunnel-access \\\n  --model openai/gpt-4.1 \\\n  --num-examples 1 --rollouts-per-example 12 --max-concurrent 1 --max-retries 0 \\\n  --sampling-args '{\"temperature\":0.6,\"max_tokens\":2048,\"response_format\":{\"type\":\"json_object\"}}' \\\n  --timeout-minutes 240 \\\n  --env-args '{\"allow_eval\":true,\"live_viewer\":true,\"public_viewer\":true,\"rubric_profile\":\"peach_strict_v1\",\"max_attempts\":12,\"sandbox_image\":\"prime/devjangid-wootzapp/democart-peach:0.5.1\"}' \\\n  --eval-name democart-prime-strict-12-v1\n~~~\n\nThe CLI currently uses the latest published environment at hosted launch: verify that latest is the reviewed 0.5.1 source, preserve its resolved version/source hash, and do not publish during the campaign. Freeze sampling and recheck GPT-4.1 availability, image support and pricing before launch. Use the explicitly selected model; do not silently fall back to another model. This is paid Prime Inference. Omitting an external API base selects the native hosted inference path. Prime documents that the evaluation-controller runtime is not billed separately on that path; the additional Android VM created by this environment remains a separate resource. For this public campaign, verify anonymous HTTPS status and PNG delivery, plus rejection of every input request, before launching. Timeouts are not dollar spending caps.\n\nThe live UI is public and read-only during model attempts. A separate interactive inspection is not scored and must use separate app state. Each attempt retains original screenshots, video, policy checkpoints and strict verdict, exports evidence, then deletes its VM. Anyone with the live URL can watch without a password; every POST is rejected, and no private logs, artifacts or transport details are served. No live URL exists until the real image, KVM boot and anonymous HTTPS frame check succeed.\n\nDirectly verified: account identity, private 0.5.0 publication and matching source hash, protected-viewer secret metadata, 143 offline tests, ten public-browser checks on a labeled historical fixture, and authenticated Prime GPT-4.1 capability/pricing metadata. Inferred: strict scoring and viewer are wired into the hosted worker. Unverified: image build, KVM, live UI, hosted secret delivery, twelve attempts and dashboard media rendering. Pass-rate range is not guaranteed.\n\nReferences: [pass@k estimator](https://github.com/openai/human-eval/blob/master/human_eval/evaluation.py), [GPT-4.1 official model documentation](https://developers.openai.com/api/docs/models/gpt-4.1), [OpenAI pricing](https://developers.openai.com/api/docs/pricing), [Prime hosted evaluations](https://docs.primeintellect.ai/tutorials-environments/hosted-evaluations).\n\n---\n\n### Public viewer patch — version 0.5.1\n\nThe user explicitly requested that anyone can open the live UI without a token. Public viewing is an opt-in mode, not public device control. The app and strict grading rules are unchanged. The campaign JSON sets public_viewer=true; the library default remains protected so older inspection callers do not accidentally become public.\n\n| File | Exact code change | Verification / remaining boundary |\n|---|---|---|\n| integrations/live_viewer.py | Adds public_readonly; unauthenticated mode/status/original-PNG endpoints, all POSTs denied, no token generated, private capture errors replaced by a generic message. Rejects public-plus-interactive/control/secret combinations. | Public HTTP unit tests pass; protected tests still pass. Only three static assets and two frame/status APIs are exposed. |\n| integrations/viewer_tunnel.py | For public mode, verifies anonymous HTTP 200 and original PNG plus control HTTP 403. Records access_mode=public_readonly and writes no viewer access-token file. Keeps Prime tunnel transport credentials private. | Offline mocked relay test passes; real Prime tunnel remains unverified. |\n| viewer/viewer.js, viewer/index.html | Discovers mode, automatically connects without a token, hides login/controls in public mode, and does not show a password prompt after transport failure. | Ten headless browser checks pass with an OFFLINE TEST historical fixture. CSP remains unchanged; a test wait was changed to Playwright locator assertions to respect CSP. |\n| integrations/hosted_env.py | Adds explicit public_viewer flag; skips only the viewer-secret requirement, never the allow_eval guard, twelve-attempt guard or no-human-control checks. | Hosted wiring and guards are tested with mocks; no model/VM. |\n| specs/campaign_12.json, integrations/hosted_session.py, pyproject.toml, __init__.py | Pins public viewing and PUBLIC Hub intent; advances package/image reference to 0.5.1. GPT-4.1, twelve attempts and strict rubric are unchanged. | Spending cap remains null; actual launch manifest not frozen yet. |\n| tests/test_public_viewer.py | Adds ten dedicated tests for anonymous viewing, blocked input, path allowlisting, redaction, freshness, secret isolation and tunnel attestation. | Full suite: 143 tests passed. No benchmark attempts were generated. |\n\nA Hub page and a live Android URL are different. Changing Hub visibility shares the source package. The live link is generated only for a running Prime Android session, and becomes unavailable after that session is deleted. Public access is read-only even for a logged-in owner.\n\n\n## Historical v0.4.0 setup — terrano09, not the current launch account\n\nThe following records the previous account's setup and funding failure. Its owner/image commands are historical; use the new-account plan above.\n\n## Secured Prime live viewer — version 0.4.0 (2026-09-09)\n\nThe default environment uses the peach-header Java/XML/SQLite app, package `com.primeintellect.amazonuidemo`. A hosted model loop creates one Prime VM, boots Android, sends the model's actions to that device, and exports screenshots, recording and policy receipts before deleting the VM. The new browser viewer relays actual Android PNGs through an authenticated Prime HTTPS tunnel. During a separate UI inspection, it supports tapping, dragging, Back and typing into a focused field; during a model evaluation, every human control is disabled.\n\nThe viewer is implemented and tested offline, but no new Prime image, Android VM, tunnel URL, smoke test or evaluation has been created. The last image-build attempt received an actual `Payment required. Check billing status.` response; the account balance was read back as USD 0.00. Creating the new viewer-access secret is also awaiting explicit approval. A Hub page is code hosting, not a live Android screen. The browser-test preview uses a labeled historical PNG fixture and is not proof of live Prime rendering.\n\n| Component / file | Exact implementation | Verified or remaining boundary |\n|---|---|---|\n| `integrations/prime_env.py` | Defaults to the hosted peach adapter. Older local ADB profile requires explicit `execution=\"legacy_local_adb\"`. | Loading alone creates no VM and calls no model. |\n| `integrations/hosted_env.py` | Requires `allow_eval=true`; live-view runs also require `DEMOCART_VIEWER_TOKEN` before any billed session. Starts a read-only viewer and logs its public URL. Stops the viewer before final export and cleanup. | A true hosted model run requires `--hosted`; no local-run/upload fallback is used. |\n| `integrations/hosted_session.py` | Creates one Prime VM: 4 CPU, 8 GB RAM, 32 GB disk, 20-minute lifetime. Frames must match the episode, immutable control mode, path, size and SHA-256. Evidence export precedes VM deletion. | Actual image readiness and usable nested KVM are unverified. Timeouts are not spending caps. |\n| `harness/hosted_worker.py` | Ordered reset/step/finalize mailbox; KVM gate; fresh API-33 emulator; installed APK hash; Android screenrecord and per-action evidence. Inspection finalization is explicitly NOT_SCORED. | Manual inspection and model evaluation use separate workers and fresh app state. |\n| `harness/live_device.py` | Captures original PNGs, retains eight recent frame pairs, and checks frame identity, age and coordinate bounds before inspection input. Writes manual-input receipts without calling them model actions. | The Android worker independently rejects controls in evaluation mode. No public ADB endpoint. |\n| `integrations/live_viewer.py` | Localhost HTTP server protects image/status/control APIs with a bearer token; no-cache responses, restricted routes and content-security headers. Accounts for capture/download latency and repeated stale frames. | This is periodically refreshed imagery, not a high-frame-rate remote-desktop stream. |\n| `integrations/viewer_tunnel.py` | Uses pinned `prime-tunnel==0.1.11` and run-local private transport configuration. Returns a URL only after checking public HTTPS, unauthenticated 401 and authenticated PNG delivery. Removes its tunnel/config/access link on shutdown. | Real Prime tunnel/authentication checks have not run. HTTP/TCP sandbox exposure is not assumed available for VM sandboxes. |\n| `viewer/index.html`, `viewer.css`, `viewer.js` | Browser login, frame age, displayed-image coordinate mapping, inspection controls, evaluation read-only display and disconnected/stale warnings. URL-fragment tokens are removed from browser history and never logged. | Headless browser checks passed with an explicit OFFLINE TEST fixture. Prime evaluation-dashboard inline video rendering remains unverified. |\n| `ui_smoke.py` | Explicit `--allow-compute` guard; boots a Prime VM, opens a temporary interactive viewer for 15–300 seconds, then exports evidence and deletes the VM. | No model calls or benchmark result upload. If launched locally, the local process is only the viewer/controller; Android runs on Prime. |\n| `Dockerfile`, `.dockerignore`, `pyproject.toml` | Current app source and SDK/recording tools are included in the private image recipe. Viewer assets ship in the wheel. Source context excludes generated runs, credentials and caches. | Publishing the wheel does not build the Android runtime image. Publish with `--runtime v0` for the classic MultiTurnEnv API. |\n| `tests/test_live_viewer.py`, `tests/test_hosted_runtime.py` | Authentication, schema, read-only enforcement, freshness, episode identity, secret preflight, private-file permissions, lifecycle and packaging regressions. | Offline tests must not be counted as a Prime smoke test or model evaluation. |\n\nThe task and reward contract are unchanged: 14 policies with five registered slots each; 34 implemented and 36 explicitly unavailable slots. Policy support is `0.5 + 0.1 × PASS count`, with inclusive `>= 0.7` and an all-INVALID override. Six workflow stages are progress diagnostics, not extra rewards. Generic interaction checks and task-verifier outcomes remain separate from viewer-health checks.\n\n### First: publish code and build the runtime image\n\nFrom this environment directory, using the authenticated terrano09 personal account:\n\n```bash\nprime --plain env push --owner terrano09 --name amazon-cart-001 --runtime v0\nprime --plain env info terrano09/amazon-cart-001\nprime --plain images push democart-peach:0.4.0 --dockerfile Dockerfile --context .\nprime --plain images list\n```\n\nKeep the image private. The intended image is `prime/terrano09/democart-peach:0.4.0`; it is not yet built. Stop if billing or image construction fails. Confirm actual Ready status before creating a sandbox. Docker cannot itself supply host KVM; the worker refuses to silently use a slow CPU fallback.\n\nFor a hosted viewer, an authorized operator must configure an encrypted environment secret named `DEMOCART_VIEWER_TOKEN` containing a newly generated random URL-safe 32-byte token. A 43–128-character URL-safe shape is enforced, but arbitrary repeated characters are not a secure password. Do not put the value in source, CLI arguments, logs, environment arguments, screenshots, receipts or the model prompt. Creation of that new credential is pending user approval. The model provider key is separate and must also be available as the intended hosted secret before a model evaluation.\n\n### Next: a compute-only interactive UI smoke test\n\nAfter image readiness, funding and permission, the manual inspection command is:\n\n```bash\npython -m amazon_cart_001.ui_smoke \\\n  --allow-compute \\\n  --image prime/terrano09/democart-peach:0.4.0 \\\n  --output-dir artifacts/prime_ui_smoke \\\n  --hold-seconds 180\n```\n\nThis prints a public tunnel URL only after real Android captures and authenticated HTTPS checks succeed. The matching private access link is stored temporarily at `artifacts/prime_ui_smoke/<run_id>/.viewer-private/access.json` with mode 0600, inside a 0700 directory. Hand that link privately to the authorized viewer, or open the public URL and enter the approved access token. If no configured token is supplied to this inspection command, it generates a session-only token; it does not create a durable Hub secret. The URL is temporary and stops working after teardown. No placeholder or historical screenshot is presented as a live session.\n\nA local invocation of this inspection utility is not itself a Prime-hosted evaluation. Its Android VM and app run on Prime; its controller and browser relay run where the command was launched. The evaluation below runs the model/controller on Prime as well and always opens a new VM, so manual inspection cannot complete the task for the evaluated model.\n\n### Only after the UI smoke passes: one real hosted evaluation\n\n```bash\nprime --plain eval run terrano09/amazon-cart-001 \\\n  --hosted --follow --allow-sandbox-access --allow-tunnel-access \\\n  --model dots-studio/dots-3-note-preview:free \\\n  --api-base-url https://openrouter.ai/api/v1 \\\n  --api-key-var OPENROUTER_API_KEY \\\n  --api-client-type openai_chat_completions \\\n  --num-examples 1 --rollouts-per-example 1 --max-concurrent 1 --max-retries 0 \\\n  --timeout-minutes 120 \\\n  --env-args '{\"allow_eval\":true,\"live_viewer\":true,\"sandbox_image\":\"prime/terrano09/democart-peach:0.4.0\"}' \\\n  --eval-name democart-peach-hosted-live\n```\n\nThis is a future launch command, not a recorded run. Recheck model availability, image input support and zero-priced routing before launch; do not silently switch to a paid model. Both `OPENROUTER_API_KEY` and `DEMOCART_VIEWER_TOKEN` must be explicitly available in the hosted environment; local shell variables alone do not prove secret delivery. The sandbox/tunnel flags grant the hosted run access to those Prime APIs. The hosted timeout is separate from the adapter's 900-second rollout and 20-minute Android VM lifetime; none is a dollar cap.\n\nWatch hosted logs for `DEMOCART_LIVE_VIEWER`. Its receipt contains the public HTTPS URL and health checks, never the password. Open it and enter the separately supplied viewer token. During evaluation the page is read-only, including direct API calls: taps cannot affect the model's task. The live viewer is a separate Prime tunnel page, not a claim that Prime's evaluation metrics page embeds an interactive Android desktop.\n\n### Saved evidence and validation boundaries\n\nInitial and after-action screenshots remain original PNGs in the model conversation. Final sample info retains policy checkpoints, the recording and a hash-checked ZIP of logs/evidence. The archive limit is 100 MiB; the embedded dashboard payload limit is 25 MiB. Export/cleanup errors remain explicit INVALID pipeline outcomes and preserve any available task verdict separately. The private viewer/tunnel credentials are not part of the worker evidence archive or uploaded sample.\n\nDirectly verified: offline security/interaction tests, real browser rendering of an explicitly historical fixture, JavaScript syntax and package builds. Inferred from code: live capture and tunnel lifecycle are wired into the hosted rollout. Unverified: Prime image build, KVM boot, actual live HTTPS viewing and Prime dashboard inline screenshot/video rendering. Do not label those pending checks as passed.\n\nReferences: [Prime tunnels and hosted permissions](https://docs.primeintellect.ai/sandboxes/tunnel), [Sandbox SDK and VM exposure limits](https://docs.primeintellect.ai/sandboxes/sdk), [Hosted environment secrets](https://docs.primeintellect.ai/tutorials-environments/secrets), [Hosted evaluations](https://docs.primeintellect.ai/tutorials-environments/hosted-evaluations), [Prime Images](https://docs.primeintellect.ai/sandboxes/images).\n\n## Legacy local profile and historical implementation\n\nSearch separately for Nimbus Wireless Headphones, Trail Steel Water Bottle and Metro Laptop Backpack. Add one of each, no extras, and open the cart. The expected subtotal is INR 4,997. The app is a controlled offline Android demo, with Material Views components. Its Java app, Python action harness, evidence collection, policy verifiers and Prime adapter are packaged together.\n\nVersion 0.2.0 uses the common task layout below. The earlier ten-action, eleven-screenshot run was `scripted-ui-validation (no model)`. It remains a scripted run. This version adds a genuine OpenRouter runner, but no new model evaluation was launched during this cleanup. Generated files and historical evaluation artifacts are excluded from Git and distribution packages.\n\n## File tree\n\n```text\namazon_cart_001/\n├── pyproject.toml\n├── README.md\n├── Dockerfile.software        # Ubuntu + Android SDK + APK provisioning recipe (software/TCG mode)\n├── app/amazon_android_ui/     # the current DemoCart app (manifest, build_apk.sh, src, res)\n├── amazon_cart_001/           # the task implementation package (harness, verification, integrations, specs)\n├── amazon_cart_2/             # Hub-slug shim re-exporting the public loader\n├── scripts/                   # provision_android.sh (Prime recipe), run_eval.sh (CLI), local KVM helpers\n└── tests/                     # offline verifier/harness suites (not shipped in the Hub package)\n```\n\nThe legacy shopping demo app and its helper scripts were removed on 2026-09-11; the current app is `app/amazon_android_ui` (peach UI).\n\nTask-specific extras are retained where real functionality exists: payment has Android instrumentation/capture tests; DemoCart has `Catalog.java`, vector product illustrations, its scripted demonstration CLI and `integrations/legacy_scripted.py` for the immutable historical scripted export. No empty matching files are manufactured.\n\n## What the files do\n\n| Files | Code responsibility | Validation or boundary |\n|---|---|---|\n| `specs/task.json` | The only task definition; its values and step budget drive the actor instruction and expected outcome. | The duplicate root task file is removed. Registry checks reject missing or incompatible configuration. |\n| `MainActivity.java`, `AppState.java`, `StateStore.java` | Render controls, validate business actions, persist state and record action acceptance. | `AppState.java` (formerly `ShoppingState.java`) handles searches, cart quantities and operation history. `Catalog.java` defines six fictional offline products. The read-only journal verifier independently replays accepted operations and compares reconstructed business state. There is no Amazon login, live listing, order, checkout or purchase. |\n| `harness/actions.py`, `acceptance.py`, `episode.py` | Validate JSON and permitted targets, execute fresh UI actions, separate tool execution from app acceptance, reject an early finish. | Malformed actions are controlled errors; the actor is not repaired or silently replaced by a script. |\n| `device.py`, `evidence.py` | Capture reset plus each action screenshot, UI tree, persisted state, read-only runtime snapshot, journal and ADB receipts. | Hashes and episode identities bind evidence; OCR is optional and cannot fabricate a missing vote. |\n| `verification/` | Evaluate 15 policies using five registered strategies each; expose reasons, evidence references and conflicting votes. | Read-only inspection; no verifier completes the task for the actor. |\n| `integrations/prime_env.py`, `prime_upload.py` | Expose the same episode through `verifiers.MultiTurnEnv`; upload an explicitly selected genuine model result with original logs/screenshots. | Loading alone starts nothing. Local KVM execution and uploaded results are not Prime-hosted compute. |\n| `pyproject.toml`, build scripts | Package the Python implementation and Android source; build the APK using pinned Gradle/Android dependencies. | SDK/JDK paths select installed tools. Cache/temp outputs default to task-local artifacts. No generated APKs, signing keys, caches or saved runs are committed. |\n\n## Exact policy and reward contract\n\nV1–V2 check specification/capability readiness and evidence integrity. G1–G4 check action format, permission, target availability and execution. S1 checks task-bound semantic selections. T1 exact three product identities; T2 quantity one each; T3 unit prices/currency; T4 three distinct search-to-add links; T5 subtotal/counter arithmetic; T6 cart review screen; T7 accepted action history.\n\nEvery verifier returns PASS = 1, INVALID = 0, FAIL = -1. A policy's support is `0.5 + 0.1 * passed_verifiers`; `>= 0.7` passes. All five INVALID returns 0; otherwise insufficient support returns -1. Two PASS votes therefore outweigh three FAIL votes. This is not 70% agreement, not a probability and not a weighted sum of stages.\n\nAny failed policy makes the episode -1; otherwise any invalid policy makes it 0; otherwise all fifteen passing gives 1. Pipeline-invalid episodes must be excluded explicitly before training normalization. Six workflow-stage indicators show progress separately from terminal reward. Evidence channels are correlated; five registrations do not imply five independent proofs. If a channel cannot establish a policy's full claim, it returns INVALID, not a guessed PASS.\n\n## Commands\n\nRun from this task directory. Keep credentials outside source files. Installing and offline tests do not start an evaluation:\n\n```bash\npython3 -m pip install -e .\npython3 -m unittest discover -s tests -p 'test_*.py'\n```\n\nFor an app build, point `ANDROID_SDK_ROOT` and `JAVA_HOME` at installed tools; `GRADLE_BIN` may point at an existing Gradle 8.9 executable. Build outputs stay in this task. Set `GRADLE_USER_HOME` and `TMPDIR` under `/data/Tirtha` when using the restricted server workspace.\n\n```bash\nbash app/amazon_android_ui/build_apk.sh\n```\n\nAfter explicitly preparing/installing on a chosen emulator, a separately approved model evaluation can be launched with:\n\n```bash\npython3 -m amazon_cart_001.cli \\\n  --confirm-eval --serial emulator-5556 \\\n  --apk app/amazon_android_ui/build/out/democart-ui.apk \\\n  --output-dir artifacts/model --model '<approved-model-id>:free'\n```\n\nThe runner requires `OPENROUTER_API_KEY` and checks current zero-price/image-input model capability at run time. It does not authorize paid-model fallback.\n\nSaved runs contain per-action frames, trajectory and ADB logs, model request receipts, policy results, progress history, final verdict, metadata and a hash manifest. Prime uploading is a separate explicit action:\n\n```bash\npython3 -m amazon_cart_001.integrations.prime_upload --help\n```\n\nOffline test fixtures are not model evaluations. Previous live UI validation does not prove the renamed/model-enabled package has completed a new live rollout. No such rollout or new Prime upload was performed as part of this cleanup.\n\n## 2026-09-10 — Readback latency repair before model attempt 3\n\nThe unscored Prime check e625726e13b7433aa8f1df79d9b1aaf9 proved that the longer startup window works: Android became ready after 189.373 seconds. Its first text action delivered the full search string, but the new confirmation query timed out after five seconds. Other successful provider reads in that same evidence took up to 18.288 seconds. This was a confirmation timeout, not proof of missing input.\n\n`amazon_cart_001/harness/peach.py` now permits each read-only confirmation query up to 30 seconds, matching ordinary state capture, within a 60-second stage deadline and at most five reads. It never repeats input or submits a search. `tests/test_text_readback_timeout.py` covers slow healthy reads and bounded incomplete-state failure with virtual time. The source suite passes all 246 tests; installed-wheel and live validation are recorded separately. Task, prompt, reward rules and the two previous failed model attempts remain unchanged. The failed unscored worker was exported and its termination verified; it is not model attempt 3.\n\n## 2026-09-10 — Exact character transport candidate before attempt 3\n\nThe readback30 unscored check on Prime worker `oj34sov5b2wwlansela6s8dh` passed startup but failed its first text action. Five successful provider reads and the final screenshot showed only `Ni` after ADB returned zero for `Nimbus Wireless Headphones`. The original evidence is in `artifacts/prime_hosted_20260910/harness054-prime-check-readback30/`. The worker was terminated after exporting two screenshots and video; cumulative observed spend was USD 1.019. No model attempt was submitted.\n\n`harness/peach.py` now dispatches each requested character exactly once with native ADB input, bounded by a 120-second dispatch deadline. Full-text readback remains required before and after keyboard dismissal. There is no correction, suffix replay, automatic search or app/database write shortcut. `verification/strict.py` recognizes the versioned `persisted_text_chars_v1` receipt and requires the exact character command sequence and hash-bound state probes; legacy receipts retain their original batch-input contract. `tests/test_character_input.py` covers exact ordering, malformed character receipts and timeout behavior. Source and installed-wheel suites each passed 249 tests. Live validation of this candidate is still pending; offline success does not establish a successful hosted evaluation. App UI, task requirements, prompt and reward mapping are unchanged.\n\n## 2026-09-10 — Bounded UI-read retry and spending-approval stop\n\nThe character-input preflight on Prime worker `biyimlctziygre38p8jd5smq` failed before sending any text: the just-before-action UI dump returned an empty root, with `ERROR: null root node returned by UiTestAutomationBridge.` A later original observation successfully captured the app. Evidence and video were exported to `artifacts/prime_hosted_20260910/harness054-prime-check-chars/`; termination was verified. The cumulative wallet delta at cleanup was USD 1.0937, subject to later settlement. This was not a model attempt and did not validate character delivery.\n\n`amazon_cart_001/harness/peach.py` now tries at most three fresh UI dumps only for that exact empty-root response; all other errors fail immediately. Each attempt cleans up its own temporary XML. No tap, typing, search or model request is retried. `verification/strict.py` explicitly audits each failed dump/cleanup pair before accepting the successful dispatch XML. `tests/test_ui_read_retry.py` covers recovery, retry exhaustion, unrelated errors and altered command receipts. Its first run caught an overescaped path regex, which was corrected before deployment. The corrected source and installed-wheel suites each pass 254 tests, with no skipped tests.\n\nThe candidate wheel is in `.agent_work/prime-hosted-check-20260909.3uLX40/dist054uiretryfixed/` relative to the workspace root. Test logs are `campaign054-uiretry-fixed-offline-tests.log` and `campaign054-uiretry-fixed-installed-tests.log` in the same work directory. The one-shot check helper is `check_harness054_uiretry.py`; the model-attempt helper requires a successful complete live check first.\n\nThe next live check was rejected by the spending-approval safeguard before its process was created. The rejection recognized the earlier USD 0.10 approval rather than the larger reservation referenced by the continuation scripts. No workaround or new VM was attempted afterward. Explicit additional spending approval is required before resuming. The current candidate remains unpublished and unvalidated live; model attempt 3 has not been submitted. Two original model attempts remain counted, both FAIL, and the task/reward settings remain unchanged.\n\n## 2026-09-10 — Additional continuation approved\n\nThe user explicitly approved up to USD 0.90 additional for one final unscored harness check and, only if that check passes, one GPT-4.1 hosted model evaluation numbered 3. The one-shot scripts record this approval and reserve USD 0.156 for the preflight plus USD 0.738912 for the model attempt. The model launcher checks the new wallet baseline as well as the earlier overall cap. No automatic replacement preflight or model retry is authorized by this approval. The earlier spending-approval blocker is resolved; live validation is still pending.\n\n## 2026-09-10 — Recorder/actor trace race found by the approved check\n\nThe approved unscored run `be7c49277c6749df8ad8e59131d43fd4` on VM `d0nw7oru4asp1ldg9h709yhc` completed seven task actions and reached four of six workflow stages. Both full search strings were delivered and confirmed. Action 7 successfully added the bottle, but its audit became INVALID with `DISPATCH_COMMAND_UNEXPECTED`. The original trace entries 138–139 were a background video-segment copy and cleanup, interleaved ahead of the UI dump. The recorder shared the actor's mutable device trace and phase; this was an instrumentation race, not an app rejection.\n\n`amazon_cart_001/harness/recording.py` now creates its own Device connection wrapper using the same explicit ADB binary and serial. Recording commands use phase `recording` and a separate trace, preserved as `command_trace` in `video/recording.json`. It does not filter or erase actor receipts. `tests/test_recording_isolation.py` tests separate endpoints/traces, actual concurrent command overlap, and retained recorder errors. The local source and installed-wheel suites each pass 257 tests. This new candidate is built in `dist054recording/` and installed in `install054recording/` under the scoped work directory; it is not yet published or live-validated.\n\nThe original run remains FAILED/NOT_SCORED and is not retrospectively declared passing. Its eight screenshots, eight checkpoints and screen recording are preserved under `artifacts/prime_hosted_20260910/harness054-prime-check-uiretry/`. Export and worker termination were verified. Wallet deductions observed during the check were USD 0.0791; cumulative deductions were USD 1.1863 at cleanup, subject to delayed settlement. GPT-4.1 attempt 3 was not launched. The approval covered one final check followed by a model run only if the check passed; a replacement check requires explicit permission within the remaining approved budget.\n\n\n## 2026-09-10 — Snapshot retry candidate; hosted validation still pending\n\nThe replacement unscored check on VM `fg4lqmsrkf5bz2ji3tt7l1ig` failed at action 1 because the read-only provider returned an empty response with `IllegalStateException: Read-only snapshot unavailable`. A subsequent captured state contained the complete requested search text. This proves a temporary evidence-read failure, not its underlying cause. Recorder commands were isolated in their own trace. Original screenshots, checkpoints, video and FAILED/INVALID verdict remain preserved under `artifacts/prime_hosted_20260910/harness054-prime-check-recording/`; export and sandbox termination were verified. No model call or benchmark attempt occurred.\n\n`amazon_cart_001/harness/peach.py` now retries only that exact empty, return-code-zero provider error within the existing five-read/60-second confirmation bound. Failed reads remain recorded; input is not replayed. `amazon_cart_001/verification/strict.py` validates the empty artifact hash, matching command stderr and unavailable-read metadata before allowing a later successful confirmation. Other errors still fail closed. `tests/test_snapshot_read_retry.py` covers recovery, exhaustion and forged error evidence. Both source and installed-wheel suites pass 260 tests each (the same suite in two installations, not 520 distinct tests). The candidate wheel is built in `dist054snapshot/` and installed in `install054snapshot/` in the scoped work directory. It has not passed a new hosted check or been published.\n\nThe latest observed wallet balance was USD 45.9301: USD 1.2454 below the original campaign baseline and USD 0.1382 below the USD 0.90 continuation approval baseline of 46.0683. These are observed deductions, subject to settlement. Approximately USD 0.7618 remains under that approval; another full preflight plus model reservation is USD 0.894912. Further spending requires a revised ceiling or scope. Actual completed model attempts remain two, both FAIL; attempt 3 is not launched. No success criteria or previous results were relaxed.\n\n\n\n## 2026-09-10 — Approved attempts 3–9 stopped by hosted preflight transport failure\n\nThe user approved USD 5.33 additional for one hosted check followed, only on success, by sequential GPT-4.1 attempts 3–9. This new approval's recorded wallet baseline is USD 45.9059; it must not be reset for subsequent launches. The preflight used VM `l9lxbn1vh2di2ewfjmx09igc`. Android booted in 481001 ms and the APK installation succeeded. Before any task checkpoint or action, Prime's Connect RPC returned an unavailable/body-read timeout. Full hosted validation therefore remains incomplete; the 260 local source and installed-package tests do not establish hosted readiness.\n\nInspection found a cleanup protocol defect: `integrations/hosted_session.py:call` advances its sequence only after receiving a response. After the ambiguous timeout, `finish` called `finalize` with the same sequence. `harness/hosted_worker.py:exchange` correctly rejected the different payload as a conflicting retry. Archive export consequently failed. A future repair must reconcile the existing request's response without replaying the action, then finalize with an unused sequence, and preserve diagnostics if reconciliation fails. The remote timeout's underlying network/service cause is not established by the receipt. No such transport repair has been implemented in this entry.\n\nTeardown was verified TERMINATED. Nine surviving startup/lifecycle receipt files were preserved in `artifacts/prime_hosted_20260910/harness054-prime-check-snapshot/`. There is no exported ZIP, task screenshot or video from this check; the exporter now records archive absence explicitly instead of implying successful media export. There were zero model calls, zero task actions and zero new benchmark attempts. Attempts 3–9 remain unsubmitted; the two historical model FAIL results remain unchanged. The candidate was not published.\n\nThe observed post-check wallet balance was USD 45.8597: USD 0.0462 deducted during this check and USD 1.3158 cumulatively against the original campaign baseline. Charges may settle later. Approximately USD 5.2838 remains under the new additional ceiling. The sequence stopped on infrastructure failure as approved; no replacement resource or model run was started.\n\n\n\n## 2026-09-10 — Short-poll mailbox and pending-request reconciliation\n\nThe user authorized fixing the transport recovery and retrying within the remaining USD 5.33 approval. `amazon_cart_001/harness/hosted_worker.py` adds an explicit `--poll` option: it publishes the immutable request once and returns a sequence-bound pending receipt immediately when no response exists. Repeated polls retain the existing conflicting-request guard and do not execute another app action. The default blocking exchange remains available for compatibility.\n\n`amazon_cart_001/integrations/hosted_session.py` records a pending request before upload, polls with individual RPC timeouts of at most 30 seconds under the existing total deadline, and records bounded transport retries. Another operation cannot overwrite that pending sequence. Cleanup first reconciles the existing response for at most 60 seconds, then finalizes at the next sequence; if reconciliation is impossible, it remains INVALID and still tears down the sandbox. This improves recoverable failures, but cannot guarantee archive retrieval when the remote service remains unavailable.\n\n`tests/test_mailbox_polling.py` adds eight offline cases: idempotent pending polls, conflicting payload rejection, temporary RPC recovery without reupload, bounded repeated failure, cleanup using the next sequence, unrecoverable pending cleanup, incorrect pending identity and deadline exhaustion. All 268 tests pass from source and the installed wheel. Candidate artifacts are in `dist054poll/` and `install054poll/` under the scoped work directory; no change to task requirements or scoring thresholds was made.\n\nThe replacement check `harness054-prime-check-poll` uses the original additional-approval wallet baseline USD 45.9059, not a fresh spending allowance. It is currently running on Prime VM `ilzdrpuq4ugkatjq809f6mre`. This entry is not evidence of a passing hosted workflow. Attempts 3–9 remain conditional on full hosted preflight success, package equality, cleanup and per-attempt remaining-budget checks.\n\n\n\n## 2026-09-10 — Hosted polling preflight passed; actual GPT-4.1 attempt 3 launched\n\nThe repaired Prime preflight completed all 15 reference actions and all six workflow stages. At the final checkpoint all six strict checks passed. Export produced 16 original screenshots, 16 checkpoints and `episodes/20260910T121209Z_c6ffa205/video/screen-recording.mp4`; 184 evidence/receipt files were preserved under `artifacts/prime_hosted_20260910/harness054-prime-check-poll/`. Sandbox termination was verified. This remains an unscored reference validation, not a successful model attempt. Observed deductions during this check were USD 0.0868, subject to settlement.\n\nVersion 0.5.4 was published to the existing private `devjangid-wootzapp/amazon-cart-001` environment. The actual native hosted GPT-4.1 evaluation for campaign attempt 3 was submitted as `jkraifceapqq70g3wj56jcoj` (https://app.primeintellect.ai/dashboard/evaluations/jkraifceapqq70g3wj56jcoj). Its worker is provisioning; this entry does not claim a model outcome. The user explicitly authorized proceeding to the task evaluation irrespective of the current reference result, but in this case the reference did pass. Source/package, pricing, previous-result, cleanup and budget guards remained active. The USD 5.33 approval baseline remains USD 45.9059 for the full sequential continuation; it is not reset per attempt.\n\n","encoding":"utf-8","truncated":false,"total_bytes":84354},"status":null}