{"data":{"kind":"file","path":"README.md","version_id":"sfn4n8mi8hz2zo0fjnu7ohmy","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":6968,"modified_at":"2026-10-02T13:37:38.708000","content_hash":"c0d08039d460db050d9b13165a7cb7779d440986bc1171f10fb99bb201ffb30a"},"entries":[],"content":"# BTC Trading Agent — V0.0.1\n\nA multi-turn BTC/USD spot-trading environment for training and evaluating agents\nunder causal market observations, adverse execution costs and verified accounting.\nIt does not place real orders or require an exchange account.\n\n## Quick start\n\n```python\nfrom btc_trading_agent import load_environment\n\nenv = load_environment()  # bundled historical data; no network downloads\nstress_env = load_environment(source=\"synthetic\", num_examples=8)\n```\n\nPublished identifier: `randall/btc-trading-agent`; package version: `0.0.1`\n(display label V0.0.1, using PEP 440 numeric version in package metadata).\nPython 3.11–3.13, `verifiers==0.1.14`, `datasets>=3.0,<5`.\n\n```sh\nprime env install randall/btc-trading-agent@0.0.1\nprime eval run randall/btc-trading-agent -m YOUR_MODEL --env-args '{\"num_examples\":4}'\n```\n\nThe exact model/provider configuration is supplied by your evaluation runner.\nNo paid training or live-model inference is launched automatically.\n\n## Market data and splits\n\nThe default environment bundles 1,096 actual daily Coinbase Exchange BTC-USD\nOHLCV candles from 2023-01-01 through 2025-12-31 UTC. Data are fetched from the\npublic `/products/BTC-USD/candles` endpoint and validated for daily continuity.\nSee `data/provenance.json` for endpoint, retrieval timestamp and SHA-256.\n\nTraining uses only 2023–2024 observations; evaluation uses only 2025 observations,\nincluding warmup. No bars overlap across splits. With defaults, training has 87\nrolling episodes (stride 8), evaluation has 14 episodes (stride 24). Evaluation\ntrade intervals do not overlap; their warmup windows may overlap. Training windows\nintentionally overlap; these are not statistically independent samples.\n\nEach episode has 16 initial observed closes and 24 decisions. A decision uses the\nlast observed close; exactly one next close is then revealed. Full replay paths\nare stored in task `info` only, not in model-visible prompts or responses. A\ntrusted runner must never serialize private `info` to the policy. As with any\npublished historical benchmark, agents with memorized history or external data\naccess can undermine the causal assumption; this is not a hidden-data benchmark.\n\nOptional `source=\"synthetic\"` generates 96 train / 48 eval scenarios across bull,\nbear, range, high-volatility, crash and reversal regimes. Generation uses local\nRNG instances and disjoint train/eval seeds. Regime and seed labels are not shown\nto the model. Synthetic results are explicitly distinguished from historical data.\n\n## Action and accounting\n\nReturn exactly `{\"target\":0.5}`: the requested BTC mid-value as a fraction of\npre-trade marked equity. Allowed range is [0,1]. No leverage, short selling,\nadditional fields, markdown, duplicate keys, booleans, NaN or infinite values.\nThe request is a pre-cost target, so realized exposure can differ slightly after\ncosts; purchase size is capped by available cash.\n\nStart with USD 10,000 and zero BTC. Buy fills are `mid * (1 + cost_bps/10000)`;\nsell fills are `mid * (1 - cost_bps/10000)`. Default one-way cost is 15 bps,\nrepresenting combined fee/spread/slippage. Orders fill immediately at the observed\nmid plus this adverse cost: an explicit simplified execution assumption, not a\nclaim that daily bars support real-time fills. No volume constraint or market impact\nis modeled. At the final close, all BTC is liquidated and its costs are included.\n\nInvalid actions leave positions unchanged but advance time. They are not free\nretries and cannot avoid terminal settlement. Every action is processed during\ntrajectory append, so the last decision is applied before max-turn termination.\nLedger and episode state are isolated per rollout.\n\nObservations: trailing close_history, remaining decisions, cash_usd, btc,\nequity_usd, realized marked drawdown, disclosed one_way_cost_bps, action errors.\nNo future OHLCV, regime, seed, or full path is exposed.\n\n## Reward\n\nLet `r = final_equity / initial_equity - 1`, `d = max_drawdown`, and\n`t = cumulative traded mid notional / initial_equity`, including liquidation.\n\n```\nutility = r - 0.5*d - 0.0005*t\nquality = sigmoid(clip(utility / 0.04, -30, 30))\nreward = quality * (1 - invalid_actions / decisions)\n```\n\nIncomplete episodes receive exactly zero reward. All reward values are finite\nand in [0,1]. Scores depend exclusively on the executed ledger, not reasoning,\nclaimed profits, reward strings, or an LLM judge. Staying in cash earns 0.5:\ncash is a valid risk-control policy, not an error. Reward coefficients are fixed\nand visible, not claimed to be optimal for every risk preference.\n\nReporting-only metrics (weight zero): net_return (can be negative), max_drawdown,\nturnover, invalid_action_rate, completion_rate, plus the runner's num_turns.\nA high reward is not evidence of real-world profitability.\n\n## Configuration\n\n- source: historical (default) or synthetic\n- history: integer 2–128, default 16\n- horizon: integer 2–128, default 24\n- cost_bps: finite 0–100, default 15\n- num_examples: positive integer or None; cannot exceed either historical split\n- seed: integer [0, 2**32), default 42, used for synthetic generation\n\nmax_turns is derived from horizon; overriding it is rejected. Additional runner\nkwargs such as timeout_seconds are passed to the multi-turn base environment.\n\n## Verification and baselines\n\n```sh\nuv pip install -e '.[dev]'\npython -m pytest tests -q\nruff check btc_trading_agent tests scripts\npython scripts/baselines.py\n```\n\nTests cover strict parsing, no-negative-balance accounting, analytical buy-and-hold,\nfinal liquidation, causal observations, temporal splitting, invalid actions,\ncrash drawdown, cash versus costly churning, and the actual verifiers rollout loop\nwith an explicitly deterministic test client. This client is NOT a language model.\n\n`reports/baselines.json` contains real executed non-LLM baseline measurements on\nboth holdouts: cash, buy-and-hold, causal momentum, and volatility guard. Each\nhistorical policy has 14 episodes and each synthetic policy has 48 episodes.\nNo learned agent evaluation or live-model profitability is claimed. The simple\ncash policy outperforms these baseline policies under this risk-adjusted reward\non these particular holdouts. Inspect per-metric results rather than reporting\ncash as a profitable trading strategy.\n\n## Limitations\n\nHistorical replay is hindsight data, not an unseen live stream. Daily-close\nexecution, fixed costs, no orderbook, no leverage, and no market impact simplify\nreality. Historical evaluations include a modest number of overlapping-context\nwindows. Regime generators are stylized, not calibrated market models. Future\nversions can add unseen periods, richer fills, and calibrated risk preferences;\nV0.0.1 does not promise Sharpe optimization, market-neutral returns or live trading.\n\nCode: MIT. Coinbase market observations retain their source attribution and are\nnot asserted to be relicensed under MIT. See LICENSE and data/provenance.json.\n","encoding":"utf-8","truncated":false,"total_bytes":6968},"status":null}