{"data":{"kind":"file","path":"README.md","version_id":"saer7vor3vwyfp19ki5jtcxn","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":2802,"modified_at":"2026-09-30T13:56:59.024000","content_hash":"8638c96692da92fcb7371fdb54d68bfe4c6203eea6dc25a34f3f046f23965ecd"},"entries":[],"content":"# Crypto Trading Decision (v1.0.1)\n\nA verifiers environment for risk-managed crypto trading decisions. The model\nreceives a synthetic BTC/USDT daily market snapshot — an OHLCV window plus\nprecomputed technical indicators (SMA-10/20, RSI-14, 5-day momentum, realized\nvolatility, volume ratio) — and must return a single JSON trade plan:\nlong / short / hold with entry, stop, target, position size, key risks, and a\ngrounded rationale.\n\nScenarios come from a seeded regime-switching random walk (uptrend, downtrend,\nchoppy). Regimes carry real drift, so the visible indicators are genuinely\ninformative about the hidden 5-day future window used as ground truth.\n\n## Version\n\n1.0.1\n\n## Output schema\n\n```json\n{\n  \"action\": \"long\",\n  \"confidence\": 0.72,\n  \"entry\": 61250.00,\n  \"stop_loss\": 59900.00,\n  \"take_profit\": 64800.00,\n  \"position_size_pct\": 2.0,\n  \"key_risks\": [\"break of sma_20 invalidates trend continuation\", \"elevated volume ratio signals distribution\"],\n  \"rationale\": \"Price 61250 sits above sma_10 60120 and sma_20 59400 with rsi_14 at 58.2 ...\"\n}\n```\n\n## Reward functions\n\n| Function | Weight | What it measures |\n|---|---|---|\n| `format_validity` | 0.20 | Required keys, valid action enum, confidence in [0,1], numeric fields, non-empty risks/rationale (5 sub-checks) |\n| `direction_accuracy` | 0.30 | Action matches the realized future move; partial credit for holding through a sub-1.5% move |\n| `risk_management` | 0.30 | Bracket structure correct for the side, reward-to-risk >= 1.5, sizing 0.5-5%, stop distance contained and proportional to realized volatility (5 sub-checks) |\n| `reasoning_and_calibration` | 0.20 | Indicator citations, concrete numbers from the snapshot, risk vocabulary, invalidation/conditional logic, anti-repetition and calibrated length, confidence-evidence consistency |\n\nEvery reward function guards against empty or malformed completions and scores\nonly within [0.0, 1.0].\n\n## Load\n\n```python\nimport verifiers as vf\nenv = vf.load_environment(\"crypto-trading-decision\")\n# or locally\nfrom crypto_trading_decision import load_environment\nenv = load_environment(num_scenarios=36, seed=20260930)\n```\n\n`num_scenarios` and `seed` are optional; defaults are deterministic.\n\n## Design notes\n\n- Ground truth is the realized 5-day forward return of the hidden tail of each\n  price series (threshold +/-1.2% for long/short, otherwise hold).\n- Indicators are computed on the visible history only; the future window never\n  appears in the prompt.\n- Reward steps are graded or sigmoid-smoothed where thresholds could otherwise\n  become discrete exploitation surfaces.\n- The rationale metric requires multiple independent signals (indicator\n  grounding + numbers + risk language + conditionality + non-repetition), so a\n  single templated sentence cannot max it out.\n","encoding":"utf-8","truncated":false,"total_bytes":2802},"status":null}