{"data":{"kind":"file","path":"README.md","version_id":"u9k12crkcmkbl78pbb90bg6u","entry":{"name":"README.md","path":"README.md","is_directory":false,"size":5029,"modified_at":"2026-09-23T21:00:37.938000","content_hash":"d748633ffc8697a075068b2b80b2bbf69b1d0571b17926645cd770a19ce53e70"},"entries":[],"content":"# fintrading-eval\n\nA verifiers v1 Environments-Hub port of the **FinMem** LLM trading-agent\nbenchmark ([arXiv 2311.13743](https://arxiv.org/abs/2311.13743); MIT-licensed\nrepo [pipiku915/FinMem-LLM-StockTrading](https://github.com/pipiku915/FinMem-LLM-StockTrading)).\n\nThe policy under evaluation plays a FinMem-style single-stock trading agent:\none seat, one verifiers interaction, one JSON decision (`buy` / `sell` /\n`hold`, single share) per trading day, driven over **bundled, deterministic\nYahoo-Finance daily OHLCV text windows** that match the paper's test periods.\n\n> Note: the task brief referenced `arxiv.org/abs/2402.20185` for FinMem; that\n> arXiv id does not exist. FinMem is arXiv 2311.13743 (verified on arXiv and\n> in the official repo README).\n\n## Scenario inventory (4 cells x 3 personas = 12 tasks)\n\n| cell | ticker | window (paper / regime) | trading days | B&H log return |\n|---|---|---|---|---|\n| `tsla_main` | TSLA | 2022-10-06 .. 2023-04-10 (paper main test window) | 127 | -0.255 |\n| `tsla_ablation` | TSLA | 2022-06-16 .. 2022-12-28 (paper ablation test window) | 135 | -0.637 |\n| `spy_bear_2022` | SPY | 2022-01-03 .. 2022-10-14 (bear regime contrast) | 198 | -0.278 |\n| `nvda_bull_2023` | NVDA | 2023-01-03 .. 2023-06-30 (AI-rally regime contrast) | 124 | +1.084 |\n\nPersonas (FinMem profiling-module character settings): `risk_seeking`\n(aggressive, high-reward), `risk_averse` (conservative, lower-risk),\n`adaptive` (self-adaptive: risk-seeking while cumulative return is\nnon-negative, risk-averse after a drawdown — the paper's rule).\n\nEpisode length defaults to **50 decision steps** (bounded episodes for hosted\nevals); the `timesteps` taskset knob can extend episodes toward the paper's\nfull windows (up to `window_days - 1`).\n\n## Data\n\nBundled in the wheel at `fintrading_eval/data/{tsla,spy,nvda}.csv` (daily\nOHLCV, 2021-01-04 .. 2023-06-30, ~122 KB total; Yahoo Finance chart API,\nfetched at build time). Prices are the split/dividend-adjusted series —\n`adjclose` is the canonical price for all returns, matching the paper's\n\"daily adjusted closing price\". No network access at episode time.\n\n## Trading mechanics (paper-faithful, eq. 7)\n\nEach decision day i (at close): the seat sees recent OHLCV (last 7 days),\nmomentum stats (3d momentum; 5d/20d cumulative returns), its current\nsingle-share position, cumulative strategy return, and its last 5 decisions.\nDecision semantics: `buy` -> position +1 (long), `sell` -> position -1 (short,\nallowed even when flat), `hold` -> keep position. Position pos_i earns\n`pos_i * ln(p[i+1]/p[i])` on the adjusted close. Unparseable replies degrade\nto `hold` and count against the `discipline` metric.\n\n## Metrics\n\nPaper metrics (Cumulative Return, Sharpe, Annualized Volatility, Max\nDrawdown) are computed from the settled ledger and mapped to documented [0, 1]\ncomponents (higher = better); the composite is their unweighted mean:\n\n- `cum_return` = clip(0.5 + CR/2, 0, 1)\n- `sharpe` = clip((Sharpe_ann + 1)/3, 0, 1)  (rf = 0, sqrt(252) annualization)\n- `stability` = 1 - clip(max_drawdown, 0, 1)\n- `discipline` = valid_actions / decision_steps\n\nRewards are recorded on the seat trace: composite `fintrading_score` at\nweight 1.0, components `fintrading_*` at weight 0.0. All raw quantities plus\nper-step series (dates, prices, actions, positions, earned returns) land in\n`trace.info[\"fintrading\"]` — audit-reaggregatable offline. Guard counters\n(empty replies, self-heals, format nudges) are in\n`trace.info[\"fintrading_seat\"]`.\n\n## Known deviations from the paper\n\n- Price-only text mode: the paper's memory layers are populated from SEC\n  filings and ranked news (Alpaca/Refinitiv, embedding-backed); this port\n  replaces them with deterministic OHLCV context + the agent's own recent\n  decisions (same decision surface, no network/embedding dependency).\n- Bounded episodes: 50 decision steps by default instead of the full\n  124-198-day windows (knob to extend).\n- No Guardrails library: action validation is a strict JSON parser with\n  bounded format nudges (the paper's Guardrails text-validation role).\n- Buy-and-hold baseline recorded in raw metrics (not a reward component).\n\n## Local dev loop\n\n```bash\nuv venv --python 3.11 .venv\nuv pip install \"verifiers==0.3.2.dev118\" numpy\nuv pip install -e .\n.venv/bin/python -m pytest tests/ -q     # model-free stub tests\n# live local smoke (direct-model seat, 3 decision steps):\n.venv/bin/vf-eval fintrading_eval -m internal/glm-5.3-fast \\\n  --env.taskset.cells tsla_main --env.taskset.personas adaptive \\\n  --env.taskset.timesteps 3 -n 1 -r 1 --no-serve --no-rich \\\n  --env.agent.runtime.type subprocess --run-name fintrading-local-smoke\n```\n\n## Hub usage\n\n- Direct-model probes: `prime eval run primeintellect/fintrading-eval --hosted\n  -m internal/glm-5.3-fast -n <tasks> -r 1 --max-concurrent 1 --max-tokens\n  131072 --timeout-minutes 60 --env-args '{\"taskset\": {\"timesteps\": 2}}'`\n- PrimeAgentHarness seats: same command with\n  `\"agent\": {\"harness\": {\"id\": \"prime-agent\"}}` added to `--env-args`.\n","encoding":"utf-8","truncated":false,"total_bytes":5029},"status":null}