Every certified SPX strategy replayed on data with no lookahead, on two independent vendors. What the old numbers got wrong, what the book really earns, and every trade behind it.
Research thread, separate from the live bot. Repo usdemo; every figure reproduces from the named study file and PREREGISTERED.md.
For the main bot: exact rules, sealed results and live-bot flags are collected in usdemo/HANDOFF_MAIN_BOT.md (adopted #19/#20/#21/#23; do not deploy the throttle; meridian time-critical; wing fallback; headline-pilot).
Lesson of the night: the losses come from sudden jumps after entry, which nothing visible at entry predicts. The protections that work are structural — wings, split entries, defined-risk singles and position size — not filters.
The original Alpaca download (archive_opts.py, line 176) asked for strikes within ±4% of that day's closing price. A fly needs wings 3.5% away. When SPX finished far from the entry, the far wing was never downloaded, so the backtest could not build the fly and quietly skipped the day — exactly the days flies lose. Alpaca's data itself was fine; the download window was not.
| Strategy | Days wings existed | Per trade | Days wings missing | Per trade |
|---|---|---|---|---|
| eleven | 197 | +$913 | 349 | −$434 |
| fly0935 | 105 | +$1,570 | 439 | −$415 |
| meridian | 235 | +$633 | 311 | −$177 |
| pin | 277 | +$339 | 263 | −$188 |
Fix 1 — download. archive_opts_v2.py centres strikes on the open (known before any entry), ±7%, into data/opt_v2/. Databento quotes were centred on the open from the start.
Fix 2 — missing wing. Use the nearest real quoted strike (owner's rule); intrinsic value only as a last resort. An earlier $0 shortcut had booked a fly at −$38,744 when its maximum loss was $5,545.
Fix 3 — broken quotes. One-minute blowouts (bid $0.50 / ask $1,362.80 on 9 Apr 2025 13:27) are ignored: spread > 50% of mid on options above $2. 216 removed that day.
Fix 4 — physics. A fly's mark is clamped to between 0 and its wing width.
Vendors agree: ATM straddles, 783 checks — median gap 0.00%, 99% within 5%. Nine common strategies 2024–26: Alpaca v2 $410,169 vs Databento $436,064.
Cumulative P&L per 1-lot set, mid prices, every day a rule fires counted. Shading = distance below the running peak.
Each column is one calendar month across the years. Bottom rows: average and how many years that month was positive. With only 4–5 years per month, one bad month is not yet a pattern — a month is only worth skipping if it loses in most years and there is a reason (e.g. year-end holiday thin markets).
All strategies, including the two dropped from book totals (fly0935, hedgeflow). Max drawdown is each strategy on its own.
| Group | Strategies | Meaning |
|---|---|---|
| Holds | orbhold | positive, t ≥ 2 in both halves, profitable in April 2025 |
| Weak | meridian, fearsurplus, ivdrift, pin, momentum, termrich, magnet, defended, wall1030, eleven | positive in both halves, individually noisy (t < 2) or hurt in April 2025 |
| Fails | fly0935, hedgeflow, magnet1500 | negative in one half |
A strategy is in the core book if its profit is at least three times its own worst drawdown over its full honest record (Databento 2022–26 or Alpaca v2 2024–26). Nothing is removed or flagged — strategies under 3:1 are simply marked Below 3:1 in the table above.
Below 3:1 on the full record: eleven 2.4, momentum 2.99 (just under), magnet 1.6, wall1030 1.2, magnet1500 0.3, fly0935 and hedgeflow ≈ 0. Choosing on the full record uses hindsight; the no-hindsight version (chosen on 2022–24 only, judged on 2025–26) is in study66: 5 strategies, sealed $125.6k/yr, drawdown −$35.3k.
Because the core is much smoother than the full book, it can run at a larger size for the same pain (Databento 2022–26, singles in XSP):
Owner brief: improve the strategies without curve-fitting. Method: one knob at a time on 2022–24 only (wing width, stop, entry time ±5–30 min, exit time); a change is declared only if its neighbouring settings are also better, it helps in every train year, and it makes economic sense; then one test on the sealed 2025–26 half.
| Change (declared before the test) | Sealed P&L | Max drawdown | Win % | April 2025 | Verdict |
|---|---|---|---|---|---|
| #19 pin: split entry into 13:50 / 14:00 / 14:10 thirds | $33,986 → $48,345 | −$18,591 → −$13,492 | 64% → 64% | $8,980 → $13,363 | Pass |
| #20 fearsurplus: wings 4.5% instead of 3.5% | $44,772 → $46,330 | −$14,443 → −$13,388 | 68% → 68% | −$2,807 → −$1,604 | Pass |
| #20b fearsurplus: wings 4.5% + stop 1.5× | $44,772 → $26,622 | −$14,443 → −$11,254 | 68% → 57% | −$2,807 → −$3,477 | Fail |
| #21 termrich: stop 2.5× instead of 2.0× | $25,593 → $25,618 | −$12,392 → −$11,117 | 65% → 65% | $11,127 → $8,922 | Pass · marginal |
Replication on SPY and QQQ (2022–24 quotes): the pin split entry (#19) repeats on both — higher profit and 26–42% smaller drawdown; the protective wing on orbhold (#23) keeps ~80% of profit on the ETFs (98% on SPX) while cutting margin. On the sealed 2025–26 ETF data (corrected Alpaca download, finished overnight): the protective wing on orbhold replicates (keeps 91% on SPY, 85% on QQQ — and orbhold itself holds on both, t 2.2 / 1.9); the pin split entry helps on SPY but not QQQ, where pin itself is weak in 2025–26 — "partially replicated", kept and monitored.
Survivorship check (after the overnight-fly trap): every adopted change traded exactly the same days as the version it replaced, on both halves — the gains are not a counting artifact.
Meridian timing check (2025–26): 12:00 is the peak again ($61.2k; 11:55 $51.6k, 12:05 $52.4k, 11:45 $22.7k, 12:15 $40.1k) — the same smooth hump as 2022–24, so it is a real noon effect, not luck. It is time-critical: 5 minutes late costs ~15%, 10 minutes ~30%. Splitting its entry would hurt.
Intraday map (2022–24): a plain fly entered every 15 minutes earns best in two windows — 13:15–14:15 and around noon — and the book already trades both (pin's split entry and meridian). No uncovered sweet spot; the one candidate (a late-morning split) moves with the book (+0.7), so it would only add size.
No change for meridian (12:00 is its peak — moving it would be curve-fitting) or ivdrift (mixed by year). Across all strategies, wings of 4.5–5% were never worse than 3.5% and narrower wings were always worse; exits before 15:50 always cost money. SPY/QQQ replication of these changes is pending the overnight ETF download.
Legs, strikes and prices as recorded by the honest harness. SC/SP = short call/put (the body), WC/WP = long wings. Prices in index points; Gross, Costs, Net and Margin in dollars per lot. ivdrift rows carry P&L only.
Owner target: $240k a year on a $500k account, with the smallest drawdown possible. Book: the 7 core strategies with tonight's passing improvements (#19 pin split entry, #20 fearsurplus 4.5% wings, #21 termrich 2.5× stop) and the two single-leg strategies turned into defined-risk spreads (#23: a protective wing 2.5% further out). Databento 2022–26, 1 lot = 1× size.
Same $240k average, with the flies delta-hedged in SPY (passed its sealed test on SPX; replicated on SPY in both periods and QQQ 2022–24; failed QQQ 2025–26 — paper pilot recommended before real money). Buying power warning (measured): at the $240k size the SPY hedge peaks at ~2,100 shares on a typical hedged day (~$1.3M), ~6,600 on 1 day in 20 (~$4.3M) and ~10,500 on 7 Apr 2025 (~$5.8M) — more than a $500k account's ~$2M intraday buying power. Practical routes: hedge with ES/MES futures (~$250–300k margin at the peak; needs a futures broker) or a partial hedge (being tested).
Trade-off: the hedge trims some of the very best months, so a typical year is a little lower (median $213k vs $236k) while bad years are the same or better. For investors, the drawdown is the headline.
The 53 real months of 2022–26 reshuffled into 10,000 possible years, at 1.58× (the size that averaged $240k).
For the same $240k a year, adding eleven costs about $20k more drawdown and a worst month about $10k deeper (profit ÷ drawdown 13.0 → 10.4). eleven stays on the page, marked Below 3:1.
Why it now fits: the naked single-leg strategies needed ~$110k margin each. With a protective wing (#23, passed its sealed test: orbhold kept 98% of profit, defended 117%) they need ~$12–15k. The same size that needed over $460k of margin now needs about $206k.
How to trade 1.6×: whole SPX lots plus XSP (1/10 of SPX) for the fraction — e.g. 1 SPX + 6 XSP per strategy.
Honest limits: margin here assumes every trade of the day is open at once (upper bound). Mid prices; real fills cost a little. SPY/QQQ replication of tonight's changes is pending. Research, not advice.
Court law: explore on train only, declare before the sealed half, one shot, SPY + QQQ replication (t ≥ 2), book-increment gate, and — new this session — a filter counts only if the days it skips are worse out of sample.
| Candidate | Verdict | Why |
|---|---|---|
| #16 ivdrift · 12:05 fly when implied vol rose since the open | Certified | SPX t=3.67, SPY 3.22, QQQ 5.58 (old archive); honest: $46.7k Alpaca v2, $61.6k Databento |
| #14 quiettape · 11:35 fly on low 0DTE volume | Decertified | skipped days did as well or better |
| #15 unspent · 13:05 fly when the implied move is unused | Killed | QQQ t=1.66 |
| #17 spotvol · broken leverage effect | Killed | SPY t=1.96; deepened drawdown |
| #18 skewhold · veto on skew flattening | Killed | vetoed days did better out of sample |
| fearsurplus / hedgeflow filters | Flagged | filters do not separate out of sample; QQQ inverted |
| termrich filter | Confirmed | fired days ~2× skipped days on all three markets |
| Owner test · entry shifted −3…+3 min | Passed | every structure strategy profitable at every shift |
| #22 tightbook · 12:05 fly when ATM bid/ask spreads are tight (quote-native) | Killed | train $272/trade, t=2.85 — sealed 2025–26 inverted: −$110/trade, April 2025 −$15.3k |
| #19 pin split entry · #20 fearsurplus wings 4.5% · #21 termrich stop 2.5× | Passed | see Improvements; #20b (stop 1.5×) failed |
| #23 defined-risk singles (2.5% protective wing on orbhold / defended) | Passed | orbhold kept 98% of profit, defended 117%; margin ~$110k → ~$12–15k |
| Owner ideas: ORB-30 / ORB-60 / trend-day credit spreads | Killed | ORB-15 (= orbhold) holds; longer windows and trend-day spreads fail on 2022–24 |
| Owner idea: defend the threatened side (exit / re-center) | Killed | exit: less drawdown but 30–50% less profit; re-center: deeper drawdown. #24 fearsurplus re-center failed sealed |
| Owner idea: skip flies on racing / high-IV days | Killed | skipped trades earned +$121k on 2022–24 and drawdown did not fall — losses come from jumps after entry |
| #25 trendhedge · buy a debit spread with a strong midday trend | Killed | 2022–24 +$100/trade, hedge corr −0.41; sealed 2025–26 −$153/trade (t=−2.13), no help on bad days |
| #26 jumpstop · close flies when SPX jumps 0.4% in 2 minutes | Neutral | sealed profit and drawdown unchanged (ratio 5.2 = 5.2); saved $7.3k on 9 Apr 2025, but false alarms on reverting jumps cost as much. Recommendation: pair with a live news filter (paper pilot) |
| Owner idea: benefit from news — sell a fly after an intraday shock settles | Not declared | shocks are rare (~15/yr); best version +$494/trade, 67% win but t=1.32 on 39 trades; bigger shocks did worse |
| Re-opened on clean data: unspent, spotvol, skewhold | Kills confirmed | skipped days did as well or better (unspent: better overall but worse in 2022) |
| Richness throttle (½× cheap / 2× rich premium) and skip-cheap, re-tested clean | Killed | sealed 2025–26: rich days no better than cheap ($151 vs $140/trade); throttle doubled drawdown for the same profit — flat sizing is best |
| Owner idea: more trend-following in credit form (opening drive, gap-and-go, second break) | Killed | best: opening drive t=1.93 but moves with the book (+0.47) — same risk as orbhold; others weak. orbhold stays the trend strategy |
| Tail insurance: buy a far-OTM 0DTE strangle every day as a hedge | Killed | pays on the worst days but costs $4–39k a year; never beats simply trading a little smaller (profit ÷ drawdown 12.7 unhedged vs ≤ 12.65 hedged) |
| orbhold on QQQ as an extra strategy | Optional | fires on the same mornings as SPX orbhold (corr +0.7) — adds profit about like adding size; not a bad-day diversifier |
| Overnight 1DTE fly (sell 15:45, close next morning) — a different risk window | Killed | first run looked great (+$213/trade, t=5.8) but silently skipped gap nights — a survivorship trap caught before declaring; honest version +$54/trade, t=1.3, 2022 negative |
| #27 iron condor at 10:00 (±0.75%, wings 1% beyond) — re-tested on clean data | Killed | 2022–24 $44/trade, every year +; sealed 2025–26 $21/trade (t=0.6), April 2025 −$11.4k, profit ÷ drawdown 0.6 |
| Forum / TradingView ideas: Bollinger break, MA-trend put spread, Ichimoku filter | Killed | BB break $54/trade (t=1.3); MA put spread wins 86% but nets $15/trade — high win rate ≠ edge; Ichimoku adds nothing |
| Owner idea #29: stop-and-reverse — at a fly's stop, buy 2 ATM options with the move | Failed | train: profit ÷ DD 7.7 → 12.4. Sealed: rescued the crash (9 Apr 2025 −$9.7k → +$36.2k) but ordinary stop-outs cost more (ratio 5.2 → 4.7) |
| Owner idea #28: breach → close the fly + credit spread against the move | Failed | sealed worst-20-days −$164k → −$32k, but profit fell to 56% of the book |
| Breach → KEEP the fly + credit spread (seen only after the sealed look) | Failed SPY/QQQ court | SPX 2025–26 looked good (ratio 5.2 → 6.3) but was seen with hindsight. Independent test: worst-20-day losses fell in all four SPY/QQQ cells (−14% to −28%), but profit ÷ drawdown fell on QQQ — not certified; a possible SPX paper pilot if worst-day protection matters more than profit |
| Dealer gamma (GEX) from pre-open open interest — the prop-desk sensor | Not declared | high-GEX days best for flies ($691/day) but U-shaped (middle worst) and high-GEX lost in 2023 — the simple GEX estimate is too crude to filter on |
| #31 delta hedge — keep the flies; once SPX is 0.3% from the body, hedge their delta with SPY every 5 min (cost $5 per delta unit) | Passed | sealed 2025–26: profit +5%, drawdown −44% (−$39.5k → −$22.1k), worst 20 days −$164k → +$37k, 9 Apr 2025 −$9.7k → +$74.6k. Replication: SPY both periods and QQQ 2022–24 pass (worst-20-day losses cut 75–85%); QQQ 2025–26 fails (trade-print data → noisy deltas suspected). 5 of 6 tests — recommend an SPX paper pilot. Hedge can reach ~2–3k SPY shares |
| Lookahead in old archive | Fixed | open-centred re-download + honest harness |