Retail Option-Income Strategies Under Honest Frictions: A Walk-Forward, Selection-Bias-Corrected Study of the Index Volatility Risk Premium at Survivable Scale, 2010–2023

Johnathan Carr Independent Researcher — nursebuilds.com

Working Paper, Version 2.1 — July 2026

Keywords: volatility risk premium; option writing; covered calls; backtest overfitting; Deflated Sharpe Ratio; household finance; retail investors; walk-forward analysis; time-series momentum

JEL classification: G11 (Portfolio Choice), G13 (Contingent Pricing; Futures Pricing), G14 (Information and Market Efficiency), G51 (Household Saving, Borrowing, Debt, and Wealth)

Disclaimer: This is an independent research working paper. It is not investment advice, and no result herein is a recommendation to buy or sell any security. Data and code availability: end-of-day option chains were licensed from OptionsDX; the backtesting engine, validation suite, strategy implementations, and logged run artifacts (including random seeds) are maintained in a private repository and are available from the author on reasonable request.


Abstract

Option-premium-selling strategies are marketed to retail investors as reliable income, even as broker disclosures indicate most retail options accounts lose money. We evaluate whether any mainstream option-income structure can fund a retirement-income mandate (~6–9% after tax) while preserving capital through crashes, using end-of-day SPY option chains (2010–2023) and a QQQ replication (2012–2023). We test 134 strategy configurations — cash-secured puts, credit spreads, iron condors, covered calls, short strangles, jade lizards, covered strangles, weekly-expiry variants, a time-series momentum overlay, and carry+trend combination portfolios — under deliberately conservative retail frictions: fills at the full quoted bid-ask spread, commissions and fees, next-bar execution, and time-varying T-bill collateral yield. Inference is strictly walk-forward and corrected for selection bias with the Deflated Sharpe Ratio (Bailey and López de Prado, 2014) and the Probability of Backtest Overfitting (Bailey, Borwein, López de Prado, and Zhu, 2017), supplemented by regime-split and circular block-bootstrap checks. No configuration survives validation on either underlying (best DSR 0.90 against a 0.95 bar). Return attribution in the framework of Israelov and Nielsen (2015) assigns 87–99% of the premium-selling family's returns to passive equity exposure. Our central measurement uses a delta-neutral, risk-managed short strangle — attribution- verified as a nearly pure volatility-risk-premium harvester — to estimate the index VRP capturable at crash-survivable position size: approximately 1% of capital per year, an order of magnitude below the income the strategies are marketed to produce. Sizing, not signal, is the binding constraint, consistent with Santa-Clara and Saretto (2009). Three extensions close the standard objections: conditioning entries on implied-minus-realized volatility richness improves the weakest configurations but degrades the best one and validates nothing; an after-tax analysis across wrappers shows that Section 1256 index treatment, while real, cannot close a two-to-one gross gap to a live income fund in a high-tax state; and cash-rate scenarios show the collateral-yield lever is symmetric across all cash-holding alternatives. Because the paper's claim rests on a measured mechanism — beta in disguise, plus a residual premium too thin to survive on — rather than on exhaustive coverage, the negative result generalizes beyond the structures tested. We document three preliminary positive results overturned by the study's own machinery, and argue the reproducible harness — not any strategy — is the transferable contribution.


1. Introduction

A large retail education industry teaches option-premium selling — cash-secured puts, covered calls, credit spreads, iron condors, strangles — as a dependable income stream, with practitioner-quoted yields of 18–30% per year circulating widely. Regulatory broker disclosures, by contrast, indicate that a substantial majority of retail options accounts lose money, and the academic record on individual investors' option trading is bleak: Bauer, Cosemans, and Eichholtz (2009) find that most retail option traders in their sample lose money, with losses driven by gambling-like turnover; Barber and Odean (2000) and Barber, Lee, Liu, and Odean (2009) document the corrosive effect of retail trading costs in equities generally; and Bryzgalova, Pavlova, and Sikorskaya (2023) show the post-2019 retail options boom concentrated in exactly the short-dated contracts with the worst expected returns.

Three mechanisms can reconcile marketed yields with realized losses: (i) yields quoted on deployed rather than total capital; (ii) equity beta earned during bull markets and misattributed to strategy skill; and (iii) survivorship in who reports. All three are testable with sufficient care. This paper tests them.

The research question is deliberately practical, phrased as a household would phrase it: after real retail frictions, honestly measured, does any common option-income structure deliver validated risk-adjusted edge — or at minimum meet an income target of roughly 6–9% after tax while losing materially less than the market in crashes? The mandate's dual objective reflects sequence-of-returns risk, the central hazard of retirement decumulation (Bengen, 1994): for a retiree, a strategy's behavior inside crashes is not a robustness check; it is the objective function.

The design principle throughout is *anti-self-deception*. Every modeling choice that typically flatters a backtest is inverted: fills pay the full quoted spread; execution is next-bar; collateral earns the historical Treasury-bill rate rather than a smooth assumption; all reported performance is out-of-sample under a rolling walk-forward protocol (Pardo, 2008); the number of strategies tried is penalized explicitly (Bailey and López de Prado, 2014; Bailey et al., 2017), in the spirit of the data-snooping literature (White, 2000; Sullivan, Timmermann, and White, 1999) and of Harvey, Liu, and Zhu's (2016) argument that discovered financial "factors" should clear a substantially raised significance bar; and every return stream is decomposed (Israelov and Nielsen, 2015) so that market exposure cannot masquerade as edge. Where the machinery produced encouraging preliminary results, we attacked them; Section 8 documents three that did not survive.

A note on the shape of the claim. A negative result over any finite strategy grid invites the reply "you didn't test my variant." This paper's claim is therefore not exhaustive coverage but a measured mechanism (Section 7): attribution shows the premium-selling family's return is overwhelmingly passive equity exposure; a frictionless re-run shows the deficit is edge, not cost; and a delta-neutral, crash-surviving harvester puts the residual volatility premium at roughly 1% of capital per year at retail-survivable size. Any index-option income structure is a portfolio of those measured ingredients, so the conclusion generalizes to variants sharing them — which is what makes a finite grid sufficient.

Contributions. We offer four:

  1. Selection-bias-corrected evidence at retail frictions. The academic option-writing literature evaluates index-level strategies at institutional cost assumptions (Whaley, 2002; Ungar and Moran, 2009) or uses frictionless decompositions (Israelov and Nielsen, 2015). The retail-investor literature documents realized losses but not whether a *disciplined* retail implementation could have succeeded. We bridge the two: a 134-configuration grid under conservative retail frictions, with inference explicitly corrected for the size of the search.
  1. A measurement of the index VRP at survivable scale. Rather than asking whether a volatility risk premium exists — it does, on average (Coval and Shumway, 2001; Bakshi and Kapadia, 2003; Carr and Wu, 2009; Bondarenko, 2014) — we ask what fraction of it a capital-constrained household can actually harvest once position size is set so the strategy survives 2008-scale moves. Our attribution-verified cleanest harvester nets ≈1% of capital per year. The binding constraint is sizing, not signal, extending Santa-Clara and Saretto (2009) from margin requirements to a retiree's survival constraint.
  1. Bear-window evaluation of undefined-risk income structures. We provide, to our knowledge, the first walk-forward accounting of short strangles, jade lizards, and covered strangles — structures ubiquitous in retail education and absent from the academic record — across four real drawdown episodes, including the March 2020 crash.
  1. A reproducible harness and an audit trail. All results regenerate from committed code and logged artifacts (seeds included), and we report the cases in which the pipeline overturned its own preliminary findings — the empirical value of treating one's own results adversarially.

The remainder of the paper proceeds as follows. Section 2 situates the study in four literatures. Section 3 describes the data. Section 4 details the methodology and the rationale for each conservative choice. Section 5 reports results, including the conditional-richness extension. Section 6 evaluates taxes, wrappers, and the rate environment. Section 7 states the mechanism and the generalization argument. Section 8 documents robustness and the audit trail. Section 9 states limitations, including a registered directional follow-up on live fill quality. Section 10 discusses implications, and Section 11 concludes. Appendix A reproduces the validation mathematics; Appendix B enumerates the strategy grid.

2. Related Literature

2.1 The volatility risk premium

That index options are, on average, priced above their subsequent realized payoffs is among the better-established facts in empirical asset pricing. Coval and Shumway (2001) document strongly negative expected returns to buying index options; Bakshi and Kapadia (2003) isolate a negative market volatility risk premium in delta-hedged option gains; Carr and Wu (2009) formalize and measure variance risk premia across indexes and single names, finding the index premium large and robust; Bondarenko (2014) shows index puts appear expensive relative to any plausible equilibrium model. Ilmanen (2012) frames the premium as the compensation flowing from insurance buyers to insurance sellers, warning that the seller's return distribution is short left-tail by construction.

Critically for our question, the existence of an *average* premium does not imply an *implementable retail income stream*: the measured premia are gross of transaction costs, assume continuous delta-hedging or institutional execution, and are earned by bearing exactly the crash risk a retiree cannot bear at scale. Where this paper sits: we take the premium's existence as given and measure what remains of it after (i) retail spreads, (ii) discrete daily management, and (iii) a position size constrained to survive the sample's worst crash.

2.2 Option-writing strategy performance

Whaley (2002) introduced and evaluated the CBOE BuyWrite index (BXM), finding equity-like returns at reduced volatility over 1988–2001; Ungar and Moran (2009) report similar risk-adjusted improvement for the cash-secured put-write index (PUT). Both studies predate the 2010s regime our sample covers, and both evaluate index methodologies without retail frictions. Israelov and Nielsen (2015) provide the decisive decomposition: a covered call is passive equity exposure plus a short-volatility position plus an uncompensated equity-timing residual, and most of its realized return is simply the equity exposure. Israelov (2019) extends the skepticism to protective puts, showing option-based downside protection is systematically overpriced relative to simple de-risking.

Closest in spirit to our sizing result, Santa-Clara and Saretto (2009) show that apparently attractive option-writing returns shrink drastically once margin requirements and the possibility of forced deleveraging are respected — the "good deals" are unavailable at implementable size. Where this paper sits: we replicate Israelov and Nielsen's attribution verdict out-of-sample at retail frictions, extend the strategy set to undefined-risk structures the academic literature has not covered, and sharpen Santa-Clara and Saretto's margin constraint into an explicit survival constraint calibrated to the retiree mandate.

2.3 Retail investor performance

Barber and Odean (2000) established that retail trading intensity predicts underperformance roughly one-for-one with costs; Barber, Lee, Liu, and Odean (2009), with complete Taiwanese market data, put aggregate individual-investor trading losses at about 2% of GDP annually. In options specifically, Bauer, Cosemans, and Eichholtz (2009) find the majority of Dutch retail option traders lose money, with the worst outcomes among the most active; Bryzgalova, Pavlova, and Sikorskaya (2023) show that the 2019–2021 retail options surge concentrated in short-dated, high-embedded-leverage contracts with poor expected returns, and that payment-for-order-flow intermediation captured a substantial share of the flow's value. Where this paper sits: this literature documents *that* retail options traders lose; ours asks whether the loss is an execution failure (fixable with discipline) or a structural property of the opportunity set at retail scale. Our answer — the premium itself is too thin at survivable size — locates the problem in the opportunity set, not merely the execution.

2.4 Backtest validity, multiple testing, and the factor context

Any grid search over strategies is a multiple-testing exercise, and the finance literature's reckoning with that fact guides our inference. White (2000) and Sullivan, Timmermann, and White (1999) established the data-snooping critique for trading rules; Harvey, Liu, and Zhu (2016) argue that, given the hundreds of factors mined from overlapping data, newly claimed discoveries should clear a t-statistic near 3 rather than 2; Harvey and Liu (2015) translate the argument to practitioner backtests. Bailey and López de Prado (2014) supply the instrument we adopt — the Deflated Sharpe Ratio, which tests an observed Sharpe ratio against the expected maximum across N trials under the null — and Bailey, Borwein, López de Prado, and Zhu (2017) supply the complementary run-level diagnostic, the Probability of Backtest Overfitting via combinatorially symmetric cross-validation. López de Prado (2018) collects the associated practice. Our walk-forward protocol follows the practitioner tradition codified by Pardo (2008); our bootstrap follows Politis and Romano (1992).

For the diversifying overlay we test, the factor literature is the motivation: time-series momentum is documented across a century and dozens of markets (Moskowitz, Ooi, and Pedersen, 2012; Hurst, Ooi, and Pedersen, 2017), its crisis-period convexity being precisely the property a premium-selling book lacks; carry and momentum are natural complements across asset classes (Koijen, Moskowitz, Pedersen, and Vrugt, 2018; Asness, Moskowitz, and Pedersen, 2013). Where this paper sits: we apply the DSR/PBO apparatus — built for institutional factor research — to the retail option-income question, and we impose the trial-count penalty on our *own* search honestly, including the widening of the grid mid-study (Section 5.5).

3. Data

Option chains. End-of-day SPY option quotes, 2010-01-04 through 2023-12-29 (3,508 trading days), and QQQ, 2012–2023, licensed from OptionsDX. We retain contracts with 0–55 days to expiry (≈19 expiries per trading day), targeting ≈35-DTE monthly entries and ≈10-DTE entries for weekly variants. Each trading day is validated before reaching any strategy (monotonicity and crossed-quote checks, delta/IV presence, positive spreads); deltas and implied volatilities are present on 100% of retained rows. SLV and AAPL chains (2016–2023) were imported for planned breadth work and are not used for inference here.

Data quality. Median at-the-money relative quoted spreads by era: SPY ≈1.3–3.0%, QQQ ≈1.2–3.8% (index-grade liquidity); AAPL ≈2.1–3.8%; SLV ≈2.1–5.5% with a 90th percentile reaching 20% (excluded from inference; retained as a live test case for spread-based exclusion rules). The sample brackets four distinct drawdown episodes — the 2015–16 correction, 2018-Q4, the February–March 2020 COVID crash, and the 2022 bear market — which serve as the mandate's principal evaluation windows.

Benchmarks. CBOE strategy benchmark indexes (BXM, BXMD, PUT, CNDR; total-return methodology) and JEPI (distribution-adjusted daily series) for the build-versus-buy comparison; total-return buy-and-hold of the underlying (dividends reinvested) and a Treasury-bill cash floor are computed inside the engine and included in every run as baselines that, by construction, can never be declared "survivors" (Section 4.3).

4. Methodology

4.1 Engine and friction model, with rationale

A deterministic backtest engine (pure Python standard library) holds strategy logic pure — strategies see a market snapshot and their own positions and emit intended orders — with every order passing an independent risk gatekeeper before simulated execution. Frictions, all active in every result reported:

Fills pay the full quoted spread per side (sells execute at the bid, buys at the ask, from that day's actual quotes). *Rationale.* Muravyev and Pearson (2020) show that effective option spreads are substantially inside quoted spreads for execution-timing-sophisticated traders; we adopt the full quoted spread deliberately as the conservative bound for a retail household that cannot time executions intraday from end-of-day data. This choice biases *against* finding edge, which is the correct direction for a study whose conclusion may be acted on by households: an edge that exists only inside the spread is not a retail edge. We verified the bias is not decisive: in a frictionless re-run (spread cost set to zero), the premium-selling book still lost money over 2017–2023 (Section 8), so the headline result is not an artifact of pessimistic fills.

Next-bar execution. Decisions formed on day *t*'s close fill at day *t+1*'s market, eliminating same-bar look-ahead. *Rationale:* same-bar fills embed a small but systematic optimism (the decision uses the same price it transacts at); switching to next-bar cost ≈0.13 DSR in our robustness pass (Section 8), confirming the bias was material.

Time-varying collateral yield. Idle cash and option collateral accrue the historical 3-month T-bill rate; long stock receives dividends; short stock pays borrow. *Rationale:* a flat assumed rate smooths the cash-flow stream and mechanically inflates Sharpe ratios (a further ≈0.13 DSR in the same pass); put-write economics in particular are materially collateral-yield economics (Ungar and Moran, 2009), so the rate path must be the historical one.

Commissions and fees per a retail broker preset (per-contract commission plus exchange and regulatory fees), split across legs; expiry settlement and daily mark-to-market throughout. Assignment is cash-settled (Section 9).

4.2 Walk-forward protocol

Rolling windows of 378 training days and 126 out-of-sample days (step 126, 252-day burn-in) yield 22 out-of-sample segments over 2010–2023; only stitched OOS performance (≈2,750 daily observations) is scored. In-sample data is used solely for selection, never for reported performance (Pardo, 2008). The walk-forward stitch, rather than a single train/test split, ensures every sub-era of the sample serves as test data exactly once and prevents the selective reporting of a favorable test window.

4.3 Selection-bias-corrected inference, with rationale

The primary grid contains N = 134 active strategy configurations per underlying (Appendix B); a post-verdict extension adds six conditional- richness rungs (N = 140; Section 5.8). N is a *ledger*, not a footnote: it includes every configuration evaluated, including those added mid-study, and including the four composite-weight variants that penalize our own weight search. Under the Deflated Sharpe Ratio (Bailey and López de Prado, 2014), each configuration's non-annualized OOS Sharpe ratio is tested against the expected maximum Sharpe across N trials under the null, with skewness and kurtosis corrections (Appendix A). Survival requires DSR ≥ 0.95 — the same order of stringency as Harvey, Liu, and Zhu's (2016) raised t-bar, applied at the strategy level. The penalty binds mechanically: when the grid grew from N = 109 to N = 134, the best pre-existing configuration's DSR fell from 0.43 to 0.19 with its returns unchanged (Table 3) — the "trial-count tax" working as designed.

Because DSR is per-configuration, we add the run-level Probability of Backtest Overfitting (Bailey et al., 2017): the T×N matrix of aligned daily OOS returns is partitioned into S = 10 contiguous blocks, and across all half/half block partitions we record how often the in-sample-best configuration falls at or below the out-of-sample median. Survival requires PBO ≤ 0.50. Final runs: PBO = 0.079 (SPY), 0.198 (QQQ) — the *ranking* of configurations is stable out-of-sample; what fails is the absolute level of performance.

Three further gates, each motivated by a failure mode we observed or anticipated: (i) regime split — a survivor must be positive in every sufficiently-sampled ATM-implied-volatility regime band (calm/normal/elevated/crisis), preventing a calm-period specialist from passing on averages; (ii) circular block-bootstrap confidence intervals (Politis and Romano, 1992; B = 1,000, block length ≈ T^⅓, seed logged), preserving the volatility clustering that i.i.d. resampling would erase; (iii) degenerate-curve exclusions — baselines can never be survivors, and neither can any configuration with zero fills in the scored window, a gate added after it caught ten "perfect" survivors whose sizing rule had silently refused every trade, leaving pure collateral-interest curves with near-infinite Sharpe ratios (Section 8). Sample-size guards (≥20 trials, ≥252 OOS observations) and an acceptance test — the machinery must flag seeded pure-noise inputs and a deliberately overfit configuration — complete the apparatus.

4.4 Return attribution

Following Israelov and Nielsen (2015), each strategy's daily option-book P&L is decomposed exactly — the identity is enforced per day — into passive equity (mean net delta × underlying move), dynamic timing (demeaned delta × move), and short volatility (the delta-hedged residual: the volatility risk premium actually harvested). Attribution is computed on gross mark-to-market so that the cost decomposition remains separate. The decomposition's purpose is disciplinary: any strategy whose "income" loads on the passive term is repackaged equity beta, whatever its marketing.

4.5 The strategy universe

Appendix B enumerates the grid; in outline: cash-secured puts (target-delta × profit-target × entry-filter presets), bull put spreads, iron condors, covered calls (bare, delta-managed, and weekly ≈10-DTE), an undefined-risk naked put-write (cautionary rung), short strangles (bare / risk-managed / weekly; symmetric deltas 0.16/0.20/0.30), jade lizards (short put plus short call spread, entered only when total credit ≥ wing width, eliminating upside risk by construction), covered strangles, eight time-series momentum rungs (1- and 12-month lookbacks × long-flat/long-short × volatility scaling; Moskowitz, Ooi, and Pedersen, 2012), and four carry+trend composite books across a capital- weight grid (Koijen et al., 2018). Management overlays: profit targets (50% / 75% / hold-to-expiry), a 2×-credit stop-loss, a 21-DTE time exit, fixed- fractional risk sizing against a 15%-overnight-crash proxy, and re-entry cooldowns. Risk-managed naked-strangle rungs run on a dedicated $150,000 notional account because honest sizing refuses entries below ≈$135,000 of crash-proxy equity per contract — itself a finding (Section 5.4), and the practical face of Santa-Clara and Saretto's (2009) margin constraint.

5. Results

5.1 Headline

Zero of 134 configurations survive validation on SPY (2010–2023); zero survive the QQQ replication (2012–2023); and zero survive the N = 140 conditional-richness extension (Section 5.8). The best trial in the program is a risk-managed short strangle at DSR 0.90 (run-level PBO 0.079/0.198; verdict runs 20260718T062122Z and 20260718T065336Z; extension run 20260718T233343Z). Table 1 reports the full-OOS ladder for the principal configurations.

Table 1 — Full out-of-sample performance, SPY 2010–2023 (stitched walk-forward OOS; costs inclusive). DSR from the N = 134 verdict run. † = item-J caveat (cash-settled assignment).

ConfigurationTotal ret. %SharpeCalmarMaxDD %Pos. months %DSRTrades
Buy-and-hold SPY (baseline)289.20.810.3834.270.5
Covered call (0.30Δ)54.80.640.2615.670.50.09202
Short strangle, bare (0.20Δ, 50% PT)54.80.310.0498.666.70.00173
Short strangle, bare (hold-to-expiry)−53.60.18−0.0796.168.20.00114
Short strangle, managed ($150k)16.01.520.851.679.50.90190
Short strangle, weekly managed17.41.170.692.173.50.62382
Jade lizard (0.30Δ put)130.30.400.1552.987.90.0154
Covered strangle (0.20Δ)†99.90.610.2131.271.20.07173
Weekly covered call71.40.680.3016.567.40.11433

5.2 The premium-selling family is repackaged equity beta

Attribution (Section 4.4) assigns 87–99% of the cash-secured-put / spread / condor / covered-call family's returns to the passive-equity component — the out-of-sample, retail-frictions confirmation of Israelov and Nielsen's (2015) decomposition. The dynamic-timing residual is uncompensated, and the short-volatility component is small (covered call: +3.7% of capital over the full period against +33.8% passive). Two corollaries follow. First, the frictionless check: with spread costs set to zero the short-volatility book still lost over 2017–2023 — the deficiency is edge, not costs. Second, the sample-artifact catch: over 2017–2023 alone, the covered call appeared to beat buy-and-hold risk-adjusted (Sharpe 0.71 vs 0.57); extending the sample through the four bears reversed the ranking (0.57 vs 0.71). Drawdown *cushioning* is real — the covered call absorbed one-third to one-half of buy-and-hold's drawdown in every episode — but it is compensation-neutral cushioning, not alpha, consistent with Whaley (2002) and Ungar and Moran (2009) read at full cycle. The undefined-risk naked put-write lost 58% in the COVID crash (72% max drawdown), the textbook left tail of Ilmanen's (2012) insurance seller.

5.3 Trend is insurance; the combination is preservation, not alpha

The 1-month long/short time-series momentum overlay is positive in all four bear episodes, carries negative net beta (≈ −0.14), and its correlation to the carry book is −0.68 *inside the bears* — the crisis convexity documented by Moskowitz, Ooi, and Pedersen (2012) and Hurst, Ooi, and Pedersen (2017), reproduced at retail frictions. It nonetheless fails DSR alone: it purchases crash convexity with bull-market underperformance. A fixed ≈70/30 carry/trend composite (weight *timing* was tested and rejected — it fits noise and loses to fixed weights out-of-sample, the small-scale echo of why factor-timing disappoints) achieves Sharpe 0.78 against buy-and-hold's 0.81, maximum drawdown ≈9% against 34%, beta ≈0.21, CAGR ≈3.3%, DSR ≈0.43. It ties the index risk-adjusted at one-quarter of the drawdown, for one-third of the total return: a capital-preservation vehicle — directly responsive to the sequence-risk mandate (Bengen, 1994) — but not an edge, and not income.

5.4 Income structures under the bear-window test, and the VRP at survivable scale

Table 2 reports the mandate's headline evaluation: behavior inside the four real drawdown episodes.

Table 2 — Bear-window performance, SPY (total return % / max drawdown % within each window).

Structure2015–162018-Q4COVID 20202022Full period
Buy-and-hold SPY−8.7 / 14.6−13.9 / 19.3−13.9 / 34.2−18.8 / 24.6+289 / 34.2
Covered call−3.2 / 4.7−4.6 / 6.8−7.2 / 15.6−9.6 / 15.1+55 / 15.6
Strangle, bare−7.5 / 15.4+1.9 / 20.1−66.5 / 98.5+7.6 / 35.9+55 / 98.6
Strangle, bare, hold-to-expiry−3.0 / 11.0−11.4 / 28.2−83.6 / 86.1+41.7 / 12.5−54 / 96.1
Strangle, managed+0.1 / 0.8−0.5 / 1.0−0.5 / 1.5+2.8 / 0.7+16 / 1.6
Jade lizard (0.30Δ)+0.1 / 0.0+3.4 / 4.4+5.7 / 52.9−3.5 / 38.2+130 / 52.9
Covered strangle (0.20Δ)†−4.1 / 6.9−5.7 / 10.5−16.5 / 31.2−10.9 / 19.3+100 / 31.2
Weekly covered call−3.0 / 5.0−6.8 / 9.5−6.9 / 16.5−9.5 / 14.9+71 / 16.5

Four findings.

(1) Unmanaged undefined-risk premium selling is ruinous. The bare strangle's COVID outcome (−66% to −84%; peak drawdowns 86–99%) is not an outlier draw but the expected left tail of short-gamma-both-sides — the realized version of the distribution Ilmanen (2012) describes. The hold-to-expiry variant *lost 54% over the full 14 years*: years of premium did not cover one crash.

(2) Risk management achieves survival — and survival caps income. The managed strangle (2×-credit stop, fixed-fractional sizing against a 15%-crash proxy, re-entry cooldown) held its maximum drawdown to ≤2.2% through every bear. The identical sizing rule prices one SPY strangle at ≈$135,000 of crash-proxy equity, capping the honest account's income at 16% total over 14 years — ≈1.1% CAGR. Its DSR of 0.90 is the best in the program and still short of the bar. This is Santa-Clara and Saretto (2009) operating at the household level: the constraint that makes the strategy safe makes it small.

(3) A direct estimate of the retail-capturable index VRP. Attribution identifies the managed strangle as the program's cleanest volatility-premium harvester: measured mean entry net delta of −0.8 shares (delta-neutral at entry), passive-equity component −0.2% of capital, dynamic +0.1%, short volatility +4.8% of capital over 14 years (≈0.35%/yr), ≈1.1%/yr CAGR inclusive of collateral yield. Read against the gross premia of Carr and Wu (2009) and Bakshi and Kapadia (2003), the gap *is* the paper's central quantity: after retail spreads, discrete management, and crash-survivable sizing, an order of magnitude of the index VRP is unavailable to the household. The premium the income industry sells is approximately one-tenth the size of the income it quotes.

(4) Structural engineering does not transfer risk away; it relocates it. The jade lizard eliminates upside risk by construction and retains the entire CSP-like downside (COVID max drawdown 52.9%); the covered strangle behaves as a levered covered call (COVID −16.5% / 31.2%); weekly expiries quadruple turnover (433 vs 114–202 trades) without changing the economics — consistent with Bryzgalova, Pavlova, and Sikorskaya's (2023) evidence that short-dated contracts are where retail expected returns are worst.

5.5 The trial-count tax, stated as a result

Table 3 — Selection penalty in action: DSR before (N = 109) → after (N = 134), returns unchanged.

ConfigurationDSR, N=109DSR, N=134
Buy-and-hold SPY0.460.21
Carry+trend composite (70/30)0.430.19
12-mo momentum, vol-scaled0.400.17
Covered call (0.16Δ, managed)0.290.10

Widening the search — even with hypotheses that all failed — lowered every incumbent's deflated significance, exactly as Bailey and López de Prado (2014) intend and Harvey, Liu, and Zhu (2016) urge. We report it as a result because it is the discipline most retail (and some professional) backtesting omits: a strategy's evidence must be discounted by everything else that was tried, and the discounting must be applied to one's own search, contemporaneously.

5.6 Build-versus-buy against live benchmarks

Over the composite book's OOS span (2012-07 to 2023-07), the 70/30 book (CAGR 3.37%, maxDD 9.2%, Sharpe 0.79) beats every CBOE premium-selling benchmark risk-adjusted at roughly one-third the drawdown — PUT 7.25% / 28.9% / 0.64; BXM 6.21% / 30.3% / 0.54; BXMD 8.74% / 32.3% / 0.65 — while earning about half their income; CNDR (−0.59% CAGR over the span) independently corroborates our iron-condor result. In JEPI's own live window (2020-05 to 2023-07) the fund wins on both axes (CAGR 13.36%, Sharpe 1.17 vs 6.16%, 0.98), with the caveats that the window is short, contains one bear, and compares a live track against a backtest. The household implication is allocational, not strategic: the income sleeve is more efficiently bought than manufactured.

5.7 Cross-underlying replication

Every qualitative conclusion replicates on QQQ 2012–2023 (N = 134, PBO 0.198, zero survivors): bare strangles −34% to −36% total with drawdowns to 85%; managed strangles ≈1% CAGR; covered calls capture 55–68 points of buy-and-hold's +330%; the best non-baseline DSR is again a trend rung, not a premium seller.

5.8 Conditional harvesting: the implied-minus-realized richness gate

A standard institutional objection to mechanical premium selling is that volatility should be sold only when implied volatility is rich relative to realized volatility — not on a calendar. Our primary grid already conditions on IV rank and volatility regime (neither survived); as an extension we added an explicit IV−RV trigger: entry requires the day's ATM-implied-minus-21-day- realized-volatility spread to sit at least z standard deviations above its own trailing distribution (z ∈ {0.5, 1.0}), applied to the cash-secured put, the covered call, and the risk-managed strangle (six rungs; N = 140; verdict run PBO = 0.095; zero survivors).

The result is instructive in both directions. The gate materially improves the *weakest* configurations — the richness-gated cash-secured put returns +6.9% to +9.5% where its ungated sibling *loses* 0.8%, and the gated covered call beats its sibling (+76.7% vs +54.8% at z = 1.0) — evidence that thin-premium periods are where mechanical selling bleeds. But the same gate *degrades* the best configuration: the delta-neutral managed strangle earns less gated (+11.2% to +14.6%) than ungated (+16.0%) and fails to beat its sibling, because once a position is delta-neutral and crash-sized, days skipped for "thin premium" cost more in foregone carry than they save. Conditioning on richness rescues bad implementations and taxes good ones; it is not the missing ingredient. The extension also re-applied the selection penalty to every incumbent (the managed strangle's DSR moved 0.90 → 0.87 at N = 140), consistent with Section 5.5.

6. Taxes, Wrappers, and the Rate Environment

Two further objections to any pre-tax negative result are that tax treatment or today's higher cash rates change the ranking. We quantify both.

6.1 After-tax comparison across wrappers

Broad-based index options (SPX/XSP) receive Section 1256 treatment — 60% long-term / 40% short-term federal capital-gains rates, marked to market annually — while ETF options and covered-call fund distributions are taxed as ordinary income. We apply calendar-year taxation with loss carryforward to the same total-return series as Section 5.6, at a top-bracket California profile (37% federal ordinary, 20% federal LTCG, 3.8% NIIT, 13.3% state; California does not conform to 60/40, and Treasury interest is state-exempt).

Table 4 — After-tax CAGR by wrapper, JEPI's live window (2020-05–2023-07). The XSP row is an upper bound: XSP quoted spreads are wider than the SPY frictions in the gross curve.

SeriesWrapperPre-tax CAGR %After-tax CAGR %
JEPITaxable CA (ordinary)13.366.18
70/30 bookTaxable CA, SPY options (ordinary)6.162.76
70/30 bookTaxable CA, XSP §1256 + exempt interest6.163.47
S&P 500 total returnTaxable CA (LTCG; deferral ignored)16.5710.19
60/40 T-bill/SPY blendTaxable CA (interest fed-only + LTCG)8.034.84
Any of the aboveRoth/IRAunchangedunchanged

Three readings. First, the Section 1256 lever is *real*: over the full 2012–2023 span it inverts a gross ranking — the PUT index (7.25% gross, §1256) nets 4.06% after tax versus BXMD's 3.93% from a higher 8.74% gross taxed as ordinary income. Second, it is *insufficient*: even at the XSP upper bound the book nets 3.47% against JEPI's 6.18% — no tax treatment closes a two-to-one gross gap. Third, the dominant "levers" are not strategies at all: the Roth wrapper strictly dominates every tax optimization, and plain index buy-and-hold dominates every income strategy after tax in a taxable account.

6.2 The cash-rate environment

Because the engine credits the historical T-bill path on collateral (sample mean 0.94%), part of any forward-looking improvement is simply today's rate level. Re-running the identical 70/30 book with the cash rate pinned:

Table 5 — The 70/30 book under flat cash-rate scenarios (full period, 2010–2023).

ScenarioCAGR %SharpeMaxDD %
Historical T-bill path (mean 0.94%)2.910.896.96
Flat 1%2.950.936.70
Flat 3%4.301.545.63
Flat 5%5.792.364.64

A 5% cash environment lifts the book by ≈2.9 points — genuine and forward-relevant, but *symmetric*: T-bills alone, an income fund's collateral, and any blend receive the same lift, so the rate environment changes absolute returns without changing any relative ranking. Claims that stack today's cash yield onto a strategy's historical return double-count whatever collateral yield the historical record already contains.

7. The Mechanism: Why the Result Generalizes

The paper's negative claim does not rest on having exhausted the space of option structures; it rests on three independently measured links that jointly constrain *every* index-option income structure:

  1. The income is beta in disguise. Attribution (Section 5.2) assigns 87–99% of the premium-selling family's return to passive equity exposure. Whatever the structure — put, call, spread, condor, lizard — the dominant return source is "own the market," obtainable more cheaply and with better tax treatment by owning the market.
  2. The deficit is edge, not cost. With spread costs set to zero the short-volatility book still lost over 2017–2023 (Section 8). No execution improvement can rescue a return source that loses frictionless.
  3. The de-betaed residual is ≈1% per year at survivable size. The delta-neutral, risk-managed strangle (Section 5.4) — mean entry delta −0.8 shares, return ≈pure short volatility — is the cleanest available harvest of the index VRP, and at the position size that survives every bear in sample it captures ≈1% of capital per year.

Any index-option income strategy is a portfolio of link 1 and link 3, priced under link 2. A structure not in our grid — a different delta, a different tenor, a ratio variant — re-weights those measured ingredients; it does not create new ones. The supporting analyses close the remaining degrees of freedom: conditioning on richness redistributes but does not enlarge the premium (Section 5.8); tax treatment cannot close the gap and is dominated by the wrapper (Section 6.1); the rate environment shifts all cash-holding alternatives together (Section 6.2). The result therefore generalizes to retail index-option income strategies as a class, at end-of-day management frequency and retail scale — the boundaries of that class, and the live-fill measurement that could still tighten the cost bound, are stated in Section 9.

8. Robustness and the Audit Trail

We report the study's self-corrections as first-class evidence, because a negative headline is only as credible as the effort spent trying to overturn it in *both* directions.

Corrections that raised measured performance. A forensic engine audit found the validation mathematics sound but the *economics* incomplete in ways that biased returns down: missing dividends, collateral interest, and borrow costs; a price-only (rather than total-return) buy-and-hold baseline; an iron condor with no exit rule. Correcting these moved the best composite's DSR from 0.12 to 0.785. Honest frictions cut both ways, and a study built to fail strategies must also repair biases against them.

Corrections that lowered it. The 0.785 proved flattered by two optimistic defaults: a smooth flat collateral rate and same-bar execution. Replacing them with the historical T-bill path and next-bar fills — each worth ≈0.13 of DSR — reduced it to 0.43. Both realistic settings became the engine defaults for all results reported here.

Three overturned positives. (i) The covered call's 2017–2023 risk-adjusted "edge" reversed when the sample was extended through the four bears (Section 5.2) — a within-study demonstration of why sub-period evidence misleads. (ii) Ten configurations briefly appeared as DSR = 1.0 "survivors" because their sizing rule had refused every trade, leaving zero-variance collateral-interest curves; the zero-fill exclusion now bars degenerate curves categorically. (iii) Two implementation defects — a days-to-expiry filter that silently emptied the weekly covered-call rungs, and the zero-fill gate misfiring on stock-rebalancing strategies — were caught by reading run output against expectations, reproduced as failing tests, and fixed before the verdict runs. The engine's acceptance test additionally verifies the apparatus flags seeded noise and a deliberately overfit configuration.

All 365 behavioral tests and 4 smoke tests pass; every number in this paper regenerates from committed code and logged artifacts, seeds included.

9. Limitations and a Registered Follow-Up

End-of-day data precludes intraday management and understates what execution-sophisticated traders achieve on spreads (Muravyev and Pearson, 2020) — our friction model is a stated conservative bound, not an estimate of best practice. Early exercise is approximated by cash settlement (flagged wherever it matters, principally the covered strangle). The tested universe is two index ETFs — deliberately the thinnest premiums; single-name premia are richer and untested here. Dividends accrue via yield on the stock leg rather than discrete ex-dates. The JEPI comparison sets a live track against a backtest. The sample, while spanning four drawdowns, is a single macro era (low-rate, largely bull, 2010–2023); walk-forward stitching mitigates but cannot eliminate era dependence. Finally, N penalizes our search, not the industry's: against the denominator of every strategy ever marketed, all published Sharpe ratios — ours included — deserve further deflation (Harvey, Liu, and Zhu, 2016).

A registered, explicitly directional follow-up on fill quality. The full-quoted-spread assumption is the one conservative bound a critic can quantitatively attack, so we have instrumented a live measurement: each paper-traded proposal logs its quoted bid/ask/mid at signal time, and realized (paper) executions are matched back to compute a capture ratio k per fill (k = 1 reproduces our assumption; k = 0 is a mid fill). We pre-commit to its limits and its handling: fills are manually placed under a fixed limit-at-mid, walk-one-tick protocol, so the human times the fill and the measurement is a *directional proxy*, not a clean capture ratio (a clean version requires automated placement with mid-quote timestamping); no headline is read below 30 matched fills; and the measured median k will be reported and, if it differs from 1, the headline grid re-run at that k — whichever direction the result cuts, including the direction that would weaken this paper's cost bound.

10. Discussion

For households. Three practical conclusions follow. First, index-option premium selling is not an income source at retail scale: its return is either equity beta (own the index directly, at lower cost and better tax treatment) or a residual VRP too thin to matter once sized for survival. Second, the validated product of premium selling is *drawdown shaping* — the covered call and especially the carry+trend composite cut crash participation by one-half to three-quarters at proportional cost to return; that is a legitimate allocation choice for sequence-risk-dominated investors, provided it is priced as insurance, not sold as income. Third, the highest-certainty levers for the income mandate are not strategies, and Section 6 now quantifies them: the Roth wrapper dominates every tax optimization; where a strategy must live taxable, Section 1256 index treatment is worth using (it inverts the PUT/BXMD after-tax ranking) but cannot rescue an inferior gross return; and collateral-yield hygiene is worth its full face value while conferring no relative advantage.

For researchers. The survivable-size VRP estimate (≈1%/yr on capital) is, to our knowledge, a quantity the literature has not directly reported: gross premia (Carr and Wu, 2009) and institutional strategy returns (Whaley, 2002; Ungar and Moran, 2009) bracket it from above, and the retail-loss literature (Bauer et al., 2009) brackets it from below, but the constrained-implementation measurement in between is the number households actually face. It suggests the retail options loss record is not primarily an execution failure but a structural property of the opportunity set — a hypothesis single-name extensions could falsify, since single-name premia are several multiples richer (Carr and Wu, 2009) against correspondingly fatter idiosyncratic tails. That extension — a liquidity- and quality-screened single-name premium basket, subject to the identical DSR/PBO bar — is this program's next registered hypothesis, with a deliberately modest prior; the directional live-fill measurement of Section 9 is the other, and the only one that could tighten this paper's own cost bound.

On method. The study's most transferable finding is procedural: a walk-forward, DSR/PBO-corrected, attribution-audited harness is buildable with commodity data and a standard library, runs a 134-configuration, 14-year evaluation overnight, and — three times in this study — protected the researcher from his own preliminary results. Given the base rates documented by the multiple-testing literature, we would argue no retail-facing strategy claim deserves attention without at least this apparatus behind it.

11. Conclusion

Under conservative retail frictions and selection-bias-corrected inference, none of 140 option-income configurations on SPY (2010–2023) or QQQ (2012–2023) exhibits validated edge — including variants conditioned on implied-versus-realized volatility richness, and under every tax wrapper and cash-rate environment examined. The premium-selling family's returns are, to first order, equity beta; undefined-risk structures fail catastrophically unmanaged and shrink to ≈1% annual income when sized to survive; and a delta-neutral, attribution-verified harvest of the index volatility risk premium at crash-survivable scale captures roughly one-tenth of the income such strategies are marketed to produce. Because these are measurements of the strategies' shared ingredients rather than a tour of their combinations, the conclusion binds the class: at retail scale, the index VRP cannot fund a retirement, and the constraint is arithmetic, not discipline. What can be validated is drawdown shaping — bought with return — and a set of non-strategy levers (the tax wrapper, bought income, collateral yield) that dominate every measured edge. The apparatus that produced these conclusions, and overturned three of its own, is the contribution we expect to outlast them.

Acknowledgments

The backtesting engine, validation suite, and manuscript drafting were built with the assistance of an AI system (Claude, Anthropic). Research design, governing decisions, and conclusions are the author's. Any errors are the author's own.

References

Asness, C. S., Moskowitz, T. J., & Pedersen, L. H. (2013). Value and momentum everywhere. *Journal of Finance*, 68(3), 929–985.

Bailey, D. H., Borwein, J., López de Prado, M., & Zhu, Q. J. (2017). The probability of backtest overfitting. *Journal of Computational Finance*, 20(4), 39–69.

Bailey, D. H., & López de Prado, M. (2014). The deflated Sharpe ratio: Correcting for selection bias, backtest overfitting, and non-normality. *Journal of Portfolio Management*, 40(5), 94–107.

Bakshi, G., & Kapadia, N. (2003). Delta-hedged gains and the negative market volatility risk premium. *Review of Financial Studies*, 16(2), 527–566.

Barber, B. M., Lee, Y.-T., Liu, Y.-J., & Odean, T. (2009). Just how much do individual investors lose by trading? *Review of Financial Studies*, 22(2), 609–632.

Barber, B. M., & Odean, T. (2000). Trading is hazardous to your wealth: The common stock investment performance of individual investors. *Journal of Finance*, 55(2), 773–806.

Bauer, R., Cosemans, M., & Eichholtz, P. (2009). Option trading and individual investor performance. *Journal of Banking & Finance*, 33(4), 731–746.

Bengen, W. P. (1994). Determining withdrawal rates using historical data. *Journal of Financial Planning*, 7(4), 171–180.

Bondarenko, O. (2014). Why are put options so expensive? *Quarterly Journal of Finance*, 4(3), 1450015.

Bryzgalova, S., Pavlova, A., & Sikorskaya, T. (2023). Retail trading in options and the rise of the big three wholesalers. *Journal of Finance*, 78(6), 3465–3514.

Carr, P., & Wu, L. (2009). Variance risk premiums. *Review of Financial Studies*, 22(3), 1311–1341.

Coval, J. D., & Shumway, T. (2001). Expected option returns. *Journal of Finance*, 56(3), 983–1009.

Harvey, C. R., & Liu, Y. (2015). Backtesting. *Journal of Portfolio Management*, 42(1), 13–28.

Harvey, C. R., Liu, Y., & Zhu, H. (2016). … and the cross-section of expected returns. *Review of Financial Studies*, 29(1), 5–68.

Hurst, B., Ooi, Y. H., & Pedersen, L. H. (2017). A century of evidence on trend-following investing. *Journal of Portfolio Management*, 44(1), 15–29.

Ilmanen, A. (2012). Do financial markets reward buying or selling insurance and lottery tickets? *Financial Analysts Journal*, 68(5), 26–36.

Israelov, R. (2019). Pathetic protection: The elusive benefits of protective puts. *Journal of Alternative Investments*, 21(3), 6–33.

Israelov, R., & Nielsen, L. N. (2015). Covered calls uncovered. *Financial Analysts Journal*, 71(6), 44–57.

Koijen, R. S. J., Moskowitz, T. J., Pedersen, L. H., & Vrugt, E. B. (2018). Carry. *Journal of Financial Economics*, 127(2), 197–225.

López de Prado, M. (2018). *Advances in Financial Machine Learning*. Wiley.

Moskowitz, T. J., Ooi, Y. H., & Pedersen, L. H. (2012). Time series momentum. *Journal of Financial Economics*, 104(2), 228–250.

Muravyev, D., & Pearson, N. D. (2020). Options trading costs are lower than you think. *Review of Financial Studies*, 33(11), 4973–5014.

Pardo, R. (2008). *The Evaluation and Optimization of Trading Strategies* (2nd ed.). Wiley.

Politis, D. N., & Romano, J. P. (1992). A circular block-resampling procedure for stationary data. In R. LePage & L. Billard (Eds.), *Exploring the Limits of Bootstrap* (pp. 263–270). Wiley.

Santa-Clara, P., & Saretto, A. (2009). Option strategies: Good deals and margin calls. *Journal of Financial Markets*, 12(3), 391–417.

Sullivan, R., Timmermann, A., & White, H. (1999). Data-snooping, technical trading rule performance, and the bootstrap. *Journal of Finance*, 54(5), 1647–1691.

Ungar, J., & Moran, M. T. (2009). The cash-secured PutWrite strategy and performance of related benchmark indexes. *Journal of Alternative Investments*, 11(4), 43–56.

Whaley, R. E. (2002). Return and risk of CBOE buy write monthly index. *Journal of Derivatives*, 10(2), 35–42.

White, H. (2000). A reality check for data snooping. *Econometrica*, 68(5), 1097–1126.

CBOE. Strategy benchmark index methodologies: BXM, BXMD, PUT, CNDR. Cboe Global Markets.

Appendix A. Validation Mathematics

Deflated Sharpe Ratio (Bailey and López de Prado, 2014). For a configuration with T out-of-sample per-period returns, per-period Sharpe SR, skewness γ₃, and excess kurtosis γ₄:

DSR = Φ( (SR − SR₀) · √(T − 1) / √( 1 − γ₃·SR + ((γ₄ + 2)/4)·SR² ) )

SR₀ = √V · ( (1 − γ)·Φ⁻¹(1 − 1/N) + γ·Φ⁻¹(1 − 1/(N·e)) )

where N = 134 is the number of active strategy trials, V is the cross-sectional (ddof = 1) variance of the N trials' per-period Sharpe ratios, γ ≈ 0.5772 is the Euler–Mascheroni constant, and Φ is the standard normal CDF. Survival requires DSR ≥ 0.95.

Probability of Backtest Overfitting (Bailey et al., 2017). The T×N matrix of date-aligned daily OOS returns is partitioned into S = 10 contiguous blocks. For each of the C(S, S/2) half/half partitions: compute per-configuration Sharpe on each half; identify the in-sample-best configuration n*; record its out-of-sample rank; set λ = logit(rank/(N+1)). PBO = the fraction of partitions with λ ≤ 0. Survival requires PBO ≤ 0.50.

Bootstrap. Circular block bootstrap (Politis and Romano, 1992), block length L = max(5, ⌊T^⅓⌉), B = 1,000 resamples, seeds logged; 5th/95th percentile intervals on total return, per-period Sharpe, and max drawdown.

Regime split. ATM implied volatility maps each OOS day to calm/normal/elevated/crisis (thresholds 15/25/35 vol points); survival requires positive total return in every band with ≥20 days.

Exclusions and guards. Baselines (cash floor, buy-and-hold) and any configuration with zero fills in the scored OOS are categorically barred from survival; verdicts require ≥20 trials and ≥252 OOS observations; an acceptance test verifies the apparatus rejects seeded noise and a deliberately overfit configuration.

Appendix B. The Strategy Grid (N = 134 primary; 140 with the extension)

FamilyAxesRungs
Cash-secured puttarget Δ {0.16, 0.20, 0.30, 0.40} × profit target {50, 75, hold} × filter preset {none, regime, IV-rank, both}, + bare/managed reference rungs50
Bull put spreadshort Δ {0.16, 0.20, 0.30} × profit target {25, 50, 75} × filter {none, both}, + bare/managed reference rungs20
Iron condorshort Δ {0.16, 0.20, 0.30} × filter {none, both}, + bare/managed reference rungs8
Covered callbare Δ {0.16, 0.20, 0.30, 0.40}; delta-managed Δ × rehedge band {0.05, 0.10, 0.20}; weekly (≈10 DTE) Δ {0.16, 0.20, 0.30}; + bare/managed reference rungs21
Naked put-write (cautionary)single rung1
Short stranglebare Δ {0.16, 0.20, 0.30} × PT {50, hold}; managed ($150k) same grid; weekly managed Δ {0.16, 0.20} × PT {50, hold}16
Jade lizardput Δ {0.20, 0.30} × PT {50, hold}4
Covered strangleΔ {0.20, 0.30}2
Time-series momentumlookback {1, 12 mo} × {long-flat, long-short} × vol-scaling {on, off}8
Carry+trend compositescarry weight {0.50, 0.60, 0.70, 0.80}4
Primary total134
IV−RV richness extension (Section 5.8){CSP, covered call, managed strangle} × z {0.5, 1.0}6
Extension total140

Management overlays where applicable: profit targets, 2×-credit stop-loss, 21-DTE time exit, fixed-fractional sizing vs a 15%-crash proxy, re-entry cooldown. Baselines (T-bill cash floor; total-return buy-and-hold) run in every walk-forward window and are excluded from the trial count and from survival by construction.