A 5.76 Sharpe means less when the book realizes only 1.82% annualized volatility. Wysocki acknowledges as much. Reaching the CBOE PUT index's 14.36% out-of-time volatility would require roughly eight times the position size, which margin and overnight gap risk make unavailable. He therefore treats the headline figure as an upper bound on live deployment. The useful question is what the 2025 hold-out actually put under test.

One daily strike on the put wing

The strategy harvests the volatility risk premium by selling index puts priced above the distribution subsequently realized. Bondarenko's figure sizes the wedge: from 1990 to 2018, the VIX averaged 19.3% while subsequent realized one-month volatility averaged 15.1%, leaving a gap of 4.2 annualized points. Passive harvesters including CBOE PUT and WPUT, the CBOE weekly put-write index, collect that premium through a static rule and retain the same exposure across regimes.

Wysocki makes a daily selection instead. At 10:00 ET, the model ranks nine candidates. Eight are short puts on SPXW, the S&P 500 weekly index options, mapped to delta targets from 0.05 to 0.45 on the shortest listed expiry. The ninth candidate is SKIP, a synthetic row scored at zero so the model can decline the trade.

A LightGBM LambdaRank ranker orders the day's nine candidates rather than forecasting each candidate's return. Each trading day forms One query group. The training label divides trade gross P&L by downside deviation calculated from one-minute bars, then bins the ratio into five grades at the 10th, 40th, 60th and 90th percentiles. A confidence gate compares the top-one and top-two scores. When their gap falls below a per-window threshold, the model abstains. Seven sizing rules translate the remaining selection into a contract count under Reg-T index-put margin and IBKR's tiered fees.

The data comprise CBOE one-minute bars and intraday quotes on SPXW from 2018 to 2025, with 2017 used as warm-up. Four expanding walk-forward windows generate predictions for 2021 through 2024. The year 2025 remains held out. Across that hold-out, the seven sizing rules deliver Sharpe ratios between 4.308 and 5.761, versus 0.1754 for CBOE PUT, 0.1312 for WPUT and 0.4622 for SPX buy-and-hold.

Edge Allocation is the headline rule. It sets margin utilization according to the percentile rank of today's implied-minus-realized volatility edge. Out of time, the rule returns 10.48% at 1.82% vol and records a 1.43% drawdown. Its walk-forward Sharpe is 3.1048.

We could not rerun any of it. Constructing the label, resolving the strike and determining the entry fill all depend on intraday SPXW quotes plus a 10:00 ET surface snapshot, while our option data are end-of-day only. Without those intraday quotes, even the daily candidate set cannot be reconstructed. The one-minute path underlying the label is further out of reach.

Much of the build is printed

As a build document, the paper is unusually complete. It gives the LambdaRank gain vector ([0,1,3,7,15]) and lays out the five-stage feature filter. The stated settings include a 0.30 null rate, a 0.05 relevance floor and a 0.85 clustering threshold, along with per-group survivor caps (position 8, vol surface 4, vix 6, calendar 8, default 5).

Each window receives Fifty Optuna trials optimizing ranking quality at the top position, and the search space is tabulated. The paper also supplies the Reg-T margin formula, fee tiers of $0.25, $0.50 and $0.65 per contract with a $1.00 minimum, the per-method sizing grids and $5,000,000 in starting capital. Appendix B catalogs the roughly 190 candidate features, block by block, including when each feature becomes observable and how many trading days it is lagged.

Where judgment still enters

Intraday marking is the first open choice. Intermediate bars are marked at the ask, the appropriate convention for a short position. Yet a one-minute ask remains a snapshot, and those marks entirely determine the downside deviation in the label's denominator. Changing the convention changes the ratio and therefore its percentile-based grade boundaries.

The gate calibration leaves another gap. Its threshold maximizes Sortino over the last six months of each training window, using what the paper elsewhere calls an eight-point trade-rate grid. We did not find the eight points listed.

Then comes the Sharpe convention, an easy source of apparent replication failure. The numerator uses geometric compounded annual return. Wysocki reports both versions: 5.761 geometric and 5.489 arithmetic. He further explains that the adjacent Probabilistic Sharpe and Deflated Sharpe cells use the arithmetic input. The deflated value therefore does not deflate the Sharpe printed beside it.

A preprocessing change can flatten 2025

Raise the correlation-clustering threshold from 0.85 to 0.90 and the confidence gate takes no position on any day in 2025, across all seven sizing methods. Walk-forward Sharpe drops to 0.281.

Wysocki explains the chain. Looser clustering allows redundant features into the model, destabilizing the top-one-minus-top-two score gap. Calibration then tightens the gate to a threshold of 0.0689, paired with a calibrated admission rate of 0.3005. Every 2025 observation falls short.

The paper also provides a pre-trade warning signal. Calibrated and realized admission rates separate on the first morning the gate is used. Given a 0.3005 admission rate, ten consecutive abstentions have probability 0.028. The suppressed year can therefore be distinguished from routine abstention within two trading weeks, using the signal alone and before any position opens.

Sizing calibration has little room to work. Theta reaches the top of its grid in 24 of 35 method-window cells and in all five Edge Allocation windows. Hitting the 16% volatility anchor would require a margin utilization fraction from 1.07 to 1.58 in four of five windows, beyond its upper bound of one. Mean realized training volatility is 0.139. Wysocki reports the mismatch and concludes that the methods do not share a common risk budget.

What did the hold-out establish?

The two-by-two ablation answers most of the question. With neither risk control, out-of-time Sharpe reaches 5.221 against 0.175 for CBOE PUT. Of the full 5.586 gap, 5.046 already exists before either control enters. Tail-risk features contribute 0.652, the gate contributes 0.012 and their interaction subtracts 0.124, leaving 0.54 between them.

Edge Allocation holds a position on 220 of its 237 out-of-time days (Table D.1), with 17 sessions flat. The walk-forward slice looks very different because the gate binds there. Neither control succeeds alone: 0.160 for features only, 0.759 for gate only and 0.623 for neither. The +2.809 interaction carries the combined result to 3.105.

Wysocki concedes the empirical limitation. Section 7.3 says the gate's 2025 calibration produces a threshold of zero in every configuration and admits every day, leaving its effect there negligible. The conclusion states that the two risk controls "earn their place on the walk-forward slice, where they act jointly".

His defence rests on complementarity. Together, the controls produce 3.105 on walk-forward. Used alone, the tail-risk features reduce Sharpe by 0.463, while the gate adds 0.136. Wysocki interprets this as the gate withholding capital on days flagged as adverse by the tail-risk features. He declines to assert day-level overlap because the grid does not establish it. The restraint is warranted, and the remaining evidence is thin: an interaction term measured across the four years used to design both components. The paper demonstrates no day-level mechanism, while every hold-out configuration leaves the controls unbound. Even so, the abstract presents "an abstention rule driven by model uncertainty" as one of three integrated novelties.

The 2025 evidence therefore covers the ranker, the nine-candidate universe, the sizing rule and the friction model. The abstention layer sits outside that test because it never bound. Almost its entire empirical case comes from the walk-forward interaction term, with a standalone contribution of +0.136. A replicator inherits a component lacking out-of-time support. Table E.1 reports the calibrated threshold for each window, though the search grid's eight trade-rate points are never enumerated. The same component can also produce a flat year after a change in feature preprocessing.

The paper contains the rest of the audit. Edge Allocation's out-of-time Deflated Sharpe is 0.8563. Diebold-Mariano fails to reject equal mean daily P&L against every passive benchmark after Bonferroni, with a corrected p of 1.000. Repricing every fill at the bid while freezing selections and contract counts removes 7.8% of the headline Sharpe, taking it from 5.761 to 5.311. The 75%-spread cell lands at 5.426. This exercise changes the fill prices and nothing else. Capacity and adverse selection remain untested, as the paper acknowledges in its limitations. The strongest internal baseline, a 30-day rolling-Sharpe selector across the same nine candidates, earns 0.613 out of time.

A second hold-out year would change my reading if the threshold calibrated positive, the gate bound and the gate-on minus gate-off difference were measured there. On the reported evidence, the reproducible claim belongs to the ranker.