A hand-coded date repairs the one failure the paper identifies on FOMC days. The network systematically underestimates implied volatility when the announcement arrives. Supply the calendar, published months ahead, and announcement-day RMSE on the SVI call surface at horizon 1 falls from 0.099 to 0.071. For puts, it falls from 0.076 to 0.061. The random walk records 0.093 and 0.071 on those same days. Elsewhere, the model's advantage is too slight to size a position around, an assessment the authors share: "we could not confirm a full statistical significance of this better performance."
The mechanism comes first. Options price uncertainty about a scheduled Fed decision before the decision arrives. Implied volatility rises into 2pm ET even though the index has yet to react. Adamski and Ślepaczuk study CBOE one-minute S&P 500 index option quotes from 9:30 to 16:00 ET. Their sample starts in January 2016, when Wednesday expiries begin, and ends with 2025, covering roughly eight scheduled meetings a year. Average ATM implied volatility for 0DTE calls starts climbing about two days before the event. By 2pm on announcement day, it reaches up to 50% above its normal level. Realized volatility, measured as a rolling annualized one-hour standard deviation of index returns, begins rising only after the conference starts. The familiar near-expiry IV explosion runs in reverse: IV peaks before 2pm, then falls.
We judged that mechanism worth building and are evaluating one. Our implementation would use SPY options, while the paper studies cash-settled SPX contracts. Any performance we report would therefore apply only to SPY, outside their universe. Our options history begins around 2020 versus 2016 in the paper, leaving less training data for a 112k-parameter network. We hold end-of-day prices and Greeks rather than one-minute quotes. The 5-minute event study and exact 2pm snapshots remain beyond our reach. Meeting dates must come from a hand-maintained public calendar.
We are not testing their method.
The paper then turns to forecasting. Cleaned quotes are placed on a fixed grid with 41 log-moneyness points in [-0.2, 0.2] and 21 maturities spanning 0 to 20 days. Three methods construct the surface: linear interpolation with linear wing extrapolation, the Ad-Hoc Black-Scholes (AHBS) polynomial, and SVI fitted separately to each maturity slice. A convolutional 2D LSTM forecasts the entire grid from daily 2pm snapshots. Its architecture has Two ConvLSTM2D layers, 32 filters, 3x3 kernels, a 1x1 Conv2D head, a five-day lookback and 112,033 parameters. Forecasts for horizons of 1, 2, 5 and 10 days are produced autoregressively. The model trains on 2016-2023, validates on 2024 and tests on 2025, with a random-walk benchmark and six seeds. The augmented model adds a separate branch receiving one input: +1 on conference day and -1 the following day. A 64-unit dense layer joins the surface branch, raising the parameter count to 168,126.
The effect fades quickly
The Mann-Whitney results are unambiguous over the part of the surface where the effect appears. For 0DTE calls on announcement day, p is 0.000 at all nine log-moneyness points at 10:00, 12:00 and 14:00. One day later at 16:00, p ranges from 0.000 to 0.078. Three days out, every result exceeds 0.31, while k=0 reaches 0.931 at 10:00. The p-values make a pyramid, with ATM becoming significant first. The authors had expected the effect to gather in the OTM wings.
Testing was restricted to |k| <= 0.02 because short-maturity wing quotes are sparse and breach arbitrage bounds. The conclusion says directly that the evidence was insufficient to support or reject a stronger OTM effect. The displayed results describe an ATM and ITM phenomenon.
Does the date beat persistence?
At h=1, the SVI-based model reports call RMSE of 0.085 (sd 0.003), compared with 0.092 for the random walk. Put RMSE is 0.077 (sd 0.003), against 0.084.
Roughly 8% and 9%.
Most Diebold-Mariano tests against the random walk then fail. For SVI calls, mean p is 0.063 (sd 0.062) at h=1, 0.269 at h=5 and 0.508 at h=10. The exceptions are SVI calls at h=2, where mean p is 0.002, and linear-interpolation puts at h=1 and h=2, where mean p is 0.000 (sd 0.000). The authors attribute those failures to the low signal-to-noise ratio in financial time series. Lower aggregated RMSE need not carry statistical significance.
Seed instability explains another test, the global comparison between the dummy model and its dummy-less sibling. SVI calls at h=1 produce mean p of 0.310 with sd 0.450. Averaging p-values over six seeds is not itself a test. With dispersion that large, the seeds disagree on which model wins.
Two details matter more than the headline RMSE. First, surface construction establishes the error level. The AHBS-based model records 0.208, while AHBS interpolation error alone is 0.21, as the paper states explicitly. Second, linear interpolation reproduces raw quotes exactly by construction yet forecasts worse than SVI: 0.116 versus 0.085 for calls at h=1. The authors write that "a perfect fit of linear interpolation in fact impacts the forecasting a little negatively, possibly due to overfitting to noisy points".
Anyone who builds surfaces should keep that sentence.
A sibling model denied the date
On announcement days, the dummy-less model never defeats the random walk. Mean DM p-values range from 0.215 to 0.961, including 0.678 for SVI calls at h=1 and 0.920 at h=5. The paper gives the reason: "which is to be expected as the model does not have access to the information regarding the upcoming announcement."
The augmented model looks decisive for calls against its dummy-less version on those days, while puts are mixed. SVI calls post 0.001, 0.001, 0.002 and 0.004 across the four horizons. Linear calls range from 0.003 to 0.005. SVI puts record 0.053, 0.039, 0.071 and 0.176, placing three of the four above 5%. The paper describes "almost all results are statistically significant at 5% level especially for call options". With the dummy, announcement-day SVI RMSE for calls at h=1 is 0.071, compared with 0.099 without it and 0.093 for the random walk. For puts, the figures are 0.061, 0.076 and 0.071.
The remaining entries weaken the claim. At h=10 on announcement days, the random walk records 0.064 for SVI puts, while the dummy model reaches 0.081. The augmented model loses. We also did not find a Diebold-Mariano test anywhere in the paper comparing the dummy model with the random walk on announcement days. The conclusion's significance claim comes from comparison with a model denied a date available on the Fed's website.
The authors themselves make the calendar central. Their conclusion says the ML advantage "is not uniform across IV surface and is concentrated mainly around short maturities and ATM point". The final paragraph goes further. Adding scheduled policy announcements "can materially improve predictive performance on days with such announcements". Without event-driven dynamics, "the advantage of ML models remains limited, especially when dealing with noisy financial time series". Yet the abstract closes by saying "our study reinforces the perspective that ML models can effectively forecast the IV surface also during abnormal days." For announcement days, that statement relies on tests against the dummy-less sibling, with no random-walk comparison.
Where the surface forecast fails
Grid-level errors cluster in the short-maturity wings. The model underestimates IV there and slightly overestimates it near ATM. Its grid-level DM map shows no machine-learning advantage for 0DTE options specifically, exactly where the pre-announcement effect reaches its maximum.
For ATM 0DTE calls, the model-implied impulse response estimates a pre-announcement effect of up to about 20% on average. The raw observations reach up to 50%. The network smooths the shock. Its impulse response also keeps IV elevated indefinitely, reflecting learned persistence and turning the h=10 figures into extrapolation.
P&L remains open
The paper is direct about the omission: "our results lack the verification in an actual trading environment. We considered only statistical significance, which we believe to be a proper starting point." There is no option structure, no spread charge, no delta hedging and no Sharpe. Bid-ask conditions enter only through quote filters used for cleaning.
This omission carries extra weight because the target is the 2pm snapshot, when the press conference starts. The paper's plots show ATM IV peaking around 2pm and falling before the conference finishes. A trader selling into the move would want that peak level. Even so, the forecast remains several decisions removed from an executed 0DTE trade.
The risk-free rate uses the 5-year CMT yield for a book dominated by short-dated options. The authors acknowledge the mismatch and argue that its effect on log-moneyness is limited, with little chance of changing the surface forecasts. We think that defence is adequate for a forecasting grid.
A narrow, inexpensive result would change my mind: a DM test comparing the dummy model with the random walk on announcement days, together with short-dated straddle or calendar P&L across the roughly eight 2025 meetings after quoted spreads. Until then, the transferable result comes from the empirical half and requires no network. It recalls our review of a 0DTE SPXW ranker whose abstention layer never bound out of sample (/articles/the-5-76-sharpe-lives-in-a-1-82-percent-denominator). The apparent innovation earns its keep on a handful of days each year. Most of the work went into the machinery surrounding it.