A single skew-stickiness ratio leaves two terms of the smile response unmeasured. Che and Das show that both matter and vary across the term structure. On SPX, the skew-transport term changes sign between 6M and 9M. The joint null that both higher terms are zero is rejected at every tenor from 1M to 24M.
Our figures require an early disclosure. We could not trade SPX fixed-tenor surfaces, so we used SPY listed options. Dividends, exercise conventions, settlement and ETF option microstructure differ. The paper's SPX estimates therefore do not transfer directly to what we traded.
The paper recasts smile dynamics cleanly. Total implied variance, w(k,T) = sigma^2 T, is the state variable, with k = log(K/F) and u = log F. Spot-driven dynamics becomes the transport equation d_u w = v(k) d_k w. Each classical convention then amounts to a choice of velocity field v. Sticky delta sets v = 0. Sticky strike sets v = 1. Bergomi's SSR takes v as a constant beta.
The jets are the Taylor data of that field at the money: beta = v(0), eta = d_k v(0), psi = d_kk v(0). Differentiate the transport equation and the ATM variance jets form a closed lower-triangular system. Level moves as beta times skew. Skew moves as beta times curvature plus eta times skew. Curvature moves as beta times a3 plus 2 eta times curvature plus psi times skew. SSR is therefore the ATM slice of a field, while one-speed transport imposes the same number on level, skew and curvature.
Daily SPX surface snapshots from 12 July 2021 to 9 July 2026 supply the empirical sample, with n = 1,118 usable dates per tenor. Seven tenors run from 1M to 24M. The strike grid contains 11 points of K/F from 0.85 to 1.15, spanning roughly minus 0.16 to plus 0.14 in log-moneyness. A cubic Taylor fit is applied to each daily cross-section and retained only when R^2 is at least 0.95. Dates with |delta u| below 0.001 are removed.
Beta falls monotonically from 1.4375 at 1M (HC3 SE 0.1240, t = 11.593) to 1.0114 at 24M (SE 0.0242, t = 41.829). Regression R^2 remains between 0.79 and 0.82. Eta is 1.9308 at 1M (p = 0.101), 0.3168 at 6M, minus 0.0309 at 9M and minus 0.3623 at 24M. Psi declines from 61.632 at 1M (SE 30.935, p = 0.046) to 0.207 at 24M (p = 0.732).
The joint Wald test of eta = psi = 0 produces 7.33 at 1M (p = 0.026) and rejects at p below 0.001 from 2M onward. Across seven expansion centres, the 1M velocity profile is U-shaped: 1.4248 at K/F = 90.5%, falling to 1.4019 at 95.1%, then rising to 1.5555 at 105.1%. At 12M, the profile decreases monotonically.
On the theory side, the paper proves local preservation of butterfly and calendar arbitrage-freeness under explicit conditions on v, within an admissibility radius for the spot displacement. The authors expressly make no claim of global preservation. They also observe that the SSR ride formula fails to preserve the Gatheral density by pure transport because the functional depends explicitly on k.
The estimator can be coded from the paper
Factorial scaling is specified, so the fitted coefficients equal the ATM derivatives directly. The 1, 1, 2 substitution multipliers are the hierarchy's binomial coefficients. Each test also names its standard-error method: HC3 for the scalar SSR regression, pairs bootstrap with B = 500 for the jets, and moving-block bootstrap with b = 5 for the Wald statistics. The reported block SEs average 16% more, and no coefficient changes significance at 5%. Out-of-sample testing uses 5-fold blocked time-series CV. Forward is defined as F(T) = S0 exp((r - q)T). The estimator itself is reproducible from those details.
The surface is the missing piece. Estimation consumes a fixed 11-point moneyness grid at seven exact tenors, yet we did not find an account of how the grid was built. That judgment enters a2 and a3 before reaching psi.
a3 carries the uncertainty
The authors draw a firm boundary around their validation. In Section 9, they describe the synthetic exercise this way: "it does not constitute evidence that SPX dynamics follow the jet-transport model." The warning applies directly to the coefficient that matters most for our reproduction.
Eleven strikes over about 0.30 of log-moneyness give a cubic little room to identify a third derivative. The paper's diagnostics expose the problem. Beyond plus or minus 5% moneyness, standard errors on the higher local jets exceed the estimates. At the 105.1% centre, v2 = 147.5 with SE 198.3. At 12M, the ATM eta estimate of minus 0.155 is about half the minus 0.31 OLS slope of v0 across the seven centres. The authors attribute the discrepancy to finite-range curvature contamination consistent with psi = 3.92, which bends the profile across the plus or minus 10% moneyness range.
Psi is structurally fragile for the reason the paper gives. Its identifying regressor is a1 times delta u, while ATM variance skew a1 is about minus 0.019 at 1M. The empirical pairs-bootstrap SE for psi exceeds the design-conditional SE by 12.9x at 1M, 7.0x at 3M and 2.7x at 6M. Empirical t-statistics are 2.0, 4.6 and 10.7 over those tenors. The authors advise caution about the 1M rejection.
Their power study matters for any rolling-window reconstruction. With eta = 1.931 at the 1M calibration, the Wald rejection rate at nominal 5% is 0.140 at N = 50 and 0.810 at N = 500. The authors' own 60-day rolling beta at 1M averages 1.4205 with standard deviation 0.3983.
Another unresolved choice appears when the R^2 filter removes dates. Because the estimator differences a0 between consecutive observations, we did not find a rule for reconnecting the difference series after a discarded day.
The strike-constancy test deserves credit. Bootstrap covariance across the seven centres is severely ill-conditioned. Mean pairwise correlation rises from 0.908 at 1M to 0.998 at 24M, with condition numbers reaching 3.0e15. The authors diagnose the problem, treat the full-rank chi2(6) magnitudes as conditioning noise, and substitute a smooth-trend chi2(2) based on linear and quadratic contrasts. They also demonstrate that the naive diagonal statistic is biased toward non-rejection in this setting, with chi2 over chi2_naive around 100 at 1M.
The repaired test rejects strike-constancy from 2M through 12M (10.1 to 32.7, all p at or below 0.006). It does not reject at 1M (3.57, p = 0.168) or 24M (5.46, p = 0.065). The conspicuous 1M U-shape therefore fails the paper's primary test, a limitation the authors acknowledge.
Do eta and psi pay for themselves?
They do on curvature moves. For delta a2, the full model lowers out-of-sample RMSE from 272.8 to 215.7 at 3M (minus 21%) and from 234.8 to 195.5 at 6M (minus 17%), in units 1e-4. The gain is 0.5% at 1M (220.4 to 219.4) and about 4% at 12M.
The full-smile improvement is tiny, and the paper reports the complete row. Full-smile OOS RMSE moves from 3.289 to 3.236 at 1M and from 5.274 to 5.191 at 3M. Results then edge in the wrong direction at 6M (6.759 to 6.761) and 12M (8.599 to 8.615). The variance decomposition accounts for this. Delta a0 contributes roughly 75% of smile RMSE, delta a1 about 17%, and delta a2 only about 6%. Curvature errors receive weight k^2/2 of at most 0.013 on this grid. A 20% curvature improvement translates into about 1.6% overall.
The authors present the distinction openly. Beta governs level risk, while eta and psi govern shape risk, which is "economically material for positions with vanna and volga exposure". Their decomposition supports that interpretation. The evidence ends with regression fits and RMSE on surface increments. The paper contains no hedge P&L and no cost analysis anywhere. Its claim about vanna/volga value follows from the decomposition and is never measured on a book.
A more useful practical result appears in the same table. Adding eta without psi makes out-of-sample curvature prediction strictly worse. RMSE of delta a2 rises from 220.4 to 281.6 at 1M (+28%) and from 272.8 to 318.7 at 3M (+17%). Half the hierarchy performs worse than none. Meanwhile, the random walk benchmark is on average 89% worse than plain SSR on the full smile. The scalar carries most of the forecast.
Our four-leg SPY curvature trade
What follows is our own adaptation of the estimator. The authors' claim remains untested by this exercise. We could not trade SPX fixed-tenor surfaces, so we substituted SPY listed options and used an SPY share delta hedge. Our implementation observes options only at end of day and cannot address intraday smile-response timing. Quoted option bid-ask spreads were unavailable. Costs are therefore assumed rather than measured: 65 cents per option contract, with no additional slippage. The paper reports no P&L and no cost analysis anywhere, leaving our minus 0.06% and minus 0.84 Sharpe without a directly comparable figure.
We estimated beta, eta and psi through the paper's forward substitution, using a rolling 60-observation window ending strictly before each decision date. Those coefficients generated a prediction for the one-day curvature change, and we standardised the innovation using lagged residuals. Whenever |z| reached 2.0, we entered a four-leg structure with two wing legs and two near the money. Entry occurred at the next close and liquidation one day later. Level and skew loadings were solved to zero, with raw vega inside 5%.
Across daily returns from 2 January 2018 to 30 December 2022, the run lost minus 0.06% and recorded a Sharpe of minus 0.84 over 60 trades, with daily volatility of 0.01%. The paper's reported edge is the RMSE improvement already described, 215.7 against 272.8 at 3M, without costs or a trading rule. Our result is a portfolio return net of assumed costs. The quantities differ, so the gap provides no evidence about the paper.
Several implementation choices help explain the result's shape. The gate compressed a 1,118-observation-per-tenor forecasting exercise into 60 trades and selected the largest residuals, which are also especially likely to be fitting artifacts. We imposed fixed-maturity transport equations on nearest listed expiries whose maturity falls daily and jumps at the roll. The measured jet changes consequently omit a maturity-direction term, whereas the paper estimates coefficients from clean fixed-tenor total variance.
We also skipped the R^2 filter, allowing poor cross-sections into the rolling regressions. Snapping to discrete SPY strikes disturbs the grid required by the cubic fit. It adds noise to a1, a2 and a3. Because psi is identified from a1 times delta u, that noise arrives where psi is already weakest. Position sizes were small enough that fixed per-contract commissions could plausibly absorb a large share of gross P&L. Yet the run's profit factor is 0.13, meaning gross winners were roughly an eighth of gross losers. Commissions alone cannot generate that result, and our output cannot distinguish cost drag from a wrong-signed forecast. The estimation window overlaps the paper's July 2021 to July 2026 sample by roughly 18 months.
Our run does not fully explain the gap. It contains no coefficient or forecast-error diagnostics, so we cannot confirm that the transport forecast reproduced before the trading rule was added. This is one automated pass and primarily evidence about our implementation. The pattern resembles an earlier note of ours about a cleanly identified quantity without an instrument that spans it, the correlation rotation premium. The desired exposure is real, while isolating it carries a cost.
I would keep the term structure of eta and psi as a risk report. I would not use it as a trade signal. Beta of 1.4375 on 1M SPX lies well beyond the classical zero-to-one range. A book riding on beta alone at 3M to 6M retains curvature-move forecast error that the full model reduces by 21% at 3M and 17% at 6M. Across total smile RMSE, the same improvement is worth about 1.6%. The case for extra jets belongs to curvature risk.
A hedge backtest on a vanna and volga book at 3M and 6M would change my view, particularly where psi has t-statistics of 4.6 and 10.7. It should use the fixed-tenor surface the estimator was built for and charge real spreads.
Our backtest stops at 2023-01-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.