Carvalho has found a real state variable and attached it to the wrong premium. The rotation index is unspanned by the level of the traded correlation surface, at most 6.7 percent, though its correlation with VIX is 0.316 and the slope of that surface does react. It forecasts the broad volatility environment two to three quarters out. The object it prices is a one-month implied variance against a twelve-month equal-weighted constituent realized leg. Match the legs and the coupling goes away. The mismatch is disclosed in the data section and again in the limitations, and the interesting question is what survives the disclosure.
The measurement
Every priced measure of co-movement in the literature is a function of correlation eigenvalues: index variance, average pairwise correlation, option-implied correlation, the absorption ratio. Carvalho's point is that the eigenvectors carry a separate piece of information, namely which firms co-move rather than how much, and that nothing prices their motion.
The construction. Estimate a twelve-month correlation matrix on S&P 500 constituents each month, work in the dual space because T = 12 is far smaller than the roughly 421-name effective cross-section, pull the loadings on modes two through four into stock space, orthonormalize to a three-dimensional subspace. Then measure the principal angles between this month's subspace and last month's, restricted to names present in both, and take the mean squared sine. Call it REC, the reconfiguration index. The market mode is dropped by construction, which is the choice that makes the whole thing work. The paper cites a market-inclusive rotation measure that correlates 0.42 with the irreversibility of a regime-aware ranking chain and near zero with its market-neutral counterpart, which is why that version tracks stress (Halperin, 2026b).
The reading is exact rather than metaphorical. One minus REC is the share of last month's co-movement structure retained. Sample mean 0.194, standard deviation 0.119, range 0.005 to 0.526 with the maximum in July 2002. A typical month rewrites a fifth of the classification and carries four-fifths forward, about a 26-degree tilt of the subdominant frame.
Data: monthly total returns, June 1994 to December 2025, including delisted names, filtered to a 379 x 430 panel of 159,742 observations. That yields 368 windows, 367 rotation readings, 365 regression observations.
The headline. Regress the log premium (one-month VIX squared minus twelve-month equal-weighted constituent realized variance) on the standardized three-month average of REC, with realized volatility and average pairwise correlation as controls, Newey-West at twelve lags. Beta +0.195, standard error 0.036, t = +5.40, R-squared 0.504. On the raw unsmoothed index, t = +4.02. Both level controls enter negatively and significantly, so the coupling is identified against volatility rather than through it.
What it is supposed to buy you: the premium is prepayment. On impact, the whole association sits in implied variance (t = +4.57) and none in realized (+0.62). Forward realized variance builds to a peak coefficient of +0.155 at eight months (t = 3.65), and the premium itself goes from +5.12 at h = 0 to +0.22 at three months and troughs at -2.68 at eight. Money in advance for turbulence that then shows up.
The two premiums
The paper's dependent variable mismatches its legs in horizon (one month against twelve) and in weighting (cap-weighted index implied against equal-weighted constituent realized). Carvalho states this as a scope condition in the data section rather than a footnote. He also reruns a companion battery to bound it. The same smoothed index gives |t| <= 1.2 against horizon-matched forward variance-swap payoffs at one, three and six months (-0.08, +0.74, +1.19), and |t| <= 1.9 against forward one-month cap-weighted index realized variance at every horizon tested.
Those two lines do most of the work in this review. The 5.40 and the -0.08 measure different objects. REC prices the spread between short-dated index fear and the slow, broad volatility environment of the persistent large-cap cross-section. It does not price the one-month variance swap.
The author's defence of that is better than it first looks, and it deserves to be quoted whole rather than clipped. The paper says the non-porting result "is not a weakness of the result but a condition of its consistency with the no-alpha boundary": a state variable that forecast the variance-swap payoff would be a trading signal, and this one forecasts the environment against which the premium is written. Internally that is coherent. The headline t of 5.40 still attaches to a quantity no desk has a book in, and the paper says as much twice: equation (2) "is not the canonical variance risk premium", and the harvest payoff "is not the payoff to a tradeable variance swap". The abstract reports the coupling to the aggregate variance risk premium at t = 5.40, with no scope condition attached to the word aggregate. The disclosure sits in Sections 2, 6.4 and 8.4 instead.
And the convention matters too. R-squared is 0.504 in logs against 0.149 in levels. October 2008 through December 2009 accounts for 46 percent of the total squared deviation of the levels premium about its mean. REC's own largest readings are March 2022, October 2018, July 2010, April 2000 and May 2020. Those months sit against levels premia between -0.015 and +0.048, versus a crisis maximum of +0.299. Rotation prices ordinary variation in the premium, not its extremes. Carvalho says this outright and links it to the no-crash result.
What the estimator survives
I expected this measure to be a laundered version of correlation intensity, and the paper anticipates that. Spectral gaps enter with the negative signs Davis-Kahan predicts (t = -3.14 upper, -2.15 lower). The full mechanical model explains only R-squared 0.121 of the index's variance. The pricing coupling holds at t = +4.92 with the whole battery as controls. Lower-gap quintile couplings run +2.02, +3.25, +3.47, +3.26, +0.97, no monotone pattern, and the calmest variance-share quintile still gives +3.96. The sorts also show where it is absent: +0.77 in the middle correlation quintile, +1.67 and +1.76 in variance-share quintiles two and three. The paper prints those.
Orthogonality to the level gauges holds. Across the correlations of the index with level gauges, no cell exceeds 0.32; the largest is the smoothed index against VIX at 0.316. Realized volatility is -0.031 raw, average pairwise correlation -0.056. Projecting REC on COR3M (Cboe's three-month implied correlation index) and COR1M recovers R-squared 0.059 on the raw index and 0.050 on the smoothed, over 240 months from 2006. Adding DSPX (the S&P 500 dispersion index) gives 0.052 raw and 0.067 smoothed over 139 months from 2014. Between 5.0 and 6.7 percent spanned, depending on smoothing and sample. Meanwhile the shape of that same surface reacts: the COR3M minus COR1M slope loads on smoothed REC at t = -3.46, beta -0.910 index points per standard deviation against a mean slope of +3.409. The level instruments do not carry the state variable, and the paper gives a Rayleigh-quotient argument for why: for the average-correlation functional the median squared overlap with the market mode is 0.750 against body-mode overlaps of 0.013, 0.004 and 0.002.
Smoothing is not load-bearing. t(REC) across averaging windows of 1 to 24 months runs 4.02, 4.00, 5.40, 5.46, 5.34, 5.03, 4.06, 3.69, 3.28, 3.09, 3.64, and the adopted K = 3 sits below the K = 4 and K = 5 maximum. Panel-filter sensitivity spans t of +2.44 to +6.11 across seven threshold pairs, again with the adopted pair not the maximum. The K = 3 subspace, the deflated-edge discard rule and the three-month window remain in-sample conventions. Carvalho concedes that part of the K = 3 to K = 5 decline may be subspace-dimension mechanics rather than bulk dilution.
One place the inference genuinely wobbles. A block-pairs bootstrap is the natural first choice for overlapping regressions. It raises 95 percent critical values to 3.63-5.59 and kills every horizon. The paper reports the failed design and argues that blocks preserve the lead-lag structure that is the alternative, so the null absorbs the effect. I find that argument persuasive. The forward result still stands on the i.i.d.-pairs null, and the h = 10 verdict flips between AR(2) and AR(4).
What rotates, in desk language
Mode two is labelled utilities in 63.6 percent of the 368 windows and energy in 17.7 percent; mode three is diffuse in 52.2 percent, mode four in 71.2 percent. One axis does most of the turning: median principal angles are 5.9, 11.8 and 42.8 degrees, and in March 2022 they were 7, 10 and 71. Two axes hold, one turns.
So the priced thing is drift in which names load on a duration-and-defensives axis, with a commodity axis as the frequent alternative. Discrete label switching does not absorb it: a switch-intensity series correlates +0.187 with the raw index and leaves it at t = +4.89 against the switch series' +1.93. Carvalho's own framing is loading drift, identified by elimination rather than direct measurement, and he flags that as a limitation. Translated: your hedge ratios and your sector buckets are being re-estimated by the market underneath you, and the pace of that re-estimation over a quarter is what carries the premium. The monthly innovation carries nothing (t = +0.30) while the three-month average carries everything (t = +4.92). The premium pays for pace, and one-month rotation drives out the twelve-month version, 4.76 against 1.38.
One more number that stuck with me. Subspaces 24 months apart still retain 0.0664 alignment against a chance level of 3/421 = 0.0071, verified by 3,000 simulated pairs. Nine times chance, and the profile is flat from twelve months out. A permanent core plus a transient with roughly twelve-month memory, and the forecast dies at h = 11-12, exactly where the subspace has decorrelated.
Where the money isn't
The attribution table is the part a variance seller will read first, and it is the part Carvalho most insistently labels as not a backtest. Selling variance at VIX squared against the paper's own realized leg, July 1998 to November 2025, 329 months, with transaction costs, margin and the VIX-squared convexity correction explicitly ignored and the payoff non-tradeable. Mean monthly payoff by real-time REC quartile: +0.0040, +0.0102, +0.0092, +0.0410. Capture rates 11.1, 27.4, 24.2, 52.7 percent. Sixty-four percent of 27 years of premium sits in the top rotation quartile, and that quartile is also the safest: loss probability 0.15 against 0.29, worst month -0.032 against -0.099. Compensation up, risk down, across the same sort. Whatever that is, it is not a risk exposure.
And it converts into nothing. The pre-registered linear overlay gives Sharpe 1.36 against 1.49 unconditional, paired t = -1.85 at equal volatility, against the overlay. On the 293-month real-time-threshold sample: unconditional 1.28, linear tilt 1.22 (t = -0.81), step tilt 1.25 (-0.94), top-quartile-only 1.02 (-2.23). Split in half, the step form gives -0.97 on 2001:07-2013:08 and -0.07 on 2013:09-2025:11. Over the 329-month payoff sample the worst 5 percent of months carry a mean standardized REC reading of -0.47. The paper states that both the 1.49 and the 1.36 are inflated by the payoff definition and should be read only as a relative comparison.
Can our adaptation settle anything?
What we could not do first. The original point-in-time S&P 500 constituent panel with delisted-name coverage from 1994 is not available to us, so we cannot rebuild the paper's 430-name survivor-tilted universe. Our coverage starts around 2010. The Cboe COR1M, COR3M and DSPX spanning tests are not available as specified, so those validation results are untested here. Most importantly, we did not test the paper's claim at all: we built a risk-overlay adaptation, while the paper's stated result is a variance-risk-premium attribution and it explicitly reports no timing alpha and no crash protection. We also traded SPY against cash rather than an index book or a variance swap.
What we ran instead. REC on a yearly point-in-time top-500 US-equity universe selected by capitalization, monthly returns, twelve-month windows, modes two to four, three-month smoothing, expanding standardization with a 36-month burn-in clipped to [-2, 2]. Then a gate. We recursively forecast one-month SPY realized variance with VIX plus lagged realized variance, and again with standardized persistent REC added. The two forecasts are scored against each other with QLIKE, a variance-forecast loss function. The overlay only acts once the augmented model's mean QLIKE beats the baseline on at least twelve completed forecast pairs. When active, SPY weight is clip(1 - 0.5 x max(z, 0), 0, 1) against cash at 0 percent. One basis point per one-way notional plus platform commissions, executed at the first SPY close after month-end. Backtest 2020-01 to 2024-07, 55 trades.
Our figures, from our run: Sharpe 0.73, max drawdown -36.5 percent, beta 0.91, annualized volatility 20.70 percent. The paper's own reported Sharpes are 1.49 unconditional and 1.36 for its scaled overlay over 329 months. The two sets measure different instruments: theirs is an uncosted, non-tradeable short-variance carry series on its own realized leg, ours is a long-only SPY-versus-cash equity overlay net of costs. The difference in level tells you about the change of traded object, the sample and the exposure rule. It tells you nothing about whether REC prices the paper's premium.
The beta is the number that explains our result.
At beta 0.91 and 20.70 percent annualized volatility, the overlay carried essentially full equity risk into a -36.5 percent drawdown through the 2020 crash. Part of that is structural in our own design. The twelve-month lookback, the three-month smoothing, the 36-month burn-in and the twelve-pair QLIKE gate consume most of the 2020-01 to 2024-07 window. So the overlay is likely sitting at 100 percent SPY for a large share of the run. That is a choice of ours, and it is the first thing I would change. Our forecast target is also cap-weighted SPY variance, and the paper reports |t| <= 1.9 for its own index against exactly that object, so a gated-off or ineffective overlay is what its evidence predicts for our construction.
One automated pass built from a paper's description is evidence about our implementation before it is evidence about anything else. What I will say is that the qualitative shape agrees with what Carvalho pre-registered: full retained beta and an unshortened drawdown are what "no timing alpha" and "no crash protection" look like once you wrap the signal in something tradeable.
What would change my mind
A direct measurement of loading drift on a churning cap-weighted panel. The index as built restricts each comparison to names present in both windows, so entering and exiting names contribute no rotation, and the identification of loading drift proceeds by elimination rather than measurement. Run the same construction where membership actually turns over, and measure the drift in loadings directly, and the mechanism claim stops resting on two ruled-out channels. The horizon-matched battery has already been run and returned |t| <= 1.2 against forward variance-swap payoffs at one, three and six months, so that door is closed. The daily-frequency version is closed too: the paper states that daily-sampled subspace geometry is a distinct object and that the companion paper finds it unpriced. As it stands the paper has established something worth knowing. The eigenvector orientation of the equity cross-section is priced information that the level of the traded correlation surface spans at most 6.7 percent of, and Carvalho has been unusually careful about the boundary of that claim. We have written before about how much of a covariance-model edge survives contact with a tradeable wrapper, in our note on MINGLE's covariance swap.
Read it as a correlation-regime diagnostic. Do not size off it.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.