AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

One crisis supports the early-warning case

Halperin's three matrix observables map 2001, 2008 and 2020; only the 2008 AUC of 0.72 arrives early.

2026-08-01 · 9 min read

Reviewing: Observable Matrix Dynamics of Stocks · Igor Halperin · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Observable Matrix Spectral Stress Sector Rotation for US Large-Cap Stocks

Backtest period 2015-01-01 to 2024-12-31 · hypothetical, net of modelled costs

Why these figures are not the paper's (1)

Our own audit found this run does not follow the paper faithfully (9)

  • Paper fixed S&P 500 membership over each ten-year period including two-year pre-roll. (invalidates: Paper full-universe sizes N=244, N=265 and N=309 no longer apply; exact Table 1 values and S&P 500 membership-specific sector attributions no longer apply.)
  • Wide panel of daily total returns. (invalidates: Exact paper correlation, market-share, PR, beta and Perron numerical levels no longer apply.)
  • Sector attribution by summing squared residual leading-eigenvector loadings within sector and dividing by sector share of names. (invalidates: Paper sector identity statements such as utilities, technology, energy, financials and REIT placement may not map exactly.)
  • Single-firm trajectory entropy production Σ_i = sum_t F, with thermodynamic force F_ab = log(μ_ab/μ_ba). (invalidates: Paper single-firm trajectory entropy-production leader results no longer apply.)

5 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2015Total 120.4%Dec 2024
Sharpe
0.69
Total Return
120.4%
Max Drawdown
-33.3%
CAGR
8.2%
Volatility
17.4%
Trades
8,003

What the paper reports for its own strategy

  • Early-warning ROC AUC (the paper's own signal, causal trailing 126-day features, label = forward peak-to-trough drawdown of equal-weight market): 2008 crisis per-period AUC ~0.72 at 63-day horizon; 2020 Covid ~0.49; 2001 dot-com ~0.46; pooled ~0.65 at 21 days (>7% drawdown) and ~0.56 at 63 days (>10% drawdown). Authors state these AUCs are in-sample and per-period, a descriptive comparison rather than out-of-sample forecast. No trading strategy, return, Sharpe or cost assumption is reported.

The forecasting case stands or falls on one figure from one crisis. The rest of the paper describes crises as they happen.

The audit starts with the machinery. Halperin calculates rolling Pearson correlations from daily returns for a fixed S&P 500 universe, then converts them into the angular distance matrix M_ij = arccos C_ij. This puts the names on a sphere. The familiar factor measures follow: market-factor share lambda_1 over the sum of eigenvalues, plus the participation ratio as an effective factor count.

Two more observables come from his earlier work on distance matrices in learning systems. The Perron eigenvalue of M follows the cross section's mean angular separation. He fits a power-law slope beta to the non-Perron eigenvalues over K in [2, sqrt(N)]. Rank-one deflation then removes the market, C_res = C minus lambda_1 V_1 V_1', after which he calculates the same diagnostics on the residual. Two trajectory-level measures compare successive snapshots. One tracks the drift of the top-K eigenspace projector from a pre-crisis anchor. The other takes the Frobenius norm of the commutator between today's correlation matrix and its anchored self.

The other two observables dispense with correlations. Each day, Halperin ranks names by rolling 10-day mean return and, in a separate chain, rolling 21-day volatility. Decile transitions become a Markov chain. Because the rankings are market-neutral by construction, they capture relative dynamics hidden by the market factor. The outputs are persistence and mixing times, entropy production against a detailed-balance surrogate (the arrow of time), and pairwise transfer entropy between sectors and names.

Daily CRSP records for index constituents cover 1996 to 2024. Halperin divides them into three ten-year windows, each centred on a crisis. At the 504-day default lookback, N = 244, 265 and 309. Survivorship is built into the construction: names must remain index members for the entire ten years and the two-year pre-roll. He says this prevents the cross section from drifting within each period. A shorter lookback requires a market-cap cap to hold N below L, leaving N = 94 at 126 days.

There is no strategy, no P&L, no Sharpe and no cost assumption. Portfolio construction appears only in the closing section as future work. An AI coding assistant performed all computation and manuscript preparation under the author's supervision. The code is public, and no independent replication is reported.

What the crisis maps get right

The crisis signatures come through clearly. During Covid, mean pairwise correlation rises from 0.29 in the 2019 baseline to 0.47 in March-April 2020. Market share moves from 0.31 to 0.48, effective factors fall from about nine to about four, and the Perron eigenvalue of M contracts from 394 to 335. The 126-day lookback gives the same episode a harsher reading: market share spikes to roughly 0.73, versus 0.52 at two years, while effective factors drop near two. For the 2008 crisis at 504 days, effective factors decline from 8.3 to 5.3 and market share rises from 0.34 to 0.43.

The 2001 bust goes the other way. At the long lookback, effective factors rise from 22.9 to 25.3 while market share slips from 0.19 to 0.17. Technology pulled away during a dispersed unwind as the rest of the market stayed together. Average correlation would have missed the episode. At 126 days, the signature reverses again, from 16.5 to 14.9 with share flat at 0.20. Window length therefore determines the sign of the fragility indicator as much as the event does.

Rotation is the paper's real addition beyond scalar concentration. Across every crisis, the raw leading eigenvector retains an overlap of 0.93 to 0.99 with its pre-crisis anchor. The market factor gets larger without changing direction. The basis rotation appears in the market-removed block. Its leading sector eigenvector falls to about 0.5 anchor overlap for Covid and below 0.2 for 2008. Meanwhile, the commutator norm rises from zero at the anchor to about 43 and about 30.

Projector drifts remain well under the random-subspace null of roughly 3.1 for K = 5, evidence that the rotation is coherent. Deflation raises the effective factor count from roughly 7-18 raw to 40-85 residual and cuts the leading sector share to 0.06-0.12. Utilities form the most persistently concentrated residual cluster across all three decades.

Lambda_1 cannot provide sector information like this. Name-level attribution makes it usable. All the Covid movers cross their half-shift within about two weeks of onset. The 2001 movers begin shifting months in advance, whereas the 2008 energy cluster reorganises gradually across 2008-2010.

Can 0.72 carry the forecast claim?

The forecast test is causal by construction. It uses only 126-day trailing signals and labels them with the forward peak-to-trough drawdown of the equal-weight market beyond a threshold over 21 or 63 days. At 63 days, per-period AUC is about 0.72 for 2008, 0.49 for Covid and 0.46 for 2001. The pooled results are 0.65 at 21 days and 0.56 at 63. Eigenvector rotation rate scores 0.49, which the paper describes as coincident.

Halperin gives the limitation directly in the summary: "Only the endogenously-building crisis, the 2008 buildup, is forecast in advance by the fragility signals," and the diagnostics are "a strong contemporaneous read of the cross section and only a partial early warning." He also concedes the statistical limit. The AUCs "are in-sample and per-period, a descriptive comparison rather than an out-of-sample forecast, so their regime-dependence is qualitative."

His defence rests on mechanism. An endogenous fragility measure can anticipate a crisis growing from the correlation structure itself, while an exogenous shock such as a pandemic lies outside that mechanism. The explanation hangs together. Yet no episode is held out when scoring the split, and the onset dates used for scoring are fixed and hand-chosen within each period.

The evidence still falls short of the abstract's phrasing. It says "forecast the endogenous 2008 crisis, though not the exogenous 2020 shock". A single in-sample per-period AUC of 0.72 supports that sentence, set against a pooled 0.56 at the same horizon.

Anyone trading an overlay faces the pooled 0.56 at the 63-day horizon, since the regime becomes clear only afterwards. The paper proposes its own remedy as future work: condition on whether the correlation structure is already tightening instead of pooling.

Signal sequence

For the 2008 and 2020 crises, the distance-matrix diagnostics trigger at onset. The rank-chain diagnostics move during the months that follow because irreversibility requires both the fast crash and the slow recovery to enter the window. The 2001 bust breaks the pattern in the distance-matrix rows, where factor count rises into the crisis. Fragility and commutator readings are therefore the timing candidates. The chains supply attribution.

Those chains still produce useful evidence. The return decile chain has a second eigenvalue near 0.86 and mixes in about seven days. The volatility chain sits at 0.97, takes 30 to 40 days to mix and climbs toward 60 after 2008. The same crash accelerates turnover in the performance ordering while making the risk ordering more persistent.

Entropy production is of order 1e-3 nats per step in both chains. Pooled across 1996-2024 and more than a million transitions, the return chain remains within error of the reversible null and lands slightly below it at decile level. The volatility chain reaches z about 8 at the 2002 bottom and 4 in 2007-08. Among single-session events, only Archegos in March 2021 leaves a clear arrow, at z about 5. Quarterly blocks dilute one-day shocks, the paper notes, which partly explains why Omicron and SVB register nothing.

Transfer entropy separates the channels. Utilities lead relative returns through 2008 and Covid, while financials lead relative risk. Truist follows in performance and leads in risk. Within-sector coherence in one-step rank changes runs 0.08-0.10 for returns and 0.03-0.05 for volatility. Persistent relative risk ranks can inform sizing. Seven-day mixing in return rankings offers no holding signal.

The paper is also candid about the failure of its strict distance-matrix theory. Beta stays near 0.7, with no shoulder at K = sqrt(N). Halperin gives three plausible reasons for the missing multiplets and shoulder: Marchenko-Pastur noise when T is of order N, a small universe where sqrt(N) is 16 to 18, and few-factor-plus-noise structure in place of a uniform sample on a hyper-sphere.

An I-BBS finite-N correction explains beta below one. The correction grows as N approaches d squared, causing the implied embedding dimension beta/(beta-1) to turn negative at these sample sizes. Halperin demotes beta to a secondary diagnostic and retains the participation ratio as the main factor count. He then uses the result as framing. The market spectrum remains within the 0.65 to 0.81 range occupied by untrained networks at initialisation, while grokking reaches 1.65 and sparse parity 2.19. The analogy is elegant and supplies no trade.

Our overlay met the wrong objective

We implemented the fragility overlay as a long-only large-cap rotation covering 2015-01-01 to 2024-12-31. Our universe contains the annual top 300 US names by capitalisation, after dropping near-duplicate return columns above 0.97. The spectral panel uses 504 days, supplemented by a 126-day probe on the top 94 names.

Market share, a negative participation-ratio z-score, mean correlation, projector drift and commutator norm feed the stress score. Above 1.0, the book turns defensive at 55% gross. Low-volatility utilities, health care and staples receive 70% of that allocation, with 30% assigned to the lowest-volatility decile. Once the score falls below 0.5 and normalisation persists for 20 days, gross returns to 100%. The top three sectors by 63-day return then receive 60%. Positions use inverse 63-day volatility, with caps of 10% per position and 25% per sector across around 60 names. We charged four tenths of a cent a share and modelled zero slippage.

These figures are ours, not the paper's: total return 120.40%, Sharpe 0.69, annualised volatility 17.41%, maximum drawdown -33.26%, and 8,003 trades. A stress overlay running at 17.41% annualised volatility and surrendering a third of capital at its worst failed at its assigned job.

The paper reports no return, Sharpe or drawdown, leaving no like-for-like result for comparison.

Direction is the closest available comparison. Halperin's fragility signals provide genuine early warning only for the 2008 crisis within his 2003-2012 window, where AUC reaches 0.72. Our test window matches his Covid period, 2015-2024, for which he reports 0.49 at the same horizon. Classification AUC and a net-of-commission equity curve measure different things. Even so, a defensive switch based on signals with no lead in this decade, followed by a -33.26% drawdown through March 2020, agrees with the 0.49 reported for the decade.

The remaining differences come from our choices. We used an annual top-300 close-price panel. His data are constant-membership CRSP total-return panels with N = 244, 265 and 309. Our participation-ratio and market-share levels therefore form a different series, and our regime dates will diverge from his Table 1. We chose the 1.0 and 0.5 thresholds and removed the volatility rank-chain features as mandatory gates.

Our defensive leg also buys the defensive leaders, close to the reverse of the outlook Halperin sketches, where followers are bought when defensives lead. Crowding into low-volatility defensives during a rate-driven correlated selloff preserves substantial equity beta. That choice is the most likely explanation for a drawdown resembling an unhedged book. Realized costs would lower 120.40% and 0.69. One automated pass gives evidence about our overlay first and cannot settle Halperin's work.

The case that remains

As a contemporaneous view of the cross section, the three observables repay the reading time. Deflation plus rotation is the part I would retain. Market direction moves very little while the sector basis rotates coherently below the null, and the changing composition identifies the crisis by industry. The early-warning case amounts to one in-sample AUC of 0.72 from one endogenous buildup, against a pooled 0.56 at the same horizon.

A narrower test could change my view. Run the 126-day fragility classifier continuously over 1996-2024 instead of dividing the data into three decade windows. Score episodes excluded from the definition of the endogenous/exogenous split. Then report AUC where the taxonomy must be assigned before the label appears. Until such evidence exists, this work is best read as a map, albeit a very detailed one.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each trading day t:
  1. Build annual large-cap universe:
     - Select top 300 US non-ADR stocks by capitalization for the year.
     - Require real close prices; exclude names with insufficient lookback data.
     - Drop near-duplicate return columns with pairwise correlation > 0.97, keeping the larger-cap name.

  2. Compute trailing return matrices:
     - Use daily close-to-close log returns through t.
     - Full spectral panel: 504-day rolling Pearson correlation matrix.
     - Fast probe: 126-day rolling panel on top 94 names by cap within the annual universe.

  3. Compute observable-matrix stress features:
     - Market-factor share = lambda_1 / sum(lambda_k).
     - Participation ratio = (sum(lambda_k))^2 / sum(lambda_k^2); use negative PR z-score as stress.
     - Mean pairwise correlation, distance-matrix diagnostics, residual geometry, projector drift, commutator norm.
     - Volatility rank-chain mixing/entropy features are included in the stress score but not required as entry/exit gates.

  4. Classify regime:
     - Enter defensive when stress_score >= 1.0 and spectral concentration confirmations fire.
     - Exit to risk-on when stress_score <= 0.5 and spectral concentration normalizes for 20 sustained days.

  5. Allocate portfolio:
     - Risk-on: deploy 100% gross; allocate 60% to stocks in the top 3 sectors by 63-day equal-weight sector return, favoring positive 126-day momentum and below-median 63-day volatility; allocate 40% broadly by inverse 63-day volatility.
     - Defensive: deploy 55% gross; allocate 70% of deployed capital to low-volatility Utilities, Healthcare/Health Care, and Consumer Staples; allocate 30% to the lowest-volatility decile of the full usable universe.

  6. Size and trade:
     - Weight by inverse 63-day volatility.
     - Enforce max 10% per position, max 25% sector weight, and target at least 60 positions where possible.
     - Rebalance at the close when regime changes or target weights drift beyond the position-cap tolerance.