Count the correlation eigenvalues below the Marchenko-Pastur floor. This paper makes a persuasive case that the count tracks market stress better than the familiar tally above the ceiling. The asymmetry is its main contribution, and the reservations that follow do little to weaken it.
The neglected end of the spectrum
The construction is simple. Start with n assets and a rolling window of w daily log returns. Standardize each column, then form the correlation matrix. Set λ = n/w. For a matrix built from i.i.d. columns, the Marchenko-Pastur law places its eigenvalues inside (λmin, λmax), where λmin = (1 − √λ)² and λmax = (1 + √λ)². Grassi, Pastorino and Uberti call the number below λmin m− and the number above λmax m+. Rolling forward one day at a time produces two integer time series.
Numerical linear algebra supplies the interpretation of m−. When a matrix has many near-zero eigenvalues, it lies close, in norm, to a lower-rank matrix. This is its numerical nullity. In portfolio terms, assets have moved close to linear dependence, leaving fewer effectively independent bets. On the sector dataset, with n = 10 GICS sub-indexes and w = 20, m− ranges from 0 to 7. At the worst points, numerical rank falls from its maximum of 10 to 3. The paper states the implication cleanly: "although the market consists of ten sector portfolios, its correlation structure behaves as if only about three effectively independent investments were available, implying a significant reduction in diversification opportunities."
The paper uses three datasets of daily log returns. Sector sub-indexes run from 1990-01-03 to 2020-09-17. The S&P 500 sample contains 377 of 500 constituents from 2005-01-05 to 2020-09-17, while the Nikkei sample retains 199 of 225 constituents over the same dates. Only names with complete histories remain. A stock that stopped trading in 2008 therefore appears in no correlation matrix. For constituent datasets where n exceeds w, agglomerative hierarchical clustering reduces the universe to k = 10 representative portfolios before m− is calculated.
On the clustered S&P stocks, m− crosses the alert level of 5 during the 2007-2009 crisis, the European debt crisis, the 2015 Greek default, Brexit and COVID. The paper's preferred out-of-sample parameters are w = 1000, λ = n/w = 0.01 and decay factor δ = 0.7. Under that configuration, conditional one-day-ahead distributions widen as m− rises. Sector standard deviation increases from 0.0092 to 0.0180, while 1% V@R rises from 0.0124 to 0.0676 as m− moves from 5 to 9. The paper contains no trading rule, and we did not find a Sharpe ratio, a return, an alpha or a drawdown.
The evidence consists of conditional moments.
Why the market mode contributes little
The upper-spectrum count has almost no value in this application. Across three decades of sector data, m+ takes only 0, 1 and 2, and remains at 1 for the large majority of periods. The authors are direct: "the few situations in which m+ ≠ 1 are not correlated with any particular market behavior." They add: "Even if m+ is a proper measure of connectedness, it does not provide useful information, at least in the present application."
The result has a limited reach. Counting eigenvalues above λmax discards the market mode's size, where the PCA literature preserves much of its information. The two indicators also use different setups. Hierarchical clustering precedes the calculation of m−, whereas m+ follows a PCA and k-fund reduction. The paper acknowledges that the indicators require different reduction methods.
Billio et al.'s cumulative risk fraction offers the stronger benchmark, and its own results compare poorly. In the sector data, the two highest CRF buckets contain 2716 and 3046 observations. Their 1% V@R figures are 0.0342 and 0.027, both below the 0.0385 recorded by the [0.895-0.91] bucket. For the Nikkei, the highest conditional standard deviation is 0.0362 in the [0.9-0.95] bucket, while the top bucket contains five observations.
Scarcity works in m−'s favor. The m− = 9 class contains 673 of 5738 sector observations, 218 of 1953 observations for the S&P stocks and 308 of 1847 for the Nikkei. Those denominators are the column totals in the paper's summary table.
A threshold that moved fifty-fold
The paper presents the Marchenko-Pastur bound as a way to remove arbitrariness from the tolerance ε. The authors also confront the threshold problem. For real data, they write, one must assess whether the theoretical threshold remains suitable for prediction. They therefore fix λ independently of w and examine the sensitivity of the results. Their explanation for the resulting gap is that real financial data are non-i.i.d., making the effective λ smaller than its theoretical value.
At w = 20, the endogenous λ equals 0.5. The paper reports "very limited out-of-sample discriminating power" at that setting, with large gains and losses distributed almost uniformly across the measure's possible values. Discrimination improves when λ is imposed exogenously at 0.17, 0.08, 0.04, 0.02 and 0.01. The strongest panels use λ = 0.01.
The authors then restore endogeneity through the window length, a step summaries often omit. They choose w = 1000, which again yields λ = n/w = 0.01, and add exponential decay factor δ = 0.7 to retain responsiveness. According to the authors, δ = 0.7 roughly corresponds to estimating the spectrum from the latest 20 observations. Every conditional statistic in the summary table uses this configuration.
The threshold therefore behaves as though it had 1000 observations for 10 series, while the spectrum effectively draws on roughly 20. For twenty observations of ten series, λ = 0.5, precisely the configuration that showed very limited discriminating power. Endogeneity returns in the accounting, yet the effective aspect ratio remains about fifty times the nominal value. Elsewhere, the authors carefully observe that linear independence and equal moments are insufficient to produce i.i.d. columns.
The winning λ and δ were selected after inspection of the full 1990-2020 sector sample and the 2005-2020 stock samples. In the out-of-sample section, I did not find a hit rate, a false-alarm rate, an AUC or a t-statistic. Crisis labels are matched ex post against published crisis lists.
The abstract says the paper "validates its effectiveness through comprehensive real data experiments in both" descriptive and predictive settings. Its predictive validation depends entirely on w = 1000, λ = 0.01, δ = 0.7. The authors selected that configuration after seeing the returns used to validate it.
Several influential cells contain very little data. Sector m− = 5 has five observations. S&P stock m− = 5 has one, leaving its standard deviation and V@R unreported. Monotonicity also breaks in places: S&P stock 1% V@R falls from 0.0288 at m− = 6 to 0.0257 at m− = 7, then rises to 0.0453 at 9.
Ten clusters standing in for 377 stocks
The timely version of m− depends on reducing 377 stocks to ten representative portfolios. The authors describe k as "an arbitrary value" and choose 10 to mirror the sector experiment. Preserving the theoretical w/n ratio gives a different result. With w = 754 for 377 stocks, the principal m− peak occurs in 2011, about two years after the beginning of the 2007-2009 crisis. The paper calls this delayed and attributes it to the three-year window, which slows the lower-spectrum indicator's response.
Clustering to k = 10 with w = 20 changes the timing. The series crosses the alert level of 5 at five events. The RMT ratio requirement and the demand for a timely signal work against each other.
The paper describes clustering once, and we did not find a statement that cluster membership is re-estimated within every 20-day window. If clustering uses the full 2005-2020 matrix, observations near the start inherit later information through the feature itself. Anyone extending the work should check this point with the authors.
The false signals are reported candidly. Sector peak 7, S&P stocks peaks 6 and 9, and 8 of 15 labelled Nikkei peaks correspond to no recognized systemic event. The authors partly attribute the Nikkei results to the index not being international enough to represent the global economy. They also set a clear scope boundary: "the present approach is designed for risk detection and is not intended to forecast market direction." Conditional means support that limit, ranging from −0.0015 to 0.0003 across sector m− classes.
Our automated pass lost 221.59%
We converted m− into an allocation rule in one automated pass. The universe comprises the 100 largest US stocks, selected again each year, with 126-day correlation windows. Hierarchical clustering reduces constituents to k of 5, 10 or 20. Entry and exit percentiles use only the previous 252 signal values. When the signal changes state, the book switches at the close between 100% SPY and 100% IEF. Costs are 5 bps one way at the rule level plus $0.004 per share in commissions, with no modeled slippage. Static SPY is the benchmark.
Our results cover 2021-01-01 to 2026-08-10. Total return was -221.59%, Sharpe -0.94, Sortino -1.13, Calmar -0.16, max drawdown -254.13% and annualized volatility 42.07%. The paper reports no strategy return, Sharpe or drawdown, leaving no paired figure for comparison. Its nearest equivalent is the conditional V@R table cited above.
Our implementation contains a scaling discrepancy against the paper's Definition 3. The code compares eigenvalues from a matrix already scaled by 1/126 with a bound divided by 126 again. Consequently, our loss does not test m−.
The book always holds one ETF at full weight. Its volatility and drawdown must therefore come from SPY or IEF rather than a blend. The defensive allocation is a duration bet, useful only when rising synchronization coincides with a Treasuries rally. Neither ETF experienced anything close to 42.07% annualized volatility over the period. An unlevered holding in either fund also cannot produce a -254.13% drawdown, which leads directly back to our code's scaling error.
Three design choices outweigh the reported result. Our sample includes none of the five US stock events caught by the paper's indicator, since every one occurred before 2021. Our λ = k/126 ranges from 0.04 to 0.16, placing it in the middle of the authors' sweep rather than at λ = 0.01, where their forecasting result emerges. We also tested an m−/k fraction variant absent from the paper, and the grid of k, variant and percentile pair was never chosen out of sample.
The signal deserves a cleaner test
The lower spectrum merits the paper's attention. Its demolition of the count-based market mode is a useful negative result. The forecasting claim carries much less weight because the decay factor and window length were chosen after examining the full 1990-2020 sector sample and the 2005-2020 stock samples. We have previously shown how lookback selection can reverse a correlation-based crisis signal outright (/articles/correlation-geometry-as-a-daily-risk-signal).
A single table would change my mind: the conditional V@R ladder alongside a false-alarm rate, using w, λ and δ fixed on 1990-2004 and statistics calculated only on 2005-2020.