RIEnet delivers the risk result. Bongiorno and Villassero report 11.17% annualized five-session volatility, versus 14.08% for the nearest competing estimator. Both figures use the same incomplete panel and include simulated commissions, fees and market impact. A Model Confidence Set removes every rival at the 0.1% level. The paper names five-session realized volatility as its primary endpoint, matching the training horizon and rebalancing frequency, while CAGR and Sharpe are secondary. The abstract goes further. It groups realized risk, risk-adjusted performance and drawdown control as if a single test supported all three.
Why incomplete panels matter
Requiring 1,200 clean daily returns before a stock enters the optimizer removes 17.15% of stock-rebalance observations. The shorter histories extend deep into the sample: 8.45% contain fewer than 600 valid returns, while 3.48% contain fewer than 252. A mandate still has to hold recent listings and names returning from temporary suspension.
Pairwise-complete estimation retains them. Each correlation entry draws on the longest overlap available for its pair, leaving different cells estimated from different samples. The resulting matrix can contain negative eigenvalues, which imply negative variance for some portfolios and make the matrix unusable in a minimum-variance optimizer. Higham-style projection onto the positive-semidefinite cone restores definiteness without removing sampling noise. Ledoit-Wolf quadratic inverse shrinkage assumes one rectangular panel with a defined observation count and degrees of freedom, neither of which pairwise estimation supplies.
Bongiorno and Villassero replace an analytical spectrum model with a network. They eigendecompose the potentially indefinite pairwise correlation matrix and retain the signed eigenvalues. Each spectral factor also receives an effective sample length, calculated as a weighted average of pairwise overlap counts using squared eigenvector loadings. With no missing observations, the measure collapses to the panel length.
Tokens containing (eigenvalue, sqrt of n, sqrt of that effective length, factor concentration ratio) enter a bidirectional GRU in eigenvalue order. It has 32 hidden units per direction, followed by a softplus head that produces a positive inverse eigenvalue.
Negative inputs still produce positive outputs.
The eigenvectors are rescaled until the reconstruction has unit diagonal, making the finished matrix positive definite by construction. Marginal volatilities go through a separate eight-unit MLP that receives each asset's own history length. The full estimator has about 7,400 parameters, regardless of asset count.
Training targets the realized five-session variance of an unconstrained analytical GMV portfolio. Live deployment instead uses a long-only GMV optimizer absent from the loss. The same optimizer runs across the seven covariance estimators, and an equal-weighted book serves as the allocation-free benchmark.
The test comprises Twenty-six expanding windows. The first uses 1990 to 1999 for out-of-sample year 2000. The final window uses 1990 to 2024 for 2025. At each date, the universe is the point-in-time top 1,500 U.S. names by capitalization, restricted to prices between USD 10 and 2,000.
Execution is simulated through a closing-auction broker. The model imposes integer-share sizing, IBKR Pro tiered commissions, SEC and FINRA fees on sales, and financing at Fed Funds plus a broker spread. Market impact shifts the execution price according to the square root of order size divided by strict-prior ten-day average volume. Goyal, Jegadeesh and Wu supply coefficients of 0.45%, 0.74% and 1.82% for NYSE-like large, small and micro caps. Nasdaq adds 0.57, 0.51 and 2.66 percentage points.
Does the comparison favor RIEnet?
RIEnet has access to more observations than QIS and MLE. Those estimators run on the complete-row subpanel available at each signal date, a choice the paper states and explains. Pairwise estimation, nearest-correlation projection and factor imputation produce matrices outside the sampling model behind QIS. Using RMT after those steps would require a separate statistical derivation. The 11.17% against 15.54% RIEnet-versus-QIS result therefore combines estimator quality with a different information set. It cannot carry a clean shrinkage comparison.
The incomplete-panel methods provide the useful test because each sees the same mask. Factor imputation records 14.08% and a 0.580 Sharpe. Bootstrap reaches 14.46%, Anderson 14.74% and Ridge 18.07%. RIEnet beats every one using like-for-like inputs.
The universe filter also deserves credit. Stocks whose 5- or 20-session log-volatility falls below the cross-sectional lower 1.5-IQR bound are removed. Near-zero-volatility stocks would mechanically improve a GMV objective, so excluding them makes the contest harder.
The confidence set has limits
The body defines the test narrowly. The authors "assess the statistical significance of the realized-risk differences" through the Model Confidence Set of Hansen et al., using a stationary block bootstrap. Ten thousand replications are run with a baseline expected block length of 11, plus sensitivity values of 5, 10, 20 and 40 blocks. Resampling leaves the volatility ordering intact at the 0.1% level.
The 9.30% CAGR carries no such test. Neither does the 0.814 Sharpe compared with 0.580 for Factor. Both cover the continuous out-of-sample span from January 4, 2000 to December 30, 2025, totaling 6,537 sessions, and we found no standard errors for either difference.
Drawdown, the third result grouped into the abstract, receives no test. RIEnet posts a maximum drawdown of -41.3%. The other covariance estimators range from -49.4% to -70.0%, while the equal-weighted book reaches -58.5%. Those are wide separations observed on one path, across twenty-six fits, with no reported seed variability for the network.
The table does answer the low-beta objection. RIEnet's five-session beta to the Russell 1000 is 0.500, compared with 0.592 for Factor and 1.125 for the equal-weighted book. A beta of 0.500 accompanies the highest CAGR in the table, 9.30%.
One fixed RIEnet specification is evaluated from portfolio construction through execution. The paper assigns the improvement to no individual component. Its lag transform, overlap conditioning, GRU and volatility MLP receive no separate tests. The authors defer attribution to synthetic benchmarks, where population covariance and the missingness mechanism are known. That setting can isolate architectural effects without repeatedly reusing the historical backtest.
The reasoning has merit because an ablation on the same 2000 to 2025 path would mine data already used once. Readers are still left unable to distinguish the contribution of effective-sample-length conditioning from the result of running a GRU on signed eigenvalues alone.
What we could reproduce
The paper describes the estimator well enough to implement. Its reported experiment is beyond our available data. Our training history does not extend to 1990, and our capitalization screen is annual, leaving a daily point-in-time top-1,500 ranking unavailable. Every figure reported above belongs to the authors.
Costs and capacity
RIEnet's median effective breadth is 57.6 names, with annual turnover of 31.53 times. Each dollar traded incurs 2.57 basis points of impact and 0.92 of fees. The estimated CAGR drag is 1.49 points.
Factor turns over 19.24 times and pays 12.33 basis points of impact, producing 3.06 points of drag. QIS, the least expensive optimized book, loses 1.17 points while spreading exposure across 754.1 effective names. The paper acknowledges that RIEnet is relatively expensive to operate. Breadth, turnover and order concentration jointly determine execution costs. The direction of the trade-off appears clearly in these rows: 19.24 annual turns at 12.33 basis points, versus 31.53 turns at 2.57.
The simulation begins with USD 1 million and compounds thereafter. A 9.30% CAGR over 26 years grows the account by roughly an order of magnitude. Market impact is imposed after optimization, using a square-root order-size model with fixed coefficients by group. The paper reports no run from a larger initial account. Before sizing the strategy, I would want the impact bill at multiples of that starting capital. The micro-cap coefficient reaches 1.82% for NYSE-like names, with Nasdaq adding 2.66 percentage points.
We have previously argued that better-conditioned covariance can explain most of a reported Sharpe improvement (our MINGLE note). This paper offers a cleaner test of that claim. One long-only GMV optimizer and one execution model remain fixed across all seven covariance estimators. The equal-weighted book remains outside that group as the allocation-free benchmark.
A component ablation and a second path would make the case convincing. The second path could come from a different market or from reported seed dispersion on the same one.