Pure beta with a 40-day hold delivers Sharpe 2.05 at 100 bps when the paper's own averaged rows are used. The promoted figure is 2.33. Both appear in the same tables. Their difference is the subject of this review.

The broader result deserves credit. The paper grounds its central claim in the leg comparison across the full grid (Table 4, Figures 1 and 8), rather than the 2.33 row. Its Limitations section also warns against treating the best cells as production parameters. The criticism here concerns the advertised row, rather than a claim the authors declined to make.

How the triggers work

Kargarzadeh, Parvini and the two Khaledians assemble a four-stage process for Russell 2000 equities: universe selection, news to sentiment, joint return-and-risk prediction, and portfolio allocation.

The main finding begins with the tradable cross-section. A stock-side set SS activates when a name's own return z-score reaches |Z| >= 2 over a 120-day window. An indicator-side set SI activates when an indicator reaches |Z| >= 2, the stock's beta to that indicator is greater than 1 in absolute value, and its tail-conditional beta is also greater than 1. The tail beta is calculated when the indicator's own z is above +1 or below -1.

Set operations produce the three legs. Pure alpha is SS minus SI, identifying a firm-specific move without a simultaneous macro trigger. Pure beta is SI minus SS, the anticipatory leg in which the macro variable has moved before the exposed stock. The beta intersection requires both triggers and matching directions. Each leg retains the top 50 names by in-sample BUY frequency. Kargarzadeh's 2024 MSc thesis, cited by the paper, supplies this decomposition.

Daily OHLCV covers October 2, 2023 to December 31, 2025. Scored news starts on October 1, 2023 and includes 451,854+ articles. The indicator set contains 58 exogenous series spanning January 2022 to January 2026. Yahoo Finance supplies fifty, among them 11 GICS sector ETFs, VIX, commodities, rates and crypto. Another 8 are FRED releases. The in-sample period ends on December 31, 2024. Its first 80% of trading days are used for training and the last 20% for validation, with a holding-period embargo at the split. The backtest is the 2025 calendar year.

The highlighted configuration combines pure beta, GPT-4o mini sentiment, a Student-t target, a 40-day hold and risk parity. At 100 bps plus fixed 5 bps one-way slippage, it reports Sharpe 2.33, annualized 95.9%, cumulative net 100.1% and a max drawdown of -18.3%. In the 50 bps 40-day allocator chart, the Russell 2000 is marked at Sharpe 0.56.

We could not run the strategy on our data. Point-in-time Russell 2000 membership is unavailable to us, as are the FX, commodity-curve and full world-index panels needed to reconstruct the indicator trigger as written. We would also need to replace the sentiment scorer with our own, either FinBERT-style or rebuilt from their code, and retrain the covariance network from scratch. Such a run would examine a different configuration from their GPT-4o mini and trained model. The analysis below therefore reads their reported figures and linked code.

What "best" costs

The 2.33 result is the maximum across four backends, two target distributions and six allocators, using the horizon and cost setting selected for that row. The paper also publishes averages. Table 4 chooses the best allocator at each horizon after averaging across backends and targets. Its 40-day pure-beta entry has Sharpe 2.05 and annualized 74.7%. The averaged panel in Figure 13 likewise gives risk parity a Sharpe of 2.05.

A different averaging order sharpens the comparison. Averaging across allocators and targets produces backend results of 1.76 (GPT-4o mini), 1.76 (FinBERT), 1.77 (Llama-2-13B) and 1.67 (Mistral-7B). The selected cell is therefore 2.33 beside a backend mean of 1.76.

The best 100 bps result for the intersection leg is Sharpe 1.01 (FinBERT, Student-t, 20 days, MVO). Once backends and targets are averaged, the same MVO cell falls to -0.33. Every allocator is negative, ranging from -0.15 to -0.86. Annualized returns span -25.9% to +0.2%, while every backend mean falls between -0.40 and -0.57.

The authors acknowledge the search issue directly. They say the best cells should be interpreted as economically motivated evidence rather than stable production parameters, and annualized figures summarize one test year. Before treating any horizon and selection pair as a persistent anomaly, they call for a longer live sample, a formal multiple-testing adjustment and event-level attribution. This review adds the arithmetic behind that request: the averaged 2.05 and the intersection results that turn negative under averaging.

Their case for the horizon pattern rests on its survival across the cost grid. The pattern reappears at 50 bps, while pure alpha's 60-day advantage grows as explicit costs rise from 0 to 100 bps. That supports the 60-day pure-alpha result. The same defence does not support the 20-day cell, where pure beta leads by 0.09 against 0.06 for pure alpha.

Three hundredths of a Sharpe on one year of daily data.

Why does the intersection lose?

SI minus SS seeks spillover before the stock responds. SS minus SI seeks slow firm-specific repricing. The intersection removes both effects.

The masking rule offers a mechanical explanation. Whenever no name has an active BUY on a rebalance date, the portfolio remains in cash until the next signal. By construction, the intersection is the narrowest set and should therefore own fewer names while carrying more cash. Exposure appears among the paper's reported metrics, though I did not find leg-level exposure or name counts in its tables and figures. I cannot distinguish weaker selection from a thinner portfolio with heavier cash holdings. That distinction governs transferability. A desk using a fallback book whenever the intersection is empty would be trading a setup the paper did not measure.

Missing ablations

Each of the six allocators receives the same model moments. This comparison isolates allocation choice while leaving the source of covariance unresolved. The paper names the relevant ablation itself: at S = 1, epistemic covariance becomes exactly zero, yielding a model that does not account for its own uncertainty. I did not find that run in the reported results. Nor did I find a historical-covariance arm or a no-news arm.

Allocator rankings create another problem. MVO directly optimizes against the predicted mean and covariance, yet finishes last in both leading legs. For pure alpha at 60 days, MVO records 0.70 against 1.87 for equal weight. Black-Litterman reaches 1.76 in the same cell. For pure beta at 40 days, MVO produces 1.16 versus 2.05 for risk parity.

The winning row ignores the predicted mean and allocates through the predicted correlation structure. Sentiment still enters because the covariance head uses the same news-plus-price representation as the mean. The measured dispersion answers how much each component matters. In the pure-beta 40-day cell at 100 bps, Sharpe ranges from 1.67 to 1.77 across the four backends, compared with 1.16 to 2.05 across the six allocators.

Those narrow backend differences persist despite class mixes of 10/21/69 positive/negative/neutral for Mistral-7B and 35/16/49 for FinBERT on the same articles. The abstract says stock-selection regime and allocator choice matter at least as much as the sentiment model. In this cell, backends account for 0.10 of Sharpe dispersion and allocators for 0.89. For the advertised row, the language model looks closer to decoration than the source of the edge.

Inside the rest of the system

Stage two assigns sentiment to a story rather than an individual article. Within a trailing 30-day window, article embeddings undergo single-linkage agglomerative clustering at cosine similarity 0.90. Only the article nearest each cluster centroid is scored. GPT-4o mini is the primary scorer, with FinBERT, Mistral-7B-Instruct and Llama-2-13B-Chat as alternatives. The output then receives an entity-prior correction, trailing winsorized demeaning and a group-day cross-sectional z-score. Articles published at or after 20:00 UTC move to the next trading day.

Stage three uses a two-branch 1D-CNN with 32 channels, kernel 3, a 64-dim hidden layer and a 30-day lookback. The network has a mean head and a rank-2 low-rank-plus-diagonal covariance head. Training minimizes Gaussian or Student-t (nu = 5) negative log-likelihood. Dropout of 0.2 remains active during inference for 50 passes. Dispersion among the predicted means becomes an epistemic covariance term, which is added to the aleatoric term.

Stage four passes the total matrix to six allocators. MVO uses delta = 2.5, a 40% position cap, long-only constraints and an optimizer turnover penalty set to zero. The other choices are equal weight, risk parity, hierarchical risk parity, Black-Litterman at tau = 0.05 and Bayesian Black-Litterman. According to the authors, the final two differ only in the allocation step applied to the same posterior. The paper claims two contributions: predicted risk as a model output instead of a historical estimate, and the story-clustering sentiment pipeline.

A forward test

I would freeze pure beta, 40 days, risk parity, GPT-4o mini and Student-t, then run 2026 forward beside the S = 1 and historical-covariance arms. Until such evidence arrives, the concentration deserves caution. Fifty long-only small-cap names, under a 40% single-position cap, generate roughly 96% annualized with an 18.3% drawdown. Trading costs are flat basis points charged on L1 turnover at segment end, rather than per-name spreads.

The authors further state that a 30-day holding period was never tested. This leaves an unobserved interval between the 0.09 result at 20 days and the 2.05 result at 40 days. At 100 bps, pure beta over one and two days records -2.15 and -2.44, with the latter annualizing at -150.7%. Pure alpha returns -1.28 and -0.15 in those same cells. No horizon shorter than 5 days survives 100 bps.

The selection decomposition is the idea most likely to transfer. Its honest price is the 1.67 to 1.77 range of backend means and the 2.05 averaged best allocator at 40 days, tested in a period other than 2025.