A 0.045 Sharpe gap over 33 months cannot rank a photonic annealer against branch-and-bound. Each portfolio was selected as the best of 16 joint configurations run for its own pipeline on the same monthly-return window. The paper's tables then reverse the ranking when the skewness term is removed.
Sahoo, Tee and Griffin acknowledge the narrow result in the abstract. Photonic hardware "can locate superior risk-return topologies within a narrow operating range", while mixed-integer programming "remains superior for risk-constrained mandates requiring tight tail-risk control and cross-seed stability". Yet the conclusion says Dirac-3 "demonstrates clear statistical advantages in out-of-sample risk-adjusted returns", and Exhibit 11 calls Dirac-3 (0,1) the "Best overall risk-adjusted configuration". That claim deserves scrutiny. Quantum supremacy is beside the point.
The paper also says "No single pipeline dominates" and describes its mandate recommendations as "starting points for further validation". No t-statistic or standard error for any Sharpe difference appears anywhere. The authors answer the short window with operational controls: composition alerts, HHI thresholding above 0.30, and post-optimization weight capping. But their proposed HHI < 0.25 ceiling would disqualify the recommended configuration because (0,1) has HHI 0.340. We did not find a backtest of the capped version.
One QUBO, three pipelines
The investable universe contains 13 value-weighted, capped US equity anomaly factors from the Jensen-Kelly-Pedersen library: value, momentum, short-term reversal, quality, profitability, investment, low-risk, low-leverage, seasonality, accruals, debt issuance, idiosyncratic volatility and size. The dataset has 164 months of monthly returns and no missing observations. Returns are already net of costs imputed at the JKP level.
Allocation is long-only across the 13 series. Weights lie within [0.00, 0.60], sum to one, and are smoothed with eta = 0.10. The objective combines a rolling volatility penalty beta1, a rolling Fisher skewness reward beta2, and a 0.002 proportional charge on L1 turnover. Any return comes from shifting the portfolio each month toward the anomaly premia favored by the optimizer.
Dirac-3, an entropy-based photonic annealer, and Gurobi share nearly the entire pipeline. A policy network produces an expected-return vector alpha, alongside a 24-month Ledoit-Wolf covariance estimate. Those inputs are assembled into a QUBO, the quadratic unconstrained binary optimization form consumed by both solvers. The solution selects a binary subset of factors. A restricted mean-variance solve then assigns weights to that subset before clipping and renormalization.
SAC, the Soft Actor-Critic reinforcement learning agent, bypasses those stages. It applies a softmax directly to the raw policy action and uses a 6-month rolling window for the penalty terms. Dirac-3 against Gurobi therefore isolates the solver while holding the pipeline fixed. SAC represents a separate investment process.
The chronological 80/20 split leaves 131 months for training and 33 for testing. According to the paper, "all reported performance metrics are computed exclusively over the test window." The abstract instead describes evaluation "across 164 months test window". Those statements differ, and every headline result rests on the 33-month sample.
Can 0.045 Sharpe identify the better solver?
Dirac-3 reaches its global optimum at beta1 = 0, beta2 = 1. Averaged over three seeds, the portfolio records Sharpe 0.760, Sortino 0.841, Calmar 0.567, MDD -3.47%, CVaR5% -1.278%, annual return 1.70% and annual volatility 2.23%. Gurobi's best-Sharpe joint configuration is (0, 50). Relative to that peak, Dirac-3 delivers Sharpe +0.045, Sortino +0.100, Calmar +0.310 and MDD 1.48 percentage points shallower, while CVaR5% is 0.367 points worse.
The other sweep changes the order. With beta2 = 0, Gurobi peaks at Sharpe 0.718 when beta1 = 5, ahead of Dirac-3's 0.715 at beta1 = 2. Gurobi also posts the best CVaR5% in that sweep, -0.983%. At beta1 = 5, the setting where Gurobi peaks, Dirac-3 produces 0.551.
Dirac-3's 0.551 is the toughest figure for the narrow-window argument. The paper's peak-metrics exhibit prints the result plainly, awarding Gurobi Sharpe and CVaR while Dirac-3 wins on Calmar and drawdown. Dirac-3 clears Gurobi at the peak only after the skewness term enters.
We found no standard error for a Sharpe difference, no t-statistic, and no adjustment for the 16 joint configurations searched per pipeline. The full search covers 24 per pipeline across both sweeps and 48 joint runs in total across the three. Reported cross-seed dispersion addresses another question: solver reproducibility. It does not estimate sampling error from 33 monthly returns. A 0.045 gap on that sample lies inside sampling noise. So does 0.003. Both figures were chosen as maxima.
The repeated configurations raise another problem. Identical setting (0, 0) appears in both sweeps with different outcomes: Dirac-3 Sharpe 0.630 against 0.646, Gurobi 0.563 against 0.603, and SAC 0.492 against 0.536. Gurobi's 0.040 discrepancy on the repeated setting nearly matches the headline advantage attributed to the solver.
An accruals bet at the sweet spot
At the Dirac-3 optimum, accruals carries 25.9% and investment 19.6%. Together they make up 45.5% of the portfolio. HHI reaches 0.340, versus an equal-weight reference of 0.077, while CR-5 is 0.895. No Dirac-3 configuration across either sweep has CR-5 below 0.825. Its mean HHI in the beta1 sweep is 0.228, compared with 0.152 for Gurobi.
Monthly L1 turnover over the tabulated beta1 grid ranges from 0.017 to 0.028 for all three pipelines. The paper flags one exception: Dirac-3 at beta1 = 0.5 reaches 0.044. Similar turnover across the three leaves composition as the source of the Sharpe and CVaR spread rather than trading intensity. The tables show it directly.
The authors recognize the concentration. Their monitoring protocol calls for post-optimization weight capping whenever HHI exceeds 0.30, naming (0, 0) and (0, 1) as examples. Their return-seeking recommendation is (0, 1). Capping would materially alter the trade because the 45.5% assigned to accruals and investment generates the Sharpe.
Both of Dirac-3's best configurations in the joint sweep occur at beta1 = 0, with no explicit volatility aversion in the objective. Its primary-sweep window appears at beta1 = 1 and 2.
Gurobi's value appears in tail risk
Gurobi's best CVaR5% in either sweep is -0.863% at (0, 1). Its Sharpe there is only 0.565 and its Calmar 0.190. Across all 16 joint configurations, HHI remains between 0.135 and 0.228, with no factor exceeding 17.7%.
At beta1 = 20, Gurobi's cross-seed Sharpe standard deviation stays strictly below 0.05. Dirac-3 remains below 0.08 throughout its beta1 grid. The classical pipelines received five seeds, while Dirac-3 received three seeds because of per-run photonic cost, as the authors state openly. That imbalance weakens the paper's stability comparison.
The paper's reproducibility recommendation gives Gurobi a return standard deviation of 2.10pp at beta1 = 20. Text discussing the same configuration reports 0.10 percentage points.
These results come from annualized returns of roughly 0.84% to 1.85% with volatility between 1.5% and 3.7%. The paper's best configuration earns 1.70% annually at 2.23% volatility. It reports no capacity or leverage analysis.
Where SAC broke
At beta1 = 0, beta2 = 20, SAC drops to Sharpe 0.088. MDD reaches -14.64%, CVaR5% -2.237%, HHI 0.593 and CR-5 0.985. Under the same setting, Gurobi retains Sharpe 0.617, CVaR5% -1.230% and HHI 0.140. SAC's CR-5 rises to 0.996 at beta2 = 50.
Another collapse occurs at (1, 20), where low-risk alone takes 45.5%, the study's largest single-factor weight. The skewness reward was meant to reduce tail risk. Instead, this configuration produces the deepest drawdown reported anywhere in the paper, -14.64%.
Reward design and constraint enforcement account for the failure. The policy network supplying alpha to the QUBO pipelines belongs to the same class of object. Dirac-3 and Gurobi place a selection stage and bounded mean-variance projection between that network and the portfolio. SAC lacks those layers.
Even when SAC functions, it trails. Its best joint Sharpe is 0.602 at (1, 0.5), and its best beta1-sweep Sharpe is 0.572. The highest nominal return, 1.85% at beta1 = 2, arrives with MDD -12.23% and CVaR5% -2.388%. Gurobi at the same setting returns 1.34% with CVaR5% -1.334%.
Our adaptation, with limits
We had no access to Dirac-3 and performed no quantum sampling of any kind. None of our results tests the photonic claim. We also lacked the JKP factor return library and did not train an SAC comparison. Our adaptation used a classical branch-and-bound solver with the same QUBO structure.
We built six long-only large-cap US equity factor sleeves: value, quality, profitability, investment, momentum and low volatility. The universe comprised the 500 largest non-ADR names. Each month, a cardinality-constrained QUBO selected 2 to 4 sleeves. A bound-constrained mean-variance solve assigned weights with a 60% sleeve cap and 10% stock cap. We applied eta = 0.10 smoothing and rebalanced monthly at the next day's close. Without a trained policy network, alpha was each sleeve's trailing 12-month mean. Costs were 0.002 on stock-level L1 turnover plus $0.004 per share.
Our run covers 2020-01 to 2024-07. It produced Sharpe 0.25, Sortino 0.27, volatility 9.94%, max drawdown -18.67%, Calmar 0.13, total return 11.36%, beta 0.28 to SPY and 4,873 trades. The paper's Gurobi peak in the beta1 sweep has Sharpe 0.718, volatility 1.81% and drawdown -4.33%. Our Sharpe is roughly a third of theirs, with five times the volatility.
The objects differ. Their allocation spans 13 factor return series with annualized volatility from 1.5% to 3.7%. Ours holds long-only equity sleeves. The substitution mechanically produces our 9.94% volatility and -18.67% drawdown, compared with the reported Gurobi MDD of -4.33%.
The paper supplies no calendar start or end date for its 33-month test window, leaving regime overlap with our 2020-01 to 2024-07 run unknown. Two choices in our adaptation push in the same direction. A trailing 12-month sleeve mean is probably much weaker than a trained signal and likely follows leadership into reversals. Selecting 2 to 4 sleeves from six is also much coarser than binary selection across 13 factors, costing us breadth.
Gurobi's HHI ranges from 0.136 to 0.170 in the beta1 sweep, with mean 0.152. Across all 16 joint configurations, it spans 0.135 to 0.228. Our figures measure our implementation. They are one automated pass rather than a verdict on the authors' work.
A repeat of the same grid on a second factor library or a held-out block of months would change my mind, provided Dirac-3 and Gurobi peaks were compared with standard errors accounting for the 16 configurations searched per pipeline. As presented, the paper uses one factor library, follows one historical path, and gives no standard error for any Sharpe difference. Its lasting results are the SAC failure map and Gurobi's HHI band of 0.135 to 0.228. Both are useful. Neither establishes an advantage for quantum hardware.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.