The paper's 0.616 Sharpe comes from the ICA fit, not from swapping CVaR for EVaR. Within its multivariate normal tempered stable model, that substitution barely changes the portfolio: gross Sharpe is 0.555 against 0.552 over twenty-six years, with a bootstrap p-value of 0.868.
Choi presents three contributions: the cumulant derivation, the portfolio constructions, and the out-of-sample comparison. The derivation is the lasting one. Its value does not depend on the empirical case for EVaR.
EVaR uses the Chernoff inequality to give a coherent upper bound on VaR and CVaR. It is the infimum over u > 0 of (1/u)ln(M_L(u)/eta), with M_L denoting the loss moment-generating function. Putting it into a portfolio optimizer creates two problems. The moment-generating function must remain finite, while both the fitted parameters of portfolio returns and the admissible range of u change with every candidate weight vector. A naive parametric method consequently refits a univariate distribution to the portfolio series at each trial weight.
Choi chooses multivariate constructions that produce the portfolio cumulant from asset-level or component-level fits. The fitted families are tempered stable Lévy processes, pure-jump processes with Lévy triplet (0, nu, mu), where exponential damping of the jump measure permits skewed heavy tails. The paper uses three: classical tempered stable (CTS), normal tempered stable (NTS) and multivariate NTS (MNTS).
Under MNTS, closure under linear projection yields the projected parameters directly. Alpha and theta remain unchanged, beta_bar = w'beta, gamma_bar = sqrt(w'Sigma w), and mu_bar = w'mu. The admissible upper endpoint becomes u_max(w) = (beta_bar + sqrt(beta_bar^2 + 2 theta gamma_bar^2))/gamma_bar^2. The ICA route instead applies FastICA, retains all N components, and fits each with a tempered stable law. It then obtains the portfolio cumulant by adding the component cumulants and an affine location term. For the inner one-dimensional infimum, the method truncates the admissible interval by epsilon, divides it into P subintervals, runs Brent's bounded minimization on each, and evaluates the finite endpoint separately.
The economic premise is straightforward. Sector returns have skew and heavy tails, so a measure that penalizes the full exponential moment could leave tail-heavy sectors sooner than variance or an average-beyond-a-quantile measure. Choi tests that premise through thirty-one portfolios formed from the eleven Select Sector SPDR ETFs. Nine have traded since December 1998. XLRE arrived in 2015 and XLC in 2018, expanding the investable universe from nine names to eleven.
Each fit uses Twelve months of daily returns, followed by a monthly rebalance. Portfolios are long-only and fully invested, with no target-return constraint. Prices are Yahoo adjusted prices, and the out-of-sample period runs from January 2000 to March 2026. Sharpe differences use the Ledoit-Wolf studentized circular block bootstrap with 4999 resamples and block length floor(T^(1/3)).
ICA plus NTS minimum EVaR supplies the best cell. It reports gross Sharpe 0.616, CAGR 8.60%, annualized vol 15.28%, cumulative return 766.96%, max drawdown 38.30%, and Calmar 0.224. Equal weight returns 0.537 with a 53.51% drawdown. Minimum variance reaches 0.559, while maximum Sharpe posts 0.476.
How much comes from EVaR?
Very little under MNTS. Holding the model constant, the EVaR-minus-CVaR Sharpe gaps are +0.003 for minimum risk, +0.001 for STAR, +0.002 for symmetric Rachev and +0.003 for the asymmetric version. Every p-value is at least 0.868. Once projection reduces the system to a single univariate NTS, replacing CVaR with its coherent upper bound produces no tradable change.
ICA is where the differences emerge. Minimum risk improves by +0.066 (p=0.216), giving EVaR 0.616 versus 0.550 for the matched CVaR twin. With NTS components, symmetric Rachev rises by +0.103 (p=0.314). Under CTS components, the corresponding four contrasts are +0.012, -0.001, +0.024 and +0.003.
The variation therefore comes from the multivariate representation and marginal family. Choi describes what the comparison combines: ICA plus NTS against MNTS "includes both the dependence structure and the restrictions on the marginal parameters". The shared subordinator makes MNTS impose a common alpha and theta, whereas ICA permits component-specific alpha_i and theta_i. Choi also finds the reverse under CVaR, where symmetric Rachev moves the other way at -0.069. Both knobs changed together, leaving the risk measure unidentified as the source.
The conclusion acknowledges the same pattern. Choi writes that EVaR-CVaR differences are generally small under MNTS and become more pronounced in the ICA specifications. His diagnosis matches the central result here and makes the paper's framing around the risk measure harder to justify.
None of the comparisons reaches statistical significance. The smallest two-sided p-value in the headline tables is 0.113, attached to a CTS-versus-NTS minimum-CVaR contrast, and no difference receives a star. Against standard benchmarks, the leading cell exceeds equal weight by +0.079 (p=0.438), minimum variance by +0.058 (p=0.320), and maximum Sharpe by +0.141 (p=0.268).
The conclusion puts the limitation plainly: the findings "are specific to the asset universe, sample period, and backtest design and do not imply that EVaR always outperforms CVaR". Choi also says they do not establish that one tempered stable model is always superior to another. Taken at face value, the empirical section offers point estimates from nine to eleven correlated sector bets.
Remark 1 contains a negative result that deserves more notice. Under the Normal assumption, E-STAR and both E-Rachev objectives give exactly the maximum-Sharpe weights. Each records an identical 604.55% cumulative return, 0.476 Sharpe and 51.74% drawdown. Normal minimum EVaR and minimum variance are twins as well, at 0.558 versus 0.559, with drawdowns of 37.27% and 37.36%. The entropic machinery earns its place only through non-Gaussian tail parameters.
A pure risk minimizer adds 271 points of return
In the ICA plus NTS minimum-risk cell, moving from CVaR to EVaR lifts cumulative return from 495.43% to 766.96%. Max drawdown edges from 37.79% to 38.30%. Expected return never enters the objective, and the feasible set has no target-return constraint. The extra return therefore arrives through sector allocation even though the optimizer seeks only to minimize risk.
Volatility increases from 14.22% to 15.28%, while excess kurtosis rises from 14.636 to 18.435. A 271-percentage-point return difference from a risk minimizer, accompanied by a one-point vol increase, reflects a sector tilt that paid over twenty-six years.
The crisis results are the paper's strongest evidence. During the GFC window, the two ICA minimum-EVaR portfolios lose 12.18% and 11.70%. Matched NTS minimum CVaR loses 17.61%, minimum variance loses 18.60%, and equal weight loses 33.77%. The GFC ordering holds even though the full-sample differences remain insignificant.
It is one crisis window among the eight regimes defined in the paper.
The cost of the allocation
ICA plus NTS minimum EVaR averages 21.67% one-way turnover each month. Its matched CVaR twin trades 9.76%, versus 6.54% for minimum variance and 1.35% for equal weight. At 25 basis points, the EVaR premium shrinks from +0.066 to +0.044. The ICA plus CTS version becomes negative at -0.018.
The ICA plus NTS cell still records the highest net Sharpe at every positive cost rate tested: 0.608, 0.599 and 0.573 at 5, 10 and 25 basis points. The paper's cost ladder ends at 25bp. Rachev portfolios are much more active, with turnover ranging from 59.18% to 64.74% per rebalance. The ICA plus NTS symmetric version falls from 0.506 gross to 0.417 net at 25bp. Choi notes that the funds are liquid and directly investable, making turnover and transaction-cost comparisons relevant for this universe. Twelve extra percentage points of monthly trading finance the whole measured edge.
Choi is appropriately careful: "Turnover alone does not show that the optimizer is unstable. It does show that the ICA and MNTS portfolios can require substantially different amounts of rebalancing." The first sentence is fair. FastICA identifies components only up to sign and permutation, while the mixing matrix is estimated again at every monthly rebalance. I did not find any sign or permutation stabilization described for those refits. My own conjecture, rather than the paper's, is that relabelling across consecutive months contributes to the turnover pattern when roughly 250 daily observations support up to eleven components.
The MNTS Rachev cells suggest similar fragility from another direction. Mean turnover is 33.01% and 32.05%, although the medians are 13.90% and 17.97%. Classical symmetric MNTS Rachev has a 0.00% median, which looks more like a corner-solution optimizer than a smooth one.
One design decision helps the paper. Choi writes: "We use the sector ETFs because the universe does not contain a low-volatility asset class in which a downside-risk minimizer can concentrate." That removes the easiest route for making a tail minimizer appear successful.
Our pass produced 0.235
We ran one automated pass based on the paper's description, trading from 2020-01-01 to 2024-07-01. This window differs sharply from Choi's. Our figures are Sharpe 0.235, Sortino 0.29, CAGR 4.24%, annualized volatility 19.42%, and max drawdown 37.6%. The paper's strongest cell, ICA plus NTS minimum EVaR, reports gross Sharpe 0.616 and Sortino 0.880 from January 2000 to March 2026. Our Sharpe is lower by more than half.
The results measure different setups. Our period differs, and our returns include commissions of four tenths of a cent a share with a one dollar order minimum. The reported 0.616 is gross.
Several documented choices explain part of the gap. Our 4.5 years include the COVID crash and the 2022 rate shock, while excluding the 2009-2019 recovery block. In the paper's regime table, that block gives the cell a 1.322 Sharpe. We also applied a 20% one-way monthly turnover cap. Choi's feasible set has only the long-only and budget constraints, with no such cap. Because his cell averages 21.67% turnover, ours binds in most months and leaves drifted weights in place of newly optimized positions.
The exact model, objective and confidence cell underlying our metrics was not recorded. It could be a 99% variant, or one of the E-Rachev cells that the paper reports between 0.403 and 0.510 gross.
Those setup differences leave some of the shortfall unexplained. Over the 2020-2023 window, the paper's ICA plus NTS minimum EVaR earns a 0.629 Sharpe with 22.35% vol, followed by 0.386 in the 2024 stub. Our 19.42% volatility is four points above the paper's full-sample 15.28% for the same objective. For a minimum-risk objective, that higher vol is the decisive figure. The turnover cap and stale weights offer the most plausible path to it.
Our portfolio also carried a 0.80 beta to SPY. A long-only, fully invested eleven-ETF portfolio with that beta behaves essentially as a market book over 2020-2024, allowing the 4.5-year window to dominate the outcome. One possible explanation remains uncheckable. We solve E-STAR through direct ratio optimization, whereas the paper uses Dinkelbach iterations. Its global convergence argument therefore does not cover the portfolio we traded, and worse local solutions could create this pattern. We could not reproduce the effect on this window in one pass. This result is evidence about our implementation before it becomes evidence about the paper.
The projection algebra and weight-dependent MGF domains are worth keeping. They make parametric EVaR inexpensive for any tempered stable book, while Remark 1 identifies exactly when the extra machinery has no use. I would treat 0.616 as one cell among thirty-one, drawn from eleven correlated names, insignificant at every tested level and funded by twelve points of additional monthly trading. It also depends on a monthly ICA refit, for which I did not find any sign- or permutation-locking described.
One result would change my view: the same +0.066 gap on a wider universe, with components locked across refits and a bootstrap p under 0.10.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.