AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

ICA carries the EVaR premium in Choi's sector backtest

Matched MNTS portfolios differ by 0.003 Sharpe while ICA adds 22% monthly turnover

2026-09-08 · 9 min read · Portfolio optimization / tail-risk-managed sector allocation · US sector ETFs

Reviewing: Entropic Value-at-Risk portfolio optimization for tempered stable Lévy processes · Jaehyung Choi · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Rolling MNTS and ICA+NTS Entropic ETF Allocation with 95% and 99% EVaR

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (1)

Our own audit found this run does not follow the paper faithfully (9)

  • Eq. (57) feasible set (invalidates: The paper's reported portfolio weights, returns, drawdowns, turnover, Sharpe ratios, regime results, and transaction-cost results no longer apply.)
  • Normal E-STAR equivalence diagnostic (invalidates: The paper's Normal E-STAR and maximum-Sharpe equivalence cannot be checked in this specification.)
  • Exact eleven State Street sector ETF universe (invalidates: All paper-supplied performance targets; all paper turnover comparisons; all paper regime comparisons; all matched EVaR-versus-CVaR Sharpe differences and bootstrap p-values.)
  • Requested 99% EVaR robustness test (invalidates: The paper's 95% numerical performance targets do not apply to the 99% runs.)

5 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 20.5%Jul 2024
Sharpe
0.24
Total Return
20.5%
Max Drawdown
-37.6%
CAGR
4.2%
Volatility
19.4%
Beta vs SPY
0.80
Trades
244

What the paper reports for its own strategy

  • ICA+NTS minimum-EVaR (EVaR_95): gross annualized Sharpe 0.616, CAGR 8.60%, annualized vol 15.28%, cumulative return 766.96%, max drawdown 38.30%, Calmar 0.224, Sortino 0.880, historical VaR95 1.38%, CVaR95 2.21%, skew 0.069, excess kurtosis 18.435 (Jan 2000 - Mar 2026, gross of costs)
  • ICA+NTS minimum-EVaR net Sharpe: 0.608 at 5bp, 0.599 at 10bp, 0.573 at 25bp; net cumulative return 630.95% at 25bp
  • ICA+CTS minimum-EVaR: gross Sharpe 0.572, CAGR 7.97%, vol 15.53%, cumulative return 645.76%, max drawdown 39.21%, Calmar 0.203; net Sharpe 0.518 at 25bp
  • MNTS minimum-EVaR: gross Sharpe 0.555, CAGR 7.08%, vol 14.13%, cumulative return 499.57%, max drawdown 37.77%, Calmar 0.187; net Sharpe 0.534 at 25bp
  • Normal minimum-EVaR: gross Sharpe 0.558, CAGR 7.12%, cumulative return 506.49%, max drawdown 37.27%; net Sharpe 0.543 at 25bp
  • MNTS E-STAR_95: gross Sharpe 0.515, CAGR 8.48%, cumulative return 743.08%, max drawdown 49.04%

The paper's 0.616 Sharpe comes from the ICA fit, not from swapping CVaR for EVaR. Within its multivariate normal tempered stable model, that substitution barely changes the portfolio: gross Sharpe is 0.555 against 0.552 over twenty-six years, with a bootstrap p-value of 0.868.

Choi presents three contributions: the cumulant derivation, the portfolio constructions, and the out-of-sample comparison. The derivation is the lasting one. Its value does not depend on the empirical case for EVaR.

EVaR uses the Chernoff inequality to give a coherent upper bound on VaR and CVaR. It is the infimum over u > 0 of (1/u)ln(M_L(u)/eta), with M_L denoting the loss moment-generating function. Putting it into a portfolio optimizer creates two problems. The moment-generating function must remain finite, while both the fitted parameters of portfolio returns and the admissible range of u change with every candidate weight vector. A naive parametric method consequently refits a univariate distribution to the portfolio series at each trial weight.

Choi chooses multivariate constructions that produce the portfolio cumulant from asset-level or component-level fits. The fitted families are tempered stable Lévy processes, pure-jump processes with Lévy triplet (0, nu, mu), where exponential damping of the jump measure permits skewed heavy tails. The paper uses three: classical tempered stable (CTS), normal tempered stable (NTS) and multivariate NTS (MNTS).

Under MNTS, closure under linear projection yields the projected parameters directly. Alpha and theta remain unchanged, beta_bar = w'beta, gamma_bar = sqrt(w'Sigma w), and mu_bar = w'mu. The admissible upper endpoint becomes u_max(w) = (beta_bar + sqrt(beta_bar^2 + 2 theta gamma_bar^2))/gamma_bar^2. The ICA route instead applies FastICA, retains all N components, and fits each with a tempered stable law. It then obtains the portfolio cumulant by adding the component cumulants and an affine location term. For the inner one-dimensional infimum, the method truncates the admissible interval by epsilon, divides it into P subintervals, runs Brent's bounded minimization on each, and evaluates the finite endpoint separately.

The economic premise is straightforward. Sector returns have skew and heavy tails, so a measure that penalizes the full exponential moment could leave tail-heavy sectors sooner than variance or an average-beyond-a-quantile measure. Choi tests that premise through thirty-one portfolios formed from the eleven Select Sector SPDR ETFs. Nine have traded since December 1998. XLRE arrived in 2015 and XLC in 2018, expanding the investable universe from nine names to eleven.

Each fit uses Twelve months of daily returns, followed by a monthly rebalance. Portfolios are long-only and fully invested, with no target-return constraint. Prices are Yahoo adjusted prices, and the out-of-sample period runs from January 2000 to March 2026. Sharpe differences use the Ledoit-Wolf studentized circular block bootstrap with 4999 resamples and block length floor(T^(1/3)).

ICA plus NTS minimum EVaR supplies the best cell. It reports gross Sharpe 0.616, CAGR 8.60%, annualized vol 15.28%, cumulative return 766.96%, max drawdown 38.30%, and Calmar 0.224. Equal weight returns 0.537 with a 53.51% drawdown. Minimum variance reaches 0.559, while maximum Sharpe posts 0.476.

How much comes from EVaR?

Very little under MNTS. Holding the model constant, the EVaR-minus-CVaR Sharpe gaps are +0.003 for minimum risk, +0.001 for STAR, +0.002 for symmetric Rachev and +0.003 for the asymmetric version. Every p-value is at least 0.868. Once projection reduces the system to a single univariate NTS, replacing CVaR with its coherent upper bound produces no tradable change.

ICA is where the differences emerge. Minimum risk improves by +0.066 (p=0.216), giving EVaR 0.616 versus 0.550 for the matched CVaR twin. With NTS components, symmetric Rachev rises by +0.103 (p=0.314). Under CTS components, the corresponding four contrasts are +0.012, -0.001, +0.024 and +0.003.

The variation therefore comes from the multivariate representation and marginal family. Choi describes what the comparison combines: ICA plus NTS against MNTS "includes both the dependence structure and the restrictions on the marginal parameters". The shared subordinator makes MNTS impose a common alpha and theta, whereas ICA permits component-specific alpha_i and theta_i. Choi also finds the reverse under CVaR, where symmetric Rachev moves the other way at -0.069. Both knobs changed together, leaving the risk measure unidentified as the source.

The conclusion acknowledges the same pattern. Choi writes that EVaR-CVaR differences are generally small under MNTS and become more pronounced in the ICA specifications. His diagnosis matches the central result here and makes the paper's framing around the risk measure harder to justify.

None of the comparisons reaches statistical significance. The smallest two-sided p-value in the headline tables is 0.113, attached to a CTS-versus-NTS minimum-CVaR contrast, and no difference receives a star. Against standard benchmarks, the leading cell exceeds equal weight by +0.079 (p=0.438), minimum variance by +0.058 (p=0.320), and maximum Sharpe by +0.141 (p=0.268).

The conclusion puts the limitation plainly: the findings "are specific to the asset universe, sample period, and backtest design and do not imply that EVaR always outperforms CVaR". Choi also says they do not establish that one tempered stable model is always superior to another. Taken at face value, the empirical section offers point estimates from nine to eleven correlated sector bets.

Remark 1 contains a negative result that deserves more notice. Under the Normal assumption, E-STAR and both E-Rachev objectives give exactly the maximum-Sharpe weights. Each records an identical 604.55% cumulative return, 0.476 Sharpe and 51.74% drawdown. Normal minimum EVaR and minimum variance are twins as well, at 0.558 versus 0.559, with drawdowns of 37.27% and 37.36%. The entropic machinery earns its place only through non-Gaussian tail parameters.

A pure risk minimizer adds 271 points of return

In the ICA plus NTS minimum-risk cell, moving from CVaR to EVaR lifts cumulative return from 495.43% to 766.96%. Max drawdown edges from 37.79% to 38.30%. Expected return never enters the objective, and the feasible set has no target-return constraint. The extra return therefore arrives through sector allocation even though the optimizer seeks only to minimize risk.

Volatility increases from 14.22% to 15.28%, while excess kurtosis rises from 14.636 to 18.435. A 271-percentage-point return difference from a risk minimizer, accompanied by a one-point vol increase, reflects a sector tilt that paid over twenty-six years.

The crisis results are the paper's strongest evidence. During the GFC window, the two ICA minimum-EVaR portfolios lose 12.18% and 11.70%. Matched NTS minimum CVaR loses 17.61%, minimum variance loses 18.60%, and equal weight loses 33.77%. The GFC ordering holds even though the full-sample differences remain insignificant.

It is one crisis window among the eight regimes defined in the paper.

The cost of the allocation

ICA plus NTS minimum EVaR averages 21.67% one-way turnover each month. Its matched CVaR twin trades 9.76%, versus 6.54% for minimum variance and 1.35% for equal weight. At 25 basis points, the EVaR premium shrinks from +0.066 to +0.044. The ICA plus CTS version becomes negative at -0.018.

The ICA plus NTS cell still records the highest net Sharpe at every positive cost rate tested: 0.608, 0.599 and 0.573 at 5, 10 and 25 basis points. The paper's cost ladder ends at 25bp. Rachev portfolios are much more active, with turnover ranging from 59.18% to 64.74% per rebalance. The ICA plus NTS symmetric version falls from 0.506 gross to 0.417 net at 25bp. Choi notes that the funds are liquid and directly investable, making turnover and transaction-cost comparisons relevant for this universe. Twelve extra percentage points of monthly trading finance the whole measured edge.

Choi is appropriately careful: "Turnover alone does not show that the optimizer is unstable. It does show that the ICA and MNTS portfolios can require substantially different amounts of rebalancing." The first sentence is fair. FastICA identifies components only up to sign and permutation, while the mixing matrix is estimated again at every monthly rebalance. I did not find any sign or permutation stabilization described for those refits. My own conjecture, rather than the paper's, is that relabelling across consecutive months contributes to the turnover pattern when roughly 250 daily observations support up to eleven components.

The MNTS Rachev cells suggest similar fragility from another direction. Mean turnover is 33.01% and 32.05%, although the medians are 13.90% and 17.97%. Classical symmetric MNTS Rachev has a 0.00% median, which looks more like a corner-solution optimizer than a smooth one.

One design decision helps the paper. Choi writes: "We use the sector ETFs because the universe does not contain a low-volatility asset class in which a downside-risk minimizer can concentrate." That removes the easiest route for making a tail minimizer appear successful.

Our pass produced 0.235

We ran one automated pass based on the paper's description, trading from 2020-01-01 to 2024-07-01. This window differs sharply from Choi's. Our figures are Sharpe 0.235, Sortino 0.29, CAGR 4.24%, annualized volatility 19.42%, and max drawdown 37.6%. The paper's strongest cell, ICA plus NTS minimum EVaR, reports gross Sharpe 0.616 and Sortino 0.880 from January 2000 to March 2026. Our Sharpe is lower by more than half.

The results measure different setups. Our period differs, and our returns include commissions of four tenths of a cent a share with a one dollar order minimum. The reported 0.616 is gross.

Several documented choices explain part of the gap. Our 4.5 years include the COVID crash and the 2022 rate shock, while excluding the 2009-2019 recovery block. In the paper's regime table, that block gives the cell a 1.322 Sharpe. We also applied a 20% one-way monthly turnover cap. Choi's feasible set has only the long-only and budget constraints, with no such cap. Because his cell averages 21.67% turnover, ours binds in most months and leaves drifted weights in place of newly optimized positions.

The exact model, objective and confidence cell underlying our metrics was not recorded. It could be a 99% variant, or one of the E-Rachev cells that the paper reports between 0.403 and 0.510 gross.

Those setup differences leave some of the shortfall unexplained. Over the 2020-2023 window, the paper's ICA plus NTS minimum EVaR earns a 0.629 Sharpe with 22.35% vol, followed by 0.386 in the 2024 stub. Our 19.42% volatility is four points above the paper's full-sample 15.28% for the same objective. For a minimum-risk objective, that higher vol is the decisive figure. The turnover cap and stale weights offer the most plausible path to it.

Our portfolio also carried a 0.80 beta to SPY. A long-only, fully invested eleven-ETF portfolio with that beta behaves essentially as a market book over 2020-2024, allowing the 4.5-year window to dominate the outcome. One possible explanation remains uncheckable. We solve E-STAR through direct ratio optimization, whereas the paper uses Dinkelbach iterations. Its global convergence argument therefore does not cover the portfolio we traded, and worse local solutions could create this pattern. We could not reproduce the effect on this window in one pass. This result is evidence about our implementation before it becomes evidence about the paper.

The projection algebra and weight-dependent MGF domains are worth keeping. They make parametric EVaR inexpensive for any tempered stable book, while Remark 1 identifies exactly when the extra machinery has no use. I would treat 0.616 as one cell among thirty-one, drawn from eleven correlated names, insignificant at every tested level and funded by twelve points of additional monthly trading. It also depends on a monthly ICA refit, for which I did not find any sign- or permutation-locking described.

One result would change my view: the same +0.066 gap on a wider universe, with components locked across refits and a bootstrap p under 0.10.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

At each monthly rebalance close:
1. Form the eligible universe from the fixed eleven sector ETFs; require 12 months of return history.
2. Using adjusted daily log returns through the prior trading-day close, take the trailing 12-month estimation window.
3. Refit the selected representation:
   - MNTS: fit NTS marginals, impose common alpha/theta, and infer Brownian dependence.
   - ICA+NTS or ICA+CTS: rerun centered FastICA and fit every independent component.
4. For each candidate long-only, fully invested weight vector:
   a. Project its MNTS parameters or ICA component exposures.
   b. Recompute the admissible MGF domain.
   c. Evaluate EVaR by partitioning the valid u-domain into 20 intervals,
      applying bounded Brent minimization, and checking any finite endpoint.
   d. Reject nonfinite or unreliable candidates.
5. Optimize the active objective: minimum EVaR, expected return / EVaR,
   or symmetric/asymmetric E-Rachev. Apply the 20% one-way turnover limit.
6. Execute target weights at the current rebalance close using real prices only;
   skip trades lacking a valid execution price.
7. Hold for one month. Deduct the selected turnover charge and brokerage costs.

E-Rachev uses multiple feasible starts and retains the best local solution. No ex post winner is selected from the experiment grid.