AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

The t-copula earns 0.06 AIC points

A careful bivariate volatility fit on resource and energy indices still awaits a forecasting test.

2026-09-08 · 7 min read · US sector and natural-resource ETFs

Reviewing: Jointly Modeling Dynamic Dependence and Volatility in Natural Resource and Energy Indices Using a Multivariate t-Copula eGARCH Model · Jenny K. Chen, Najmeh Nakhaei Rad and S. Yaser Samadi · Read it on openalex

Our backtest of this idea

Our automated quick test, not the paper's

Liquid ETF Pairwise Copula-eGARCH Tail-Risk Allocation

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (2)

Run on a different market than the paper

The paper analyzes resource and oil-and-gas indices that are not identified as directly holdable instruments in the available universe. Implement the same dependence-aware volatility and portfolio-risk mechanism using liquid US ETFs representing those exposures, such as XLE for energy and a broad natural-resources/materials ETF where coverage permits; asymmetric volatility, heavy-tailed marginal modeling, and cross-asset tail dependence are price-return mechanisms that survive this ETF substitution. The paper's reported estimation and simulation results apply only to its original index series, not to the ETF implementation.

The paper's own figures describe its universe and do not carry over to ours.

Our own audit found this run does not follow the paper faithfully (7)

  • Student-t eGARCH magnitude centering: The implementation replaces the paper's stated standard-normal center √(2/π) with the distribution-specific expectation for variance-standardized Student-t innovations. (invalidates: Exact reproduction of the paper's Table 8 AIC and BIC values, Table 9 Copula-eGARCH-MVT parameter estimates, and Table 11 Monte-Carlo results under the paper's printed eGARCH recursion.)
  • Inverse-forecast-volatility allocation: ETF weights are set in proportion to inverse forecast marginal volatility and then capped at 10%. (invalidates: The paper's Table 8 model ranking and Table 9 parameter estimates cannot be treated as evidence for this strategy's allocation weights, annualized return, Sharpe ratio, turnover, or maximum drawdown.)
  • Pair-sleeve construction and cross-pair tail aggregation: Each simulated pair uses inverse-forecast-volatility within-pair weights, and strategy-level VaR and Expected Shortfall are the maxima across five separately fitted pairs. (invalidates: The paper's Table 4 dNRI and dOGI 5% VaR and Expected Shortfall values and its NRI-versus-OGI downside-risk comparison do not apply to the strategy's weighted ETF pair returns or maximum-across-pairs risk forecasts.)
  • Rolling percentile exposure overlay: The strategy compares daily next-day maximum pair VaR and Expected Shortfall with trailing 80th percentiles and halves exposure when either trigger fires. (invalidates: The paper's empirical Copula-eGARCH-MVT fit results do not establish the overlay's VaR exception frequency, exposure path, drawdown reduction, or risk-adjusted trading performance.)

3 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 42.5%Jul 2024
Sharpe
0.47
Total Return
42.5%
Max Drawdown
-27.9%
CAGR
8.2%
Volatility
20.0%
Beta vs SPY
0.52
Trades
9,937

The Student-t copula in Chen, Nakhaei Rad and Samadi's model improves AIC by 0.06 points over a Gaussian copula fitted to the same eGARCH margins. In the paper's model-selection summary ranked by AIC and BIC, the scores are 48303.56 and 48303.62. That thin in-sample margin is the measured value of the joint tail dependence for which the framework is named.

The rest of the model carries the argument.

Start with the margins

The sample contains two daily series, a Natural Resource Index and an Oil and Gas Index, from 30 April 2014 to 27 June 2023. There are 2,306 observations, leaving 2,305 after differencing. Augmented Dickey-Fuller p-values on the levels are 0.1832 and 0.3684, which leads the authors to first-difference both series.

In index points, the differenced series have standard deviations of 30.12 and 124.56. Kurtosis is 8.32 and 7.30, while skewness is -0.72 and -0.47. The paper describes skewness of -0.72 and -0.47 as negligible, an awkward characterization beside a model selected partly for asymmetry. An ARCH LM test using 12 lags produces 639.82 and 319.13, both p < 2.2e-16. Fat tails and volatility clustering are plainly present.

Estimation proceeds in two stages. Each series receives an AR(1)-eGARCH(1,1) margin with Student-t innovations. The specification models log-variance directly and lets negative shocks affect next-day volatility differently from positive shocks of equal size. Each conditional CDF then transforms its fitted margin into uniforms. A Student-t copula joins those uniforms, with its correlation matrix evolving through a DCC(1,1) recursion. The copula's degrees of freedom determine how frequently the two markets reach their extremes together.

A desk could use the output to forecast tomorrow's conditional covariance, simulate from the joint law, calculate 5% VaR and Expected Shortfall for a book holding both exposures, and adjust size. The parameter estimates for the best-fitted model offer no return signal, a claim the paper does not make. The two means are 0.221 (t = 0.445) and -0.844 (t = -0.245), while the second series has an AR(1) coefficient of -0.041 (t = -1.266). Risk timing is the proposed edge.

Volatility persistence comes in at 0.987 and 0.991. Sign effects are negative, at -0.072 (p < 0.001) and -0.020 (p = 0.044), versus magnitude effects of 0.140 and 0.105. Leverage asymmetry is present, though ordinary shock size matters more. The DCC coefficients are a = 0.021 and b = 0.957. Their sum is 0.978, describing a correlation process that responds slowly and retains shocks for a long time.

Tail-risk users should look closely at the next estimates. The marginal Student-t shapes are 8.867 and 11.092, while the copula degrees of freedom are 12.546 (SE 3.148). The estimated joint tail is thinner than either marginal tail. Its finite value supports the authors' statement that co-exceedance is non-zero. Yet the direction weakens the strongest version of their case, in which the two exposures crash together more severely than their separate fat tails would imply.

What wins the eighteen-model contest?

The authors fit eighteen specifications over the full sample. These combine DCC(1,1) and DCC(2,1), three margin families under multivariate normal and multivariate t, and six copula variants. Best and worst AIC span 48303.56 to 48465.63.

The winner's components tell the useful story. With the copula family held fixed, replacing sGARCH margins with eGARCH margins improves AIC by 10.65 points, from 48314.21 to 48303.56. With the margins held fixed, replacing the Gaussian copula with the t-copula improves it by 0.06. Nearly all the gain comes from the asymmetric variance equation. At this sample size, the dependence change barely registers.

BIC reverses the ranking. Its minimum of 48400.35 goes to the simpler Copula-sGARCH-mvnorm, compared with 48401.19 for the selected model. The authors acknowledge this directly: "Given the small BIC differences, the selection of Copula-eGARCH-MVT is based on its lowest AIC, its ability to represent asymmetric volatility and tail dependence, and its superior simulation performance."

Credit is due for saying so. The reasons still need separating. The first relies on a criterion with a lighter parameter penalty, using one in-sample fit in which all eighteen models see the same 2,305 differenced observations. Tail dependence, the second reason, contributes 0.06 AIC points in the paper's own selection table. The third relies on a Monte Carlo where every model is estimated on data generated by its own process.

Section 3.3.1 contains the qualification. The abstract omits it and says the model "provided the best fit".

Graded on self-generated data

For each model, the Monte Carlo treats its empirical estimates as truth. It simulates 1,000 replications of length T = 2,305 from that model's data-generating process, then re-estimates the same specification. eGARCH recovers itself closely, with biases no larger than 0.0005, MSE around 1e-4, and 95% Wald coverage from 0.946 to 0.962 across all nineteen parameters.

This establishes that a correctly specified estimator can recover its own process. It cannot show which model survives when the true process is unknown, because the exercise never fits one model to data simulated by another.

The reported comparisons make the distinction harder to ignore. Under its own DGP, sGARCH produces an unconditional variance bias of 4.89e19 when the true value is 918.12. Its omega bias is 21,171, and coverage for the joint copula shape is 0.141. The gjrGARCH column reports a variance bias of 1.60e19 and coverage of 0.997. Reading these as economic differences between models would be a mistake. The authors attribute the deviations to "finite-sample instability and numerical issues rather than differences in true parameter values".

Their explanation should govern how the table is read. It reports numerical stability for the log-variance parameterisation: coverage of 0.946 to 0.962 across all nineteen eGARCH parameters, versus 0.141 for the sGARCH joint copula shape. The table cannot support a forecasting horse race. One discrepancy remains unresolved: the simulation uses 9.6912 for the joint copula shape, whereas the fitted table gives 12.546.

We did not find a rolling out-of-sample forecast comparison, a VaR exceedance test, or any portfolio result in the paper. Its conclusion places evaluation under "alternative distributions, higher-frequency data, and risk measures such as Expected Shortfall" among future work, so the authors do not claim those tests were run.

The reported 5% VaR and ES estimates are in-sample. They are -1.983% and -3.091% for the resource series, and -4.151% and -6.087% for oil and gas. Those figures appear in percent, while the volatility tables use index points. Mean conditional volatility is 28.370 for the resource series and 118.556 for oil and gas. Levene's test on the differenced conditional volatilities gives F = 557.08. Consequently, part of the cross-series risk difference reflects the scale created by differencing index levels instead of logging returns.

An ETF adaptation

We applied the mechanism to instruments outside the paper because the two indices cannot be held directly among the ETFs available to us. A point-in-time annual screen selects the 100 US ETFs with the highest dollar volume. After removing duplicates whose trailing return correlation exceeds 0.95, it retains ten. Those ten form five bivariate sleeves, including energy and broad resource exposure. None of the paper's estimates describe these instruments.

Each pair receives AR(1)-eGARCH(1,1) Student-t margins and a DCC(1,1) Student-t copula. Estimation uses 756 days, with a minimum of 504, and the model is refitted every five trading days while states update daily. Base weights use inverse one-day forecast volatility. Each ETF is capped at 10%, with any clipped weight remaining in cash.

For every pair, we simulate 10,000 next-day joint returns and calculate 5% VaR and ES. The daily risk reading is the worst across the five sleeves. It is compared with its own trailing 252-day 80th percentile, using only prior forecasts. A breach of either threshold cuts gross exposure in half to 0.5. The implementation rebalances daily at the close, charges $0.004 a share, assumes zero modelled slippage, and runs from 2020-01-01 to 2024-07-01.

Across that period, the book returned 42.53% in total. Its Sharpe was 0.47, Sortino 0.65, Calmar 0.29, maximum drawdown -27.94%, and beta 0.52 against SPY. A 0.47 Sharpe paired with a 27.94% drawdown is weak for a book designed solely to reduce exposure before tail losses arrive. This was one automated pass constructed from the paper's description, rather than a verdict on the authors' work.

Two limitations deserve weight. A 10% cap across ten ETFs, with clipped weight held as cash, creates a cash allocation by construction. Comparing its returns with a fully invested resource sleeve would therefore mislead. Zero modelled slippage on daily rebalancing also overstates the top line.

The paper supplies no return-based performance measures of its own. It reports AIC/BIC and simulation coverage, so we make no performance comparison. Our adaptation also chooses sizing rules, refit timing and thresholds absent from the paper. Its parameter estimates and coverage findings do not transfer to this implementation.

The missing test

A rolling-window exercise would answer the central question. Fit the model, issue out-of-sample forecasts for one-day 5% VaR and ES, and test the exceedance rate for unconditional coverage as well as exceedance independence. The comparisons should be a t-DCC and a plain EWMA benchmark. Extra parameters would earn their place if the copula-eGARCH exceedance rate came closer to 5% with less clustering.

Until such evidence exists, the defensible claim remains narrow. Under its own DGP, the log-variance specification delivered 95% Wald coverage between 0.946 and 0.962 for all nineteen parameters. In a separate data-generating process, sGARCH coverage for the joint copula shape was 0.141, while its unconditional variance bias reached 4.89e19 against a true value of 918.12. Those separate processes support a numerical-stability comparison, rather than a like-for-like contest. On these two series, the tail-dependence layer buys 0.06 AIC points, which is nothing.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each allocation date t after the close:
  1. Select the applicable calendar year's 100 ETFs ranked by dollar volume.
  2. Using returns available through t, classify exposures over 252 days.
     Require at least 200 observations and remove global duplicates whose
     trailing-return correlation exceeds 0.95.
  3. Retain 10 eligible ETFs and partition them into five bivariate sleeves.
  4. For each pair, use 504-756 complete observations to fit:
       - AR(1)-eGARCH(1,1) margins with variance-standardized Student-t errors;
       - a DCC(1,1)-driven bivariate Student-t copula.
     Refit parameters every five trading days; update states daily.
  5. Forecast each ETF's t+1 conditional volatility. Set base ETF weights
     proportional to inverse forecast volatility, normalize, and clip each
     position at 10% without redistributing clipped weight; residual capital
     remains in cash.
  6. Within each fitted pair, form fixed inverse-forecast-volatility weights.
     Draw 10,000 next-day joint returns and calculate 5% VaR and Expected
     Shortfall for the pair sleeve.
  7. Define daily joint-tail summaries as the maximum VaR and maximum ES
     across the five pairs.
  8. Compare each current summary with its trailing 252-day 80th percentile,
     using strictly prior forecasts and requiring 126 prior observations.
     If either threshold is exceeded, multiply gross exposure by 0.5;
     otherwise use 1.0.
  9. Submit the resulting rebalance at close t for the subsequent
     close-to-close holding return. Skip invalid fits or unavailable prices;
     do not fabricate or forward-fill execution prices.