AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

NEPSE Persistence Halves After July 2022 as GARCH and EGARCH Tie

The full-sample 18-day half-life fits neither regime; forecasts explain 5.8% of an absolute-return proxy

2026-08-14 · 7 min read · US equity ETFs

Reviewing: Volatility Dynamics and Forecasting of the Nepal Stock Exchange Index: A Comparative Analysis of GARCH and EGARCH Models · Nischal Shrestha and Jeevan Pokhrel · Read it on openalex

Our backtest of this idea

Our automated quick test, not the paper's

Regime-Adaptive Student-t GARCH Volatility Targeting for the NEPSE Index

Backtest period 2022-01-01 to 2026-08-12 · hypothetical, net of modelled costs

Why these figures are not the paper's (2)

Run on a different market than the paper

The paper models a Nepalese equity index, which is not available in the platform. The implementable version would forecast daily volatility for liquid US equity ETFs such as SPY, QQQ, and IWM; conditional-volatility clustering, asymmetric shock responses, and regime changes are price-return mechanisms that can be tested on these ETF return series, but the paper's Nepal-specific estimates and forecast results do not transfer.

The paper's own figures describe its universe and do not carry over to ours.

Our own audit found this run does not follow the paper faithfully (9)

  • deviation left undescribed by the audit (invalidates: The paper's 230-forecast rolling experiment; the reported GARCH and EGARCH RMSE, MAE, and Mincer–Zarnowitz R²; the paper's full-sample and subperiod persistence and half-life comparisons)
  • deviation left undescribed by the audit (invalidates: The paper's reported GARCH and EGARCH Mincer–Zarnowitz R² values and any direct ranking based on those values)
  • Regime-adaptive rolling estimation and exposure changes: The strategy can replace the paper's fixed 919-return estimation window with a 504-return window and halve exposure when operational instability diagnostics trigger. (invalidates: The paper's 230-forecast GARCH and EGARCH RMSE and MAE values, its Mincer-Zarnowitz R² values, and its reported full-sample and pre-/post-break persistence and half-life comparisons whenever the adapted regime action is active.)
  • Operational realized-volatility proxy: Operational model selection compares conditional-variance forecasts with squared log returns, while the paper's Mincer-Zarnowitz regression uses absolute returns. (invalidates: The paper's reported GARCH and EGARCH Mincer-Zarnowitz R² values and its ranking based on those R² values for the operational squared-return evaluation.)

5 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2022Total 37.3%Aug 2026
Sharpe
0.83
Total Return
37.3%
Max Drawdown
-11.4%
CAGR
7.2%
Volatility
9.8%
Beta vs SPY
0.45
Trades
1,150

What the paper reports for its own strategy

  • Out-of-sample RMSE 0.009552 (GARCH(1,1)) vs 0.009670 (EGARCH(1,1)), 230 one-step-ahead rolling forecasts, 30 May 2021-15 May 2026 sample
  • Out-of-sample MAE 0.007521 (GARCH) vs 0.007406 (EGARCH)
  • Mincer-Zarnowitz R² 0.058319 (GARCH) vs 0.045760 (EGARCH) against absolute daily return proxy
  • Volatility half-life 18.0 trading days (GARCH, persistence 0.9622) and 12.0 trading days (EGARCH, persistence 0.9438), full sample

A trader should care about the gap between 0.9622 and 0.9211. The first is full-sample GARCH persistence for the Nepal Stock Exchange (NEPSE) index; the second comes from the 880 observations after 13 July 2022. Persistence before the break is 0.9396, estimated from 269 observations. Shrestha and Pokhrel report half-lives of 18.00 trading days for the full sample, 11.12 days before the break and 8.43 days afterward. A book sized on an 18-day decay from the full-sample GARCH therefore uses a parameter suited to neither part of the sample.

The study covers a market that rarely receives this kind of volatility analysis. Its data are daily NEPSE index closes from nepsealpha.com, spanning 30 May 2021 to 15 May 2026: 1,150 prices and 1,149 log returns. ADF rejects a unit root in returns (-9.2908, p<0.01), while the index level is not rejected (-2.3801, p=0.4174). Returns have skewness of 0.4917, excess kurtosis of 2.0621 and a Jarque-Bera statistic of 249.8829. Daily standard deviation is 0.013882, around 22% annualized.

Auto-ARIMA selects an AR(3) mean equation. The ARCH-LM statistic on its residuals is 42.2621 at lag 5 (p<0.001), giving the authors grounds to model conditional variance. They estimate GARCH(1,1) and EGARCH(1,1) by maximum likelihood with standardized Student-t innovations, whose shape is around 5.05. A GJR-GARCH(1,1) provides an additional check on asymmetry, and inference relies on White sandwich standard errors.

Forecasting uses a fixed-length rolling window. The initial estimation window contains the first 919 returns. With each new observation, the parameters are estimated again and the oldest return leaves the window. This yields 230 one-step-ahead forecasts, evaluated against absolute daily return as the realized-volatility proxy. GARCH records RMSE of 0.009552 and Mincer-Zarnowitz R² of 0.058319. EGARCH records RMSE of 0.009670 and R² of 0.045760, although its MAE is lower, at 0.007406 against 0.007521. Within the sample, GJR has the highest log-likelihood (3390.569) and the lowest AIC (-5.8861). GARCH retains the lowest BIC (-5.8475). The BIC-selected Bai-Perron procedure locates a single break on 13 July 2022.

The evidence on asymmetry remains unsettled, as the authors acknowledge. EGARCH estimates a signed-shock coefficient of -0.0379 with p=0.146 and a magnitude coefficient of 0.2611 (p<0.001). The GJR asymmetry coefficient is 0.0986 (p=0.026). Their summary reads: "the leverage effect remained model-dependent rather than conclusive." Elsewhere they report skewness of +0.4917 without tying it back to this result. The heavier tail is on the right, whereas an equity index governed by the leverage story should skew in the opposite direction.

Forecasting variance for one day

The out-of-sample target is one-day-ahead conditional variance. The AR(3) mean is estimated, and within the GARCH mean equation AR(1) (0.0985) and AR(3) (0.0726) are significant. μ (p=0.079) and AR(2) (p=0.076) are insignificant. Out-of-sample return forecasts are never scored.

The deliverable can therefore support exposure scaling toward a volatility target or a VaR budget. The analysis ends with statistical loss functions. We found no position rule, transaction costs, VaR backtest or utility calculation anywhere in the paper. That scope is legitimate for a frontier-market econometrics study. A practitioner still has to provide the whole return-generating layer, leaving any claim that the GARCH-EGARCH distinction pays beyond the paper's evidence.

Conflicting loss functions

The RMSE difference is 0.000118, about 1.2% of the level, and favors GARCH. MAE favors EGARCH.

The rankings conflict.

The authors are direct about what their metrics can establish. The paper says "the reported R² values alone do not establish forecast unbiasedness or statistical significance." It follows with: "the forecasting conclusions should be based primarily on the relative RMSE and MAE values and should not be interpreted as evidence of strong absolute predictive power." Their claim is already narrow. The remaining problem is practical: the two loss measures order the models differently on this sample, and we found no Diebold-Mariano test or other test of the loss differential. Someone selecting a production model receives little direction from the comparison. That comes close to the paper's own conclusion that the models are broadly comparable.

Another result deserves attention because it runs against intuition. The symmetric model fails the sign-bias diagnostic (2.990, p=0.0029; joint 9.319, p=0.0253). The authors answer by fitting EGARCH and then GJR. Even so, the symmetric model wins two of the three out-of-sample measures.

Only 5.8% of a noisy proxy

The higher Mincer-Zarnowitz R² is 0.058319. Absolute daily return is a noisy measure of latent volatility, making a low R² partly mechanical, as the paper notes. The reported table omits the quantities that matter for sizing: the estimated intercept, the slope and a joint test of a=0 and b=1. An R² by itself cannot distinguish a forecast with a stable positive slope from one with no slope. Those cases require very different treatment inside a volatility target. We raised the same issue in another setting, where a reported accuracy measure failed to explain the model's equity curve at all (/articles/where-finsmart-s-returns-come-from).

Where the recipe ends

The estimation procedure is specified unusually well for reproduction. It uses R 4.6.0, rugarch 1.5-5, a hybrid solver, Student-t innovations, White sandwich covariance and the exact 919/230 fixed-window protocol. Three details still required judgment.

Can a fixed 919-day window absorb the break?

The break analysis is the paper's strongest contribution and its most obvious unfinished piece. Persistence declines after the break in both models, and both half-lives contract. GARCH moves from 11.12 days before the break to 8.43 afterward; EGARCH moves from 13.52 to 7.00. Full-sample GARCH has an 18.00-day half-life, longer than either subperiod. This is precisely the Lamoureux-Lastrapes effect cited by the paper. EGARCH's full-sample half-life of 12.00 days falls between the regimes and conceals the post-break decline.

The authors state the problem plainly: "the full-sample GARCH persistence and half-life exceed both subperiod estimates, indicating that ignoring the break overstates persistence." The abstract likewise says "a structural break was identified on 13 July 2022, after which both models exhibited lower persistence and shorter volatility half-lives." Yet the forecasting study uses a fixed 919-observation window spanning the break. It advances 230 times, once for each out-of-sample forecast, while the pre-break subsample contains 269 observations. Pre-break returns remain in the estimation window even for the final forecast.

The imbalance is large: 269 observations before the break and 880 after it. The authors themselves warn that the pre-break estimate is fragile. They make no causal attribution for the break and identify regime-switching specifications as a next step. The next paper needs a break-aware or two-state benchmark, scored on realized drawdown and turnover.

Our adaptation cannot settle the paper's claim

We could not trade the paper's asset because NEPSE is unavailable on our platform. Our implementable adaptation applies the same conditional-volatility sizing idea to SPY, QQQ and IWM. The Nepal-specific parameter estimates do not transfer.

Our automated pass created a long-only overlay that scales exposure toward a 12% annualized volatility target using a one-day Student-t GARCH or EGARCH forecast. Selection between the models uses trailing 60-forecast RMSE. A failed stability check halves exposure and reduces the estimation window to 504 returns. From 2022-01-03 to 2026-08-12, the run returned 37.30% in total, with a Sharpe of 0.83, a maximum drawdown of 11.44% and 1,150 trades. Realized volatility was 9.79% against the 12% target. The overlay carried a 0.45 beta to SPY, which explains why the loss ended at 11.44%. The paper claims only "practical implications for risk assessment," while warning against treating volatility parameters as constant through time.

Those P&L figures are provisional. The run recorded 1,150 trades without attached symbols and 1,156 daily rows across a window containing 1,162 trading days. We therefore cannot verify that the exposure reported was the exposure held. The p-value thresholds governing which model may trade were also never fixed. As a result, the traded model path is not pinned down and probably changes across runs.

The paper offers no P&L for comparison. Its own performance figures concern forecast loss: RMSE 0.009552 versus 0.009670, MAE 0.007521 versus 0.007406 and R² 0.058319 versus 0.045760, all measured over 230 forecasts. A Sharpe of 0.83 and an RMSE of 0.009552 measure different things; neither verifies nor disputes the other. Every return-producing element in our run is ours: the 12% target, the position cap, the 2bp one-way turnover charge, dynamic switching and regime de-risking.

Our protocol also departs from theirs. Model selection used squared returns in our run, while the paper uses absolute returns. We omitted the ARCH(1) through ARCH(3) candidates listed in their Table 4. Evaluation covered the full 2022-2026 period instead of their 230-day holdout. This was one automated pass, with unresolved implementation gaps, rather than a verdict on the authors' work.

Two results would change my view of the paper itself. A Diebold-Mariano test could establish whether the 0.000118 RMSE gap supports a ranking. A sizing comparison could show whether break-aware persistence beats the fixed 919-day window on realized drawdown. The former would resolve the near-tie; the latter would turn 13 July 2022 into a trading rule instead of a caveat.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each NEPSE trading-day close t:
  1. Compute log returns r_t = log(close_t / close_{t-1}).
  2. Using observations dated no later than t, update the structural-break,
     rolling forecast-error, and parameter-shift diagnostics.
  3. If any diagnostic is triggered, maintain the unstable flag for at least
     20 trading days; use 504 returns and multiplier 0.5. Otherwise use 919
     returns and multiplier 1.0.
  4. Fit AR(3)-Student-t GARCH(1,1) and EGARCH(1,1) models on the active window.
     Run residual and sign-bias diagnostics. Fit GJR-GARCH(1,1) only as an
     asymmetry robustness check.
  5. Generate each valid candidate's one-step-ahead conditional-variance
     forecast h_(t+1).
  6. Select GARCH or EGARCH by the lowest trailing 60-forecast RMSE against
     squared log returns. Break an exact RMSE tie by trailing MAE, then GARCH.
     Before 20 evaluated forecasts, use cumulative RMSE; with no evaluated
     forecast, default to a valid GARCH fit.
  7. Compute raw_weight = 0.12 / (sqrt(252) * sqrt(h_(t+1))).
  8. Set target_weight = raw_weight * regime_multiplier, constrained to a
     long-only 100% position cap and the 4.0 gross-leverage platform limit.
  9. Submit the target as a market-on-close rebalance at close t for the
     close-t to close-(t+1) return interval.
 10. Skip the signal or trade when required history, a valid fit, or the
     execution close is unavailable; do not forward-fill execution data.