A volatility forecast that loses to the unconditional mean of its own target cannot size a position. Every model in this paper loses to the unconditional mean of its own target. The ranking table carries the R-squared column, and that is the number a risk-targeting desk has to deal with before anything else in the study matters.
Oyedele compares five machine-learning forecasters against two asymmetric GARCH specifications. The data is daily equity prices for six SADC markets (Southern African Development Community): Botswana, Mauritius, Namibia, South Africa, Zambia and Zimbabwe. The machine-learning side is Random Forest, LSTM, GRU, XGBoost, and a weighted-average Ensemble of those four. The econometric side is EGARCH and GJR-GARCH, both of which let negative shocks move conditional variance more than positive ones of the same size. Data comes from Yahoo Finance, 2 January 2015 to 8 May 2026. Prices are converted to log returns and diagnostics are run: Jarque-Bera, ADF, Ljung-Box, plus spectral peak power and dominant frequency. The sample is then split into training and test sets. Forecasts are scored on RMSE and MAE, with MAPE and R-squared entering a composite average rank.
The economic case behind this kind of work is volatility timing. If you can forecast next-period variance better than a trailing window can, you scale exposure inversely to that forecast. Realized risk stays closer to target and you shed less capital in stress. Better forecast, tighter risk control, and the improvement has to survive the turnover that rescaling generates. The paper claims implications for risk management and portfolio optimization on the strength of the error metrics alone. We found no return, Sharpe, drawdown, hit rate or t-statistic anywhere in the text.
Random Forest posts the lowest average MAE at 0.0121, against EGARCH at 0.0155 and GJR-GARCH at 0.0158. Table 2 reports EGARCH's per-country MAE as 0.0152 in each of five markets and 0.0172 in South Africa. Those average to 0.0155, ranked sixth of seven. Measured against those two averages, Random Forest's error is 23% below GJR-GARCH's 0.0158 and 22% below EGARCH's 0.0155. The table labels it "Best Overall".
The winning margin is one ten-thousandth
Random Forest's lead over the deep-learning models is 0.0001 of MAE (0.0121 against 0.0122 for LSTM, GRU and the Ensemble alike). On RMSE it loses to all three, 0.0166 against 0.0164. The abstract calls 0.0166 the lowest RMSE in the study. Table 3 ranks it fourth of seven. And the discussion says the opposite of the abstract: "While Random Forest dominated the overall ranking, LSTM and GRU achieved superior RMSE and R-squared outcomes." We did not find a Diebold-Mariano test or a Model Confidence Set anywhere in the methods. So a 0.0001 gap is being read as an ordinal ranking with nothing behind it.
The composite rank is also not monotone in the errors it aggregates. EGARCH beats GJR-GARCH on both reported error measures (RMSE 0.0205 against 0.0209, MAE 0.0155 against 0.0158). It also beats it on R-squared, -0.7288 against -0.7908. Table 2 and Table 3 both rank EGARCH sixth of seven. In the overall ranking table it lands seventh, below GJR. MAPE is the only column where GJR-GARCH wins: 1,223,823.2 against 1,274,145.4. That gap of about fifty thousand sits inside figures above a million. A MAPE above 100,000 implies denominators near zero, and the paper never says how the target was built, so the cause cannot be confirmed. The paper itself says Random Forest achieved the highest overall rank "due primarily to its superior MAE and MAPE outcomes". By the paper's own account, the broken column is the stated basis of first place.
To the paper's credit, the discussion says plainly that no single metric gives a complete assessment. It also argues that machine learning should complement rather than completely replace econometric models. The same passage defends the multi-criteria average by citing Poon and Granger, Akgun and Gulay, and Leushuis and Petkov on the value of several evaluation criteria. Multi-criteria averaging helps when each criterion is informative. A MAPE column running from 296,501.0 for Random Forest to 1,274,145.4 for EGARCH is not informative, and averaging it in contaminates the composite rather than diversifying it. The concession sits badly beside a table that assigns "Best Overall" on a composite whose loudest component is broken.
Negative explanatory power, undiscussed
The R-squared column: LSTM, GRU and Ensemble at -0.1033, Random Forest at -0.1258, XGBoost at -0.1669, EGARCH -0.7288, GJR-GARCH -0.7908.
Every one negative. A constant forecast equal to the sample mean of the target would have beaten all seven models on the paper's own reported metrics. On the split, the paper says only that "the sample was partitioned into training and testing datasets". It never states which partition the Table 4 errors are measured on, nor the split dates or proportions. The discussion refers to R-squared once, writing that "LSTM and GRU achieved superior RMSE and R-squared outcomes", without noting that the whole column is below zero.
Part of the difficulty is that the forecast target is never defined. We did not find any statement of how daily realized volatility was constructed: squared returns, absolute returns, a rolling standard deviation, something else. Without that, 0.0121 has no interpretable scale. The near-zero denominators implied by the MAPE figures cannot be diagnosed either.
The pre-estimation table has a related problem. Table 1 reports means of 2.8351 for five markets and 8.3502 for South Africa, kurtosis 1.2189 and 1.6608, and ADF p-values of 0.628 and 0.724. Those are price-level quantities, non-stationary at conventional levels. The text nonetheless reads "kurtosis values exceeding unity indicate departures from normality and the existence of fat-tailed distributions". Kurtosis of 1.22 is thin-tailed. The fat-tail argument does not follow from that number.
Five markets, one row of numbers
Botswana, Mauritius, Namibia, Zambia and Zimbabwe report identical figures on every statistic Table 1 reports, to every digit printed. Mean 2.8351, SD 0.1670, skewness 0.6157, kurtosis 1.2189. Jarque-Bera 356.8751 with p 0.000, ADF -1.303 with p 0.628, Ljung-Box 27085.1 with p 0.000, peak power 46.064, dominant frequency 0.0011. The duplication carries into the forecast tables, where Random Forest's MAE is 0.0120 in each of those five markets and 0.0129 in South Africa. Same for every other model.
The paper never states which index or ticker represents each market. Five separate national price series cannot produce identical statistics on every reported line, so the tables either use one series or repeat one set of results. Functionally the cross-section is one series plus South Africa, which removes the regional comparison the paper is built to deliver. Separately, the paper is dated "Received: 1/16/2026" and the sample runs to 8 May 2026, so part of the sample postdates submission.
Would this change an ETF position?
We cannot trade SADC equities. Any implementation here forecasts next-day variance for liquid US equities and broad and sector US equity ETFs. That is a different universe and not a replication. We would run the same lagged-return and lagged-variance inputs on US names, which need no futures curve and no local microstructure. None of the paper's reported error figures carry over, and nothing here should be read as a US result. We have made a version of the substitution point before, about a premium whose legs no traded instrument spans.
A rolling-window volatility estimate is the benchmark such an implementation would have to clear, and the paper offers no figure against one. Rescaling exposure on a forecast pays turnover. A forecast with R-squared of -0.1033 has no measured edge over a constant to pay it with. The paper's own numbers give no help in choosing a model either. The Ensemble of four algorithms reproduces LSTM on all four metrics in the ranking table (MAE 0.0122, RMSE 0.0164, MAPE 351,472.2, R-squared -0.1033). There is no diversification across the forecasters to collect.
Four things would move me: a stated definition of the volatility target, per-country series that actually differ, a Diebold-Mariano test on the 0.0001 gap, and at least one model with R-squared above zero.