One coefficient carries the trading case: 0.110 on a lagged headline residual, from an equation explaining 3.7% of the variation in absolute EU Allowance returns.
Most of the paper is taxonomy.
Mallik, Megaritis and Rodgers argue that web datasets form a third information category alongside public and private information. They call it open information, meaning widely available content that is not a designated form of public information. The financial mechanism is familiar: news arrival drives volatility. Their input is Mitchell and Mulherin's transformation of headline counts. If headline flow from the open web contains information beyond regulated disclosure, a carbon volatility book should get a better one-day-ahead forecast of |r| than persistence provides. That would be tradable.
The construction is unusually clean for a study combining text and prices. There are no keyword lists and no sentiment model. The authors take the GDELT Events Database, retain headlines with the ENV secondary role code, then divide the daily count by CAMEO actor role. Headlines coded BUS, MNC or GOV are classified as public information, P. All other environmental headlines become open information, O. They regress log(O) on log(P) without an intercept. The coefficient is 1.08 (SE 0.002), and its residual is named log(Ou), the portion of open information unexplained by public information. EUA spot end-of-day closes come from the International Carbon Action Partnership, sampled daily from 01-04-2013 to 31-12-2023. The regression contains 2,803 observations and the VAR contains 2,753, after 106 missing values were filled through a randomised bootstrap.
The results begin with the proposed categories. O and P are correlated, with Pearson 0.68 in levels and 0.76 in logs, which settles the paper's first proposition. The contemporaneous correlation between the residual log(Ou) and |r| is 0.040.
A VAR(1) on (log(Ou), |r|) puts the lagged residual into the |r| equation at 0.110 (SE 0.052, p<0.05). Its own lag is 0.18 (SE 0.091, p<0.10), and adjusted R-squared is 0.037. When the lagged residual is added to the mean equation of an AR(1)-GARCH(1,1), both AIC and BIC improve across 253 of 253 expanding in-sample windows. The same holds for fGARCH. In one-step-ahead leave-one-out forecasts, Mean Error shifts from -0.05112 to -0.04284, RMSE from 1.94575 to 1.94543, and MAE from 1.52465 to 1.52311.
A daily bias gain of 0.0083
Mean Error is a bias measure. RMSE falls from 1.94575 without the regressor to 1.94543 with it, a difference of 0.00032, or 0.01%. MAE changes by 0.1%. The Mean Error improvement amounts to 0.0083 return points a day. A one-day forecast with 0.0083 less bias and a 0.00032 RMSE improvement will not change position sizing.
The paper gives no trading strategy, no Sharpe, no hit rate and no transaction cost assumption. Statistical fit therefore carries the abstract's promise of material significance for pricing.
A footnote says the Diebold-Mariano tests were statistically insignificant in most samples at 10%. The authors then compare p-values under two opposing one-sided alternatives and report that this comparison favoured the specification containing the extra term. Such a comparison is no test at all; the reported test was already insignificant. They also say the forecast figures come from one sample. Changing the sample alters the values, while leaving the order of percentage improvement similar. Accept that claim as written. The ~10% reduction in bias and the ~0.01% reduction in RMSE are both stable. Economic interest attaches to only one of them.
The paper itself labels the AIC/BIC exercise as in-sample, leaving the 253/253 count silent on forecasting performance.
We have covered this pattern before: an EGARCH asymmetry term wins on information criteria without producing anything worth trading (our earlier note).
Where the effect disappears
The authors disclose four failures in footnotes and Table 9:
- using signed returns instead of |r| in the VAR made the log(Ou) coefficient statistically insignificant;
- putting the term in the variance equation of the GARCH-X rather than the mean equation also made it insignificant;
- iGARCH improved 0 of 253 windows on AIC and 0 of 253 on BIC;
- EWMA improved 248 of 253 on AIC and 0 of 253 on BIC.
All four failures appear in the paper. The authors write that model selection matters and that an integrated-volatility specification may override the influence. Across the four GARCH variants in Table 9, the X term improves both criteria in 253 of 253 windows for GARCH(1,1) and for fGARCH. The surviving setup is narrow: the regressor is lagged, placed in the mean equation and tested against absolute returns.
What remains inside the residual?
log(Ou) is produced by a univariate log-log regression with no intercept and a reported adjusted R-squared of 0.992. Two series sharing a scale can give a no-intercept log-log fit of 0.992. The residual is small, and the VAR finds persistence in it. Its own lag is 0.34 (SE 0.018, p<0.01), with adjusted R-squared 0.121. Persistence in a 2013-2023 GDELT count series admits two readings. It may reflect day-to-day news surprise. It may equally reflect decade-long drift in GDELT coverage. The reported tests do not separate those explanations.
GOV creates another problem. The paper's code table labels GOV as "Public or Open". Because the sub-categories are difficult to distinguish, the authors assign it to P. They then argue that "Such a misclassification weakens OI", making a significant weak-form result evidence for a stronger true result. The paper asserts that inference without demonstrating it.
The concession and its escape hatch
The abstract says the authors "demonstrate their existence and justify material significance for pricing". Later, the discussion says the materiality outcomes "have been mixed" and describes the finding as "a significant yet feeble effect". Each sentence refers to the same coefficient.
The dispute concerns how the authors treat its small size. They suggest that during extreme events such as a bank run, the impact "could be materially larger", on the possibility that correlations rise under stress. For a reader considering carbon vol, this sentence carries the argument. Yet the sample comprises 2,753 daily observations from 01-04-2013 to 31-12-2023, and no tail or subsample test is reported.
Their other defence is that the tests use one type of web dataset, leaving open the possibility that another design would produce different outcomes. Evidence confined to one dataset, one asset and two of the four GARCH variants tried gives no reason to expect more from a better dataset.
We could not run this ourselves. Our price data includes no EUA spot or carbon futures series, and we lack GDELT Events history carrying CAMEO role codes. Our news coverage begins around 2020, too late to reconstruct a 2013-2023 ENV headline count.
The evidence that would change our position is the same residual, constructed the same way, producing a loss reduction that survives Diebold-Mariano on a second asset or a second web dataset. Until then, the paper offers a definition with a coefficient attached, and by the authors' own arithmetic that coefficient is a tenth of volatility persistence.