A causal claim about KSE-100 volatility rests on a daily text sample averaging 18.256 tweets, with a minimum of one. The empirical pattern makes sense. The direction of causality remains an assumption.
The construction is the paper's most interesting feature. Muhammad, Bashir and Shah wrote a Python scraper around a two-layer keyword rule. A qualifying tweet must contain at least one market-anchored term, such as Pakistan Stock Exchange or PSX, and at least one expectation term, such as economy, policy, IMF or policy rate. This filter aims at market-wide expectations rather than firm-specific chatter. It comes closer to Baker and Wurgler's idea of sentiment than the event-window tweet samples common in this literature. Five years of scraping, from July 1 2020 to June 30 2025, produced 33,309 English-language tweets.
FinTwitBERT labels each tweet positive, negative or neutral according to its highest-probability class. The model is a BERT variant pretrained on financial tweets, so the labels come from a deep learning language model rather than a word list. The authors make this their main methodological claim. Earlier Pakistan studies used VADER, SentiWordNet, or Google Trends searches for the single word "Pakistani"; FinTwitBERT is presented as an advance over all three.
Daily labels feed three constructs. Intensity equals positive plus negative tweets divided by the total, making it one minus the neutral share. Disagreement multiplies the positive and negative shares and peaks at a 50-50 split. Attention is the raw daily tweet count. The authors place all three in the variance equation of a GARCH(1,1) for KSE-100 daily log returns. The mean equation is AR(1). EGARCH and TGARCH, the threshold specification also called GJR in the paper, handle asymmetry, while KIBOR and the PKR/USD rate enter as controls.
Table 3 contains the reported results. Intensity has a positive variance coefficient of 1.3556 by itself and 1.8176 in the full joint model with controls. Disagreement comes in at 1.7317 and 1.8323. Attention starts at -0.1599 and is insignificant, then becomes negative and significant at -0.3946 when intensity and disagreement enter beside it. Across the nine specifications in Table 3, volatility persistence declines from 0.8802 in the baseline to 0.8548 in the joint model with controls. The abstract puts the claim plainly: "We conclude that it is the sentiment and not mere volume of discussion that translates into market risk."
The authors describe this as the first frontier-market study to place intensity, disagreement and attention together in a conditional-variance equation. Its unit is a market-wide index, without a cross-section. The paper includes no forecast evaluation, out-of-sample test or trading strategy. A practitioner therefore gets a possible daily volatility-monitoring input.
A denominator too small
Mean daily tweet volume is 18.256, with a standard deviation of 13.833. The minimum is one.
A single qualifying tweet makes intensity either 0 or 1, while disagreement becomes exactly 0. Such arithmetic yields a share with almost no sample behind it. Table 2 shows that these days are no remote corner of the data. Across 1,235 observations, intensity ranges from a minimum of 0 to a maximum of 1, and disagreement does the same. The minimum daily tweet count is one, against a mean of 18.256 and a standard deviation of 13.833.
Measurement error rises as tweet count falls. Tweet count also serves as one of the three regressors, which means intensity and disagreement are noisiest on the very days when attention is low. The effect on the attention coefficient could matter, and we did not find the issue addressed.
The authors acknowledge the thin volume. They write that "investor attention is limited with average value of 18 tweets per day". That figure can reasonably describe the Pakistani retail conversation captured by their filter. Extracting a daily sentiment share asks more of it. So does asking a GARCH variance equation to price daily changes in a series with a mean of 0.538 and a standard deviation of 0.174 (Table 2).
Part of the problem appears in the paper's own limitations paragraph. The authors say that "the key limitation of our study is that it considered only a single social media platform for extracting investor sentiment". They also write that "we relied on only English language tweets", although the Pakistani investor community is mostly bilingual. The English-only filter directly contributes to the 18.256 tweets a day at the center of this review. Other platforms and Urdu posts are left for future work. The current conclusion comes from the English-only series.
Shared counts, entangled regressors
Intensity, disagreement and attention all come from the same daily positive, negative and neutral counts. Intensity is the complement of the neutral share. Disagreement multiplies the positive and negative shares that jointly form intensity. Attention sums those same counts.
The attention result needs correspondingly cautious treatment. On its own, the coefficient is -0.1599 and insignificant. In the joint model it reaches -0.3946 and becomes significant. A coefficient that changes in both magnitude and significance after correlated regressors enter carries the familiar mark of collinearity. The paper supplies no variance inflation check, correlation matrix among the constructs or orthogonalisation. We did not find a discussion of their mechanical dependence.
The authors give the reversal a substantive reading. Attention facilitates price discovery, they argue, so more eyes pull prices closer to fundamentals and lower conditional variance. Ballinari and coauthors and Wen and coauthors are cited in support. The account is coherent. Yet the same account would fit a sign flip produced by the specification, and the evidence does not separate those explanations.
Tables 3 and 4 create another gap. They print point estimates and significance markers, using an asterisk for insignificant coefficients in Table 3 and 1% stars in Table 4. We did not find standard errors, t-statistics or p-values in either table. Log-likelihoods and information criteria are also absent across the nine specifications.
Model selection instead turns on persistence falling from 0.8802 in the baseline to 0.8548 in the joint model with controls. Alpha plus beta shifts by 2.5 points, and the paper treats that movement as evidence of a better specification. Yet alpha1 remains between 0.2046 and 0.2159, while beta1 stays between 0.6403 and 0.6703 across all nine models. The GARCH dynamics move very little.
Which direction survives?
The paper's conclusion depends on exogeneity. When introducing the sentiment measures, the authors assume investor sentiment is "an exogenous shock which effects the market volatility and not the other way around". They describe this treatment as common in the sentiment literature and cite Yang, Fernandez-Perez and Indriawan (2025), the high-frequency price-impact study.
An earlier passage complicates that choice. The literature review summarises a companion paper by the same three authors, their 2024 spillover study, which found that information "mainly spillovers from volatility to sentiments rather than the reverse, implying that a significant portion of investor sentiment is volatility driven". The current paper reports that result accurately. Two pages later, its methodology assumes the opposite direction.
The reverse mechanism is straightforward. When the index falls 6%, as it does at the sample's worst point (Table 2 min return -0.0607318), the day may generate more tweets and more emotional language than a flat session. Intensity and disagreement then rise with realised volatility. A contemporaneous variance regression captures that co-movement with the reported sign, while leaving the arrow unresolved.
The paper does not attempt the tests needed to distinguish the directions. We did not find lead-lag tests, Granger causality in either direction, an event-timing split, an instrument, or an out-of-sample forecast comparison with the baseline GARCH. Daily data and 1,234 return observations make a lag structure test inexpensive.
Asymmetry does not match the summary
Both EGARCH and TGARCH assign positive sentiment a negative coefficient, -0.0106 and -0.194, both at 1%. Negative sentiment receives a positive coefficient of 0.008 and 0.095. Under these estimates, optimistic tweet counts actively lower conditional variance.
Two parts of Table 4 conflict with the paper's summary. The first is magnitude. EGARCH shows positive sentiment at -0.0106 and negative sentiment at 0.008. TGARCH reports -0.194 against 0.095. The positive-sentiment coefficient is roughly twice the size of the negative-sentiment coefficient in both specifications. Even so, the conclusion says the results "confirmed the classic leverage effect where negative sentiment drives volatility more than positive sentiment". Table 4 gives the reverse ordering. The second conflict is the asymmetry parameter, which is negative in all four models: -0.105, -0.0899, -0.191, -0.124.
The prose and table also disagree. In the results text, the "positive and significant coefficient of negative sentiment is 0.0189" for EGARCH. Table 4 gives NS = 0.008 for that model. The paper leaves the discrepancy unresolved.
Comparing the columns presents a separate problem. Positive sentiment equals -0.0106 in EGARCH and -0.194 in TGARCH. EGARCH models the log of conditional variance, whereas TGARCH models its level. Table 4 places the coefficients beside each other without rescaling. A direct comparison of -0.0106 and -0.194 therefore has no interpretation.
The conceptual contribution still matters. Raza and coauthors, like most comparable studies, use the previous day's return sign as a proxy for good and bad news. Their approach measures asymmetry generated by the market itself. Positive and negative tweet counts improve on that design by supplying an external text measure, though the exogeneity assumption remains underneath it.
The descriptive statistics carry another quiet result. Negative tweets average 5.748, compared with 5.188 positive tweets. Their dispersion is 5.850 against 4.399. During the same period, mean daily return was +0.10%. The sample's 33,309 tweets were net gloomy through a rising market.
We could not run this against our own data. The study covers the KSE-100, while we hold US equities and ETFs with no Pakistan index constituents. More decisively, the signal requires the authors' keyword-filtered Twitter corpus. Our text sources consist of news, transcripts, press releases and subtitles. Replacing that corpus with US news sentiment would change the market and the information channel, leaving nothing here to replicate.
A monitoring signal, for now
For someone tracking Pakistani market stress, the disagreement construct is worth building as a monitoring input. It arrives daily, and its coefficient has the sign predicted by theory: 1.7317 alone and 1.8323 in the joint model with controls.
The evidence that social media sentiment drives KSE-100 volatility remains an in-sample conditional-variance estimate. Its causal direction is assumed, and its three regressors share the same daily counts. The paper's literature review notes that Behrendt and Schmidt found high-frequency Twitter sentiment and activity added little for DJIA stocks. None of these nine specifications eliminates the possibility that Pakistan would look similar after the direction is tested.
One test would change my view. Re-estimate intensity and disagreement lagged one and two days, then place beside it a reverse specification where lagged realised volatility predicts the sentiment constructs. Report both. If the forward relation survives at daily lags while the reverse does not, the conclusion stands on much firmer ground. We have written before about signals that collapse under stricter specifications, in our note on the correlation rotation premium. The estimation already exists. The test is a week's work.