The 43.04% versus 41.45% Bitcoin contrast cannot identify amplification because the sample period and VAR size change together. The stress-index result deserves attention.
Erben Yavuz builds a Diebold-Yilmaz connectedness study. A VAR is estimated across a set of financial series, followed by the generalized forecast error variance decomposition in the Pesaran-Shin version, which removes sensitivity to variable ordering. The exercise measures how much of each variable's forecast error comes from shocks elsewhere in the system. It then tracks which markets send shocks, which receive them, and how those roles move week by week across asset classes.
Adding the off-diagonal shares and dividing by the number of variables produces the Total Connectedness Index. The asymmetric version gives directional spillovers: how much each variable sends (TO), receives (FROM), and the difference (NET). A positive NET marks a net transmitter of shocks.
All variables are weekly. The St. Louis Fed Financial Stress Index (STLFSI) enters in levels. The paper calls it an index "which integrates information from credit, money, and equity markets" into a single standardized stress measure. The remaining inputs are S&P 500 log returns times 100, the weekly log difference of VIX, the first difference of the U.S. 10-year Treasury yield in percentage points, gold log returns, and, in the second sample, Bitcoin log returns. The settings are VAR lag p=4, forecast horizon H=10, and a rolling window fixed at 104 weeks.
The long sample covers January 2000 to December 2024 without Bitcoin. The subsample covers 2010 to 2024 with it. Stress data come from FRED and everything else from Bloomberg. Augmented Dickey-Fuller unit-root tests reject a unit root for all six transformed series, with the strongest result reported for S&P 500 returns at -29.865.
The paper contains neither a portfolio nor an out-of-sample test. We found no Sharpe and no turnover figure anywhere in it. This is a monitoring exercise and should be assessed on those terms.
A stress sender and a volatility receiver
The ranking of net spillovers is the result worth keeping. It barely changes across the two specifications. STLFSI is the largest net transmitter, at +34.26 in the sample without Bitcoin and +29.19 in the sample with it. The 10-year yield change follows at +11.63 and +14.28. S&P 500 returns are the largest receiver at -24.78 and -26.02, followed by VIX changes at -15.69 and -19.25. Gold records -5.42 and -3.58.
The VIX entry needs a careful reading. Across the full system, the stress index sends while implied volatility receives: +34.26 against -15.69 without Bitcoin, then +29.19 against -19.25 with it. Tables 3 and 4 report these system-wide totals, averaged across every other variable in the VAR. They do not print a pairwise STLFSI-to-VIX cell, so the ranking says nothing about which series moves first.
The paper describes the roles this way: "the equity market and the VIX primarily act as net receivers of risk, while gold assumes a modest risk-absorbing role." Its substantive claim is that a broad credit and money-market stress index occupies the transmitting side of the U.S. cross-asset system, while implied equity volatility occupies the receiving side. That claim matters.
A mechanical caveat complicates the interpretation. STLFSI enters the VAR in levels, whereas every other series appears in differences or returns. My reading is that a persistent level series beside near-white-noise return series will explain some forecast error variance because it is autocorrelated and the others are not. The paper defends the levels treatment by noting that STLFSI is published as a standardized measure. Standardization does not settle the variance-decomposition issue. Part of the +34.26 may reflect stress, and part may reflect persistence. The paper does not separate them.
Standard errors are also absent for every net spillover estimate. I therefore cannot say how much of the difference between +34.26 and +11.63 would remain inside a sampling-error band.
What does the Bitcoin comparison identify?
The Bitcoin conclusion comes from one cross-sample comparison. Rolling maximum connectedness is 83.74% versus 80.27%, with both peaks dated to COVID. The amplification claim relies directly on that maximum: "The higher maximum connectedness observed in the sample including Bitcoin (83.74%) further implies that cryptocurrencies may amplify systemic interactions during crisis episodes."
Three features change between 80.27% and 83.74%. Bitcoin enters the system. The start date shifts from 2000 to 2010, removing the 2008 crisis and its 2009 aftermath entirely. The VAR also expands from five variables to six, changing denominator N in the connectedness index and the number of off-diagonal cells being averaged. A 3.47-point difference in a rolling maximum is carrying all three changes. The 1.59-point difference in the rolling averages has the same problem.
The paper recognizes the timing issue and invokes availability: "Bitcoin is available only from 2010 onward; therefore, two separate samples are analyzed to ensure methodological consistency." Availability determines the 2010 start for the Bitcoin subsample. The 2000 start for the baseline remains a choice. The confound enters there, and the paper could have avoided it.
Descriptive statistics show how different the periods are. In the long sample, STLFSI has a mean of 0.044, a standard deviation of 1.114, and a max of 9.630. Over the post-2010 subsample, its mean is -0.237, its standard deviation is 0.573, and its max is 5.646. The second period is calmer by half on the dispersion measure, while its worst observation lies four points below the first period's maximum.
Assigning the TCI difference to Bitcoin requires the missing 2008-2009 window to have generated the same connectedness path. The paper's own account of crisis spikes suggests otherwise.
A clean comparison is inexpensive. Estimate the five-variable system over 2010-2024, then compare it with the six-variable system over 2010-2024. We did not find that specification in the paper.
Crisis dates through a 104-week window
The paper states that "the index reaches a peak of approximately 80% during the COVID-19 period (July 2020)." A July 2020 estimate from a two-year rolling window uses observations beginning around mid-2018. The peak exists, though the attached date is the window-end label rather than a point observation.
The same timing qualification applies to the 2022 tightening episode. Figure 3 shows Bitcoin's net spillover becoming positive, and the text calls this "a second period of positive net spillovers... during the global monetary tightening cycle beginning in 2022."
Convention supplies the stated rationale for the window length. According to the paper, "the choice of a two-year rolling window is consistent with previous studies employing rolling-window connectedness methods and is widely regarded as appropriate for identifying medium-term changes in financial spillover dynamics." The choices p=4 and H=10 are asserted without a comparable argument. No sensitivity check is reported for window length, lag order or horizon.
The authors reject TVP-VAR on transparency grounds and cite prior evidence that the approaches agree, specifically Gabauer and Gupta (2018) and Antonakakis et al. (2020). Citation makes the choice defensible. The paper never estimates both methods to show that their paths overlap.
One descriptive statistic should halt a referee. Long-sample gold returns have a reported standard deviation of 145.808 and a range from -694.853 to +695.790. The subsample gives 196.746 with the same range. These columns supposedly show weekly log returns times 100. Bitcoin returns have a standard deviation of 148.053 and a minimum of -869.087.
Weekly gold does not move 695%.
The transformation behind these columns differs from the one described in the text. If those same series entered the VAR, the variance decompositions carry the problem forward. The paper's S&P 500 column provides the internal comparison: its standard deviation is 2.487, the scale expected for a correctly transformed weekly return series in this table.
Bitcoin after the confound is removed
One Bitcoin finding survives without the cross-sample contrast. Inside the 2010-2024 system, Bitcoin's average net spillover is +5.38. The positive sign makes it a net transmitter on average, which fits poorly with an unconditional safe-haven interpretation. Bouri et al. (2017) appears on both sides of the same sentence, first as the safe-haven study from which the result "partially diverges" and then among the studies with which it agrees.
The size is small. Against STLFSI at +29.19 and the 10-year yield at +14.28, Bitcoin's +5.38 is roughly a fifth of the stress index's transmission. Bitcoin is the system's weakest transmitter. Gold at -3.58 is a weaker receiver than Bitcoin is a transmitter. Without a confidence interval, I cannot tell whether +5.38 is distinguishable from zero.
The abstract already concedes the fair comparison. Full-sample connectedness is 21.34% without Bitcoin and 21.17% with it, leaving a gap of 0.17 percentage points. The abstract says that "while Bitcoin does not fundamentally alter long-run average connectedness, it significantly amplifies connectedness dynamics during short- and medium-term stress episodes."
The clause after the comma carries more weight than the analysis can support. "Significantly" has no reported standard errors, confidence intervals or significance tests behind it for any TCI or net-spillover estimate. The concession rests on a 0.17-point gap measured consistently in both samples. The amplification claim relies on rolling figures drawn from different periods and different N.
We could not reproduce any of this. STLFSI is the pivot of the entire system, and our macro data feeds do not contain it. Our price history also begins around 2010. The 2000-2024 baseline and its 2008 results remain out of reach regardless.
Estimate both systems over the same 2010-2024 window, add bootstrapped bands around the TCI difference, and correct the gold column. If the gap survives those changes and Bitcoin keeps its sign, the amplification claim can be judged.