The tradeable claim rests on 26 tail observations, and its left-tail direction changes with the proxy. The paper's identification result is stronger. When the cause has the lighter tail, the reverse coefficient is point-identified at 0.5 instead of confined to the interval (0.5, 1). A trader should treat the S&P 500 to Bitcoin application as a diagnostic.

A disclosure belongs before the finance results. The paper pairs Bitcoin with the S&P 500 index, which we cannot hold. We would use the index level as the signal, trade SPY for the equity leg and trade Bitcoin directly. All tail quantities below belong to the paper and were measured on its index sample. Nothing here is a replication.

The coefficient and its limits

The Causal Tail Coefficient, from Gnecco and coauthors, measures the limiting expected rank of one variable when the other lies in its extreme upper tail. Its rank construction makes it invariant to monotone marginal transformations. Within a linear structural causal model with heavy-tailed innovations, the coefficient equals 0.5 when the variables have no directed path between them. It equals 1 when one drives the other.

Direction comes from the asymmetry between the two coefficients, Delta. The estimator uses the top k order statistics and is nonparametric.

Earlier theory typically gave every variable the same tail index. Leimenstoll and Schienle allow unequal tails, producing three identification cases under no confounding. A lighter-tailed cause (alpha1 > alpha2) forces the reverse coefficient to exactly 0.5, with Delta = 0.5. Equal tails leave the reverse coefficient somewhere in the interval (0.5, 1). When the cause has the heavier tail, both coefficients equal 1 and direction disappears.

Hidden confounders are then divided by tail weight. Light-tailed confounders vanish asymptotically. Heavy-tailed confounders (alpha_H below both) drive both coefficients toward 1. The abstract acknowledges the problem directly: such confounders "can induce extremal dependence patterns that are observationally indistinguishable from direct causal effects".

Following Pasche and coauthors, the proposed remedy is a proxy-adjusted coefficient. It fits a generalized Pareto scale that is linear in an observed proxy H. The accompanying Confounder-Test statistic is sqrt(k)(hatGamma - 0.5), with an N(0, 1/12) null. A Hill-based pre-test checks for unequal tails.

The Causality-Test records size of 4.3 to 5.8 percent against a nominal 5. With n = 10,000 and k = 39, power under a causal link ranges from 87.0 to 100.0 percent. Confounding makes the plain test over-reject. Rejection reaches 19.8 percent with a light-tailed confounder and 34.1 to 54.4 percent with a heavy-tailed one. Adjustment reduces those ranges to 3.7 to 7.0 and 8.2 to 16.8 percent respectively.

Where the evidence comes from

Three applications follow. The climate examples pair Swiss train delays with Zurich precipitation across 3,994 observations, then Danube and Main discharge with local rainfall across 3,072 and 2,948. Finance uses ARMA-GARCH residuals from open-to-close log returns on the S&P 500 and Bitcoin. The sample runs from 18 July 2010 to 4 December 2024 and contains 3,620 observations after non-trading days are removed.

The finance evidence comes down to 26 tail observations.

In the right tail, the coefficient from S&P to Bitcoin is 0.7689, versus 0.5088 in the other direction. Delta = 0.2601 exceeds a bootstrap critical value of 0.1550. The Confounder-Test is 0.0449 and indicates no heavy-tailed confounder.

The unadjusted left tail is weaker. Delta = 0.0862 falls below 0.1856, while the Confounder-Test rejects at 0.5283. Conditioning on MSCI Europe residuals raises Delta to 0.2125 against 0.1678. Classical LiNGAM finds nothing here at all. Pairwise LiNGAM agrees with the direction.

Unequal tails earn their place

The paper's lasting contribution is the sharper null. A lighter-tailed cause pins the reverse value to 0.5, where the equal-tail literature could identify only the interval (0.5, 1). The climate analysis also shows the full procedure operating rather than leaving it as theory.

For the Main, the plain test produced Delta = 0.1860, above the 0.1373 critical value. The Confounder-Test then detected confounding at 0.8853. Conditioning on the upstream catchment at Schweinfurt raised Delta to 0.3387, above 0.1719, and lowered the confounder statistic to -0.3000. The authors found the failure and repaired it. Every judgment that follows concerns the repaired estimator.

The Danube result separates tail propagation from mean propagation. Delta drops from 0.2717 at one day to 0.0025 at four. Across the same lags, LiNGAM continues to report structural weights from 25.55 down to 12.45 out to eight days. Tail propagation is faster than mean propagation.

Why the left tail stays open

The adjusted coefficient assumes its proxy captures everything relevant. The authors describe conditioning on H as accounting for all relevant confounding effects, calling it a rather strong proxy of the true confounder. Their defence is the Confounder-Test. They present it as the way to assess whether adjusted CTC-type tests work, since those tests require the proxy to capture all relevant confounding.

The finance application exposes the weakness in that defence. MSCI Europe produces a confounder statistic of 0.1142, while VIX produces 0.2548, so both proxies pass. Their directional answers diverge. With MSCI Europe, causal asymmetry reaches 0.2125 against a critical value of 0.1678 and clears it. With VIX, the corresponding result is 0.1549 against 0.1834 and fails. The certification device cannot distinguish the proxy that produces the finding from the proxy that removes it.

The rate condition adds another problem. Theorem 4.4 covers a causal link where the cause is the lighter-tailed variable (alpha2 < alpha1). It requires k = floor(n^nu), with nu below 2rho/(alpha2 + 2rho). For the two left tails, the paper reports Hill estimates of 0.2301 for Bitcoin and 0.0804 for the S&P 500. These estimate the observed series. As the authors state, the innovations are unobservable, leaving alpha1 and alpha2 to be inferred from the tails of X1 and X2.

Our own arithmetic on those two figures places alpha2 near 4.35. The gap then exceeds one, so rho = 1, giving a bound of roughly nu < 0.32. Every application test uses nu = 0.4. Leimenstoll and Schienle identify this exact mismatch in simulation for the (4,3) case, where type-I error approaches nominal only as k shrinks. The rejection that prompts their left-tail adjustment comes from that same statistic.

Simulation offers little comfort. Under heavy-tailed confounding, the adjusted test's size was 8.2 to 16.8 percent even with the true confounder observed. The reported left-tail Delta of 0.2125 clears a critical value calibrated to 5. The authors also concede that causal tail tests may be less reliable in small samples or when confounding effects are strongly asymmetric. The finance application has both features.

Use on a trading desk?

Both return series measure the same day's open-to-close move. The finance result therefore describes contemporaneous extremes, ordered through tail asymmetry. The paper runs no lead-lag test between the return series; only the VIX proxy is lagged a day. Its conclusion also identifies the lack of methods for quantifying causal-effect magnitude as a significant gap.

Direction without magnitude or horizon belongs in stress monitoring. It could inform a prior for crypto exposure caps when equity left tails open up, or set the ordering used in scenario design. The paper supplies no sizing rule.

The evidence would become more persuasive if the asymmetry survived a one-day lag on the Bitcoin side with k inside the theorem's stated rate bound. The left-tail claim also needs to hold under a proxy fixed in advance, rather than one certified by the same Confounder-Test that triggered the adjustment.