AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

A coincident stress index that ties the effective rank

The 0.273 win is over the Absorption Ratio; against the sharper baseline it is +0.038, p=0.10

2026-08-12 · 10 min read · US equities and US ETFs, with optional crypto

Reviewing: The Triadic Stress Index in Financial Markets · Alberto Acedo · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Point-in-Time Triadic Network-Stress Defensive Overlay

Backtest period 2016-01-01 to 2026-08-11 · hypothetical, net of modelled costs

Why these figures are not the paper's (3)

Run on a different market than the paper

The paper includes FX, sovereign debt, and commodity markets that are not directly available as the studied instruments. A backtest would restrict the correlation-network mechanism to US equities and liquid US ETFs (including commodity, Treasury, and credit ETFs where desired), because the mechanism relies on cross-asset return-correlation concentration rather than futures carry, FX basis, or bond-specific pricing; reported results from the paper's five-market universe do not transfer to this restricted universe.

The paper's own figures describe its universe and do not carry over to ours.

This is not a replication of the paper

  • The platform cannot reproduce the paper's full cross-market sample because it lacks FX spot pairs, sovereign bonds, and commodity futures. The available price history begins around 2010, so the 2006-2009 crisis-period tests cannot be replicated. Crisis labels and any external stress index used for the paper's detection F1 evaluation would need to be independently reconstructed; the proposed backtest instead evaluates a tradable de-risking overlay.

The figures below measure what we could run, not the paper's own method, so they are not evidence for or against its claim.

Our own audit found this run does not follow the paper faithfully (8)

  • deviation left undescribed by the audit (invalidates: All reported raw and filtered TSI magnitudes; F1 and unexplained-alarm results; correlations and alarm overlap with effective rank and the Absorption Ratio)
  • deviation left undescribed by the audit (invalidates: Paper cross-market crypto results including the Terra/Luna peak and MATIC-USD attribution; direct comparability with paper cross-market figures)
  • Hysteretic point-in-time allocation gate: The traded overlay uses a strictly historical expanding 90th-percentile entry threshold and an invented 75th-percentile exit threshold, whereas the paper evaluates contemporaneous alarms at each complete series' 90th percentile with no allocation state machine. (invalidates: The paper's F1@p90, precision, recall, unexplained-alarm rates, and exact 10% alarm-budget results as descriptions of the overlay's traded risk-off episodes.)
  • Screened US stock/ETF network versus the paper's datasets: The strategy constructs a yearly capitalization-ranked US stock/ETF network rather than the paper's banking, AI, OFR-component, seven-sector, crypto, FX, commodity, and ten-instrument sovereign-debt panels. (invalidates: The paper's banking and AI scenario readings, 42-asset robustness z-scores, OFR-universe F1 results, Terra/Luna and MATIC-USD attribution, FX and commodity findings, sovereign-debt peak and MBB attribution, and five-market scope conclusions as results for this screened portfolio.)

4 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2016Total 229.7%Aug 2026
Sharpe
0.87
Total Return
229.7%
Max Drawdown
-46.7%
CAGR
11.9%
Volatility
25.0%
Beta vs SPY
0.67
Trades
1,304

What the paper reports for its own strategy

  • Out-of-sample F1@p90 = 0.447 (TSI with memory, 2016-2026, 897 windows, OFR 23-window crisis list; no transaction costs — this is a detection test, not a trading strategy)
  • F1 gap vs Absorption Ratio = 0.273 (0.447 vs 0.174) out of sample 2016-2026, block-bootstrap 95% CI [0.095, 0.392], one-sided p<0.0005
  • Full-history F1@p90 = 0.389 (filtered TSI, 2000-2026, 2,232 windows, 23-window list) vs 0.342 effective rank (tie, p=0.07) and 0.175 Absorption Ratio (p<0.0005)
  • Precision 69%, recall 32% out of sample (2016-2026) at the 90th-percentile threshold on the 16-episode list
  • Precision 1.00, recall 0.29 out of sample on the 23-window list (90 alarms, none outside a labelled window)
  • Unexplained-alarm rate 4.0% (9 of 224 alarms), full history 2000-2026, 25-episode list, top-decile alarm budget

If you already run an effective-rank monitor on your rolling correlation matrix, nothing in this paper gives you a detection reason to switch. Acedo says it outright. From his Discussion: "A practitioner already running an effective-rank monitor would not switch on the strength of the detection numbers alone, and we do not suggest they should." His conclusion puts the same verdict on the scoreboard: "TSI matches a properly sharpened spectral measure at detecting stress rather than beating it." He has a rejoinder in the same sentence, and it is the reason to keep reading. What such a practitioner would gain, he says, is "the attribution layer, which their current tool supplies only after they commit to a component count and get it right." His own real-data test dissolves that layer into the plain node degree. The rank correlation is 0.996 on 464 S&P 500 names, and I come back to it below. What survives is worth pricing, and it is much smaller than the abstract's three favourable comparisons make it sound.

How it is built

The Triadic Stress Index is a single composite number computed on a correlation network. Take 20 trading days of daily log-returns for n assets. Build the correlation matrix, set the adjacency to the absolute correlations with a zero diagonal, and recompute every three days. TSI multiplies four things: weighted clustering (a normalised count of closed triangles, Tr(A^3)), connection density, degree-sequence variance, and the inverse spectral gap of the graph Laplacian, which sits in the denominator. The construction was lifted unchanged from soil-microbiome co-occurrence networks, with the orientation flipped. A dense weakly modular community is health in soil and everything-falls-together in a market. On top of the raw series sits an asymmetric exponential filter, fast up (alpha 0.6) and slow down (alpha 0.08), so a spike becomes a plateau.

There is no trade in the paper and Acedo says there is not one: "everything in this paper is a backtest of a state reading." The claim is a monitor. Detection is scored as F1 at the 90th percentile of each series against labelled crisis windows, which fixes every metric at the same 10% alarm budget. Intervals come from a block bootstrap (block roughly a quarter, 2,000 resamples), with Holm-Bonferroni over a pre-declared family of 24 tests. The main sample is the 8 components of the OFR Financial Stress Index, 2000-2026, 2,232 windows, with 897 windows held out for 2016-2026. Four other panels run alongside it. Three are equity: 10 US banks 2006-2011, 11 AI-sector stocks 2021-2026, and 42 stocks across 7 sectors 2006-2026. The fourth is a cross-market set of 10 crypto plus 7 FX plus 10 commodities plus 10 sovereign-debt instruments, 2021-2024.

Five baselines. The Absorption Ratio is the share of return variance sitting in the top n/5 principal components. The effective rank is the exponential of the entropy of the correlation eigenvalues, a count of how many directions the matrix effectively has. Then the Vendi score, Ollivier-Ricci curvature and the balance index of signed correlation networks.

The reported headline is threefold: an out-of-sample F1 of 0.447 against 0.174 for the Absorption Ratio, a parameter-free per-node attribution (diag(A^3)) naming the asset carrying the concentration, and the cleanest alarms of anything tested. The register is declared in the same abstract: peak cross-correlation with the OFR index at zero lag, so a coincident state index rather than a forecast.

The 0.273 is a gap against the weakest thing on the table

That headline number pairs filtered TSI at 0.447 with the Absorption Ratio at 0.174, which is the AR's raw reading. Acedo also reports the like-for-like column, where every series gets the same filter. AR falls to 0.134. The bootstrap difference is +0.219, 95% CI [+0.070, +0.410], p<0.0005. Either way the AR loses badly, and either way the reason is the AR. It scores 0.134 to 0.204 across the six labelling-by-sample cells. Out of sample the paper puts it at an F1 of roughly 0.17 to 0.19, "roughly half of every other measure tested". The benchmark section calls it "a weak baseline on this data."

Unpack the 0.447 before deciding what it buys. On the 23-window list, out of sample, the same filtered series fires 90 alarms and none of them falls outside a labelled window: precision 1.00, recall 0.29. That list calls 34.9% of windows stress. So the headline number describes a monitor that misses roughly seven in ten labelled stress windows and is never wrong when it does speak.

Two further dents. The AR win disappears under the narrowest of the three crisis lists. The gap stays positive there (+0.170 and +0.121), but the interval touches zero. It is also one of three results that Holm-Bonferroni strips out of the family of 24, which goes from 16 significant before correction to 9 after. On a wide panel the benchmark cannot be scored at all. With a 60-day window and 464 assets the covariance rank is 59, so the top n/5 = 93 eigenvalues capture the whole trace and the ratio is identically one.

Acedo attributes that last failure to the configuration rather than to the measure.

Does the topology add information the spectrum already carries?

Against the effective rank the answer is no, and the paper is straight about it. Out of sample, filtered, 0.447 against 0.377: delta F1 = +0.038, 95% CI [-0.010, +0.111], p=0.10. Raw, the pair is 0.367 against 0.347, a lead of 0.020; filtered the lead is 0.070. Neither interval excludes zero, so the tie does not depend on the filter. Full history, 0.389 against 0.342, p=0.07. The tie holds in all six cells, with TSI ahead in five and behind by 0.023 in the sixth (p=0.59). TSI correlates 0.85 with the effective rank (Spearman 0.93) and shares 53% of its alarm windows. Acedo also declares that the benchmark family is thinner than four columns suggest. The Vendi score is mathematically identical to the effective rank on a correlation matrix. The AR and the effective rank on returns are near-duplicates at Pearson 0.91.

The unexplained-alarm result reads well at first: 4.0% of TSI's alarms (9 of 224) match no documented episode, against 14.7% for the effective rank and 59.4% for the AR over the full history. Acedo then says the quiet part out loud. At a fixed alarm budget the unexplained-alarm rate is exactly 1 minus precision. F1 at a fixed alarm count is a monotone function of the same quantity. So this is the F1 test in different clothing, and it carries the same verdict. The 4.0% is also computed over the full history on the laxest of his three lists.

The sharpest criticism in the paper is his own. A shape-normalised variant replaces degree variance by a Gini coefficient and density by mean degree, stripping out the scale of the degree sequence. It scores 0.098 against a chance level of 0.152 on the bank panel. His reading: "a substantial part of what TSI detects in this setting is the level of correlation rather than the shape of its distribution." If the detector is largely reading average absolute correlation, a tie with a spectral entropy is the expected outcome.

Attribution that holds on synthetics and dissolves in the S&P 500

On synthetic networks with known epicentres (120 simulations per regime, 16 nodes, 40-observation windows), diag(A^3) scores 0.975 to 0.994 across one to four epicentres. A fixed k=2 spectral read gets 0.756 when there is one epicentre. The Kaiser rule collapses to 0.000 at one epicentre (0.071, 0.457, 0.808 at two, three, four). The node degree gets 0.933 to 0.961. Against the local balance index of Bartesaghi et al., the only published per-node alternative on correlation networks, triangles score 0.992/0.996/0.972/0.966 versus 0.329/0.552/0.724/0.841. Rebuilding the generator with negative loadings so frustrated triangles exist does not move the ordering.

The abstract's defence of this layer is narrower than accuracy. It says spectral attribution "must first choose how many components to read and collapses under a standard but wrong choice." True of the Kaiser rule. False of the Marchenko-Pastur edge, where the spectral rule scores 0.959 to 0.997 and ties everywhere, picking k=1.71 on average. So the surviving claim is that there is no selection rule to get wrong.

Then the real data. Across 245 windows of 464 S&P 500 constituents, the cross-sectional rank correlation between diag(A^3) and the plain weighted degree averages 0.996. A cubic in the degree explains R^2 = 0.992 of the triangle count. The 0.8% residual carries nothing. Its rank correlation with forward five-day returns is +0.002, with a bootstrap interval spanning zero. Its apparent +0.187 with forward volatility falls to +0.016 ([-0.003, +0.035]) once current volatility is partialled out. Acedo reports all of this against his own claim, and tells practitioners to use the degree, which costs one matrix row-sum. Credit where it is due, and note what it leaves: the per-node layer buys convenience and nothing else.

Register decides the use case

Raw TSI's cross-correlation with the OFR index peaks at lag zero, rho = +0.203 on 2,232 windows; the filtered version peaks at k = -2. On first differences the peak sits exactly at k=0 (+0.094). At k=+1, where a leading indicator would show something, the correlation is -0.009 and +0.008 against a significance threshold of plus or minus 0.041. On the 464-asset panel the index's orientation flips sign between contemporaneous and forward-span labels. A monitor, then, and Acedo frames it that way throughout.

We built the trade the paper refused to build

We could not reproduce the paper's sample. Our platform has no FX spot pairs, no sovereign bonds and no commodity futures, so the cross-market panel is out of reach. Usable price history begins around 2010, so the 2006-2009 crisis tests cannot be run at all. What the paper reports across its five markets does not transfer to a US-equity-and-ETF-only universe, and nothing below should be read as if it did. Crisis labels and the OFR series behind the detection scores would have to be independently reconstructed. Instead of a detection test we built the thing the paper explicitly declines to build: a tradable de-risking overlay on US equities and liquid US ETFs. This is an adaptation. We could not test the paper's claim.

We implemented TSI from the paper's definition, filling in several construction details it does not pin down. Two deviations in our build remain unresolved. The universe is 42 point-in-time capitalisation-ranked US stocks and ETFs. We rank the filtered series against an expanding point-in-time history, go 25% equity / 75% MBB when it breaches the 90th percentile, and return to 100% equity below the 75th. Execution carries a one-day lag. 2016-01-01 to 2026-08-11, 10 bps per one-way turnover plus $0.004 a share, zero modelled slippage.

Our numbers: +229.67% cumulative, Sharpe 0.87, max drawdown -46.66%, Calmar 0.26, beta 0.67 to SPY, 1,304 trades. Acedo's published figure is a detection F1@p90 of 0.447 on 897 out-of-sample windows with no costs by construction. His 0.447 and our 0.87 measure different quantities on different universes, and no mapping between them exists. He says a trading rule "is a different exercise with a different evidential standard."

What we can explain about the difference is definitional first. Everything that turns a state reading into a book (the thresholds, the hysteresis band, the lag, the defensive leg) is ours. Beyond that, the paper's own zero-lag result predicts what our drawdown shows. A coincident reading plus a one-day lag and a three-day recompute cycle de-risks after the correlation spike is in prices. Hence the 46.66% peak-to-trough at a beta of 0.67. Our gate is an expanding percentile, so early in the sample the 90th-percentile trigger is set on very few observations. As we read his protocol, the threshold there is the 90th percentile of the whole series, which cannot be computed live. Our universe includes SPY, IVV, VOO, QQQ and VUG, whose overlapping holdings mechanically raise pairwise correlation, density and clustering. So our alarm dates are not his alarm dates. And the defensive leg is MBB rather than cash. In our judgement that removed the hedge during the 2022 rate stress, the same bond selloff his own sovereign-debt panel flags at z=3.6. We cannot close the gap quantitatively. We have no F1 for our own series, and two implementation deviations remain unresolved in our run. This is evidence about our build before it is evidence about anything else.

It could still earn a slot as a second monitor

On a concentrated book, alongside an effective-rank read, on the strength of the 47% of alarms it does not share with its nearest relative. Use the filtered version. Raw TSI collapses on a basket concentrated in the epicentre, reading z = -1.09 on the 10-bank panel during the European debt crisis. The memory filter repairs that. Acedo shows the driver is concentration rather than size, since a 9-asset diverse basket reads +5.35 and the 42-asset panel +2.13. The pattern replicates at z = 1.35 in COVID and z = 1.41 in the 2022 selloff. Two scope failures are declared: the September 2022 gilt and sterling crisis, never flagged, and the 2022 Ukraine invasion, about three months late.

The calibration rests on one 26-year series with roughly 20 episodes. The filter parameters were chosen by inspection, then grid-searched on that same series. In-sample and out-of-sample F1 correlate at 0.96 across all 36 combinations. One 26-year series is a thin base for a live threshold. Raw magnitudes are also not comparable across networks of different size, and only within-series z-scores are usable.

One result would change my view of the triangle count: a real-data attribution test in which the part of diag(A^3) not explained by degree carries forward information. Acedo lists that test as open. After his own 0.996 he says the expected outcome is fairly clear.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

At each yearly universe refresh:
  Select 42 point-in-time US stocks/ETFs by descending capitalization.
  Exclude ADRs, nonpositive-liquidity names, and candidates lacking real price history.
  Replace exclusions with the next eligible candidate; do not impute returns.

Every third trading day:
  1. Calculate 20 daily log returns for each of the 42 assets.
  2. Form Pearson correlation matrix rho and adjacency A = abs(rho), with diag(A)=0.
  3. Compute weighted degrees k_i = sum_j A_ij.
  4. Compute:
       C = Tr(A^3) / [n(n-1)(n-2)]
       D = sum_{i != j} A_ij / [n(n-1)]
       Coex = population_variance(k)
       L = diag(k) - A
       gap = lambda_2(L)
       M = 1 / gap
       TSI = (C * D / M) * Coex
     If gap is nonpositive or unavailable, retain the existing portfolio state.
  5. Apply asymmetric memory:
       alpha = 0.60 if TSI rises, otherwise 0.08
       memory_t = alpha * TSI_t + (1-alpha) * memory_(t-1)
  6. Rank filtered TSI against only prior and current recomputation observations using an expanding percentile.
  7. If risk-on and percentile &gt;= 90%, schedule risk-off.
     If risk-off and percentile &lt; 75%, schedule risk-on.
     Otherwise retain the current state.

At the scheduled close after the one-trading-day execution lag:
  Risk-on: 100% equal-weight eligible equities, 0% MBB.
  Risk-off: 25% equal-weight eligible equities, 75% MBB.
  Cap each equity at 10%; MBB is separately capped at 75%.

Between recomputations, carry the last signal and allocation state forward.