Changing every ingredient in HRP's clustering step leaves the crypto book's P&L unmoved. Across 547 tuning cells, the clustering stage is inert, a negative result more useful than most positive findings. One disclosure belongs ahead of the figures: we could not rebuild the paper's point-in-time Binance universe, so the results below do not test the authors' method.

Start with the calculation. Hierarchical Risk Parity turns a correlation matrix from a rolling window of daily returns into a distance. The baseline uses the standard sqrt((1-rho)/2). Hierarchical linkage then rearranges the covariance matrix, placing correlated assets next to one another. The tree divides capital from the top down, using inverse-variance weights at each bisection, with no matrix inversion. Crypto makes the attraction easy to see. Correlations shift, covariance estimates are noisy, and a method that depends on ordering rather than an inverse should degrade gracefully.

Matys, Rodriguez and Delfau vary the opening stage in eight ways. They test Marchenko-Pastur eigenvalue denoising, spectral detoning, a graphical-lasso partial correlation, EWMA correlation at lambda = 0.94, empirical lower-tail dependence at q = 0.05, Ledoit-Wolf shrunk covariance, returns standardised by trailing 30-day vol, and a tail-dependence plus shrinkage hybrid. Spectral detoning removes the top eigenvector, which is the market mode in long-only crypto.

They also introduce four learned distances. Each one embeds an asset's return path, then replaces correlation distance with cosine distance between embeddings: level-3 path signatures, node2vec on a k-nearest-neighbour correlation graph, an NT-Xent contrastive encoder, and TS2Vec. Across all twelve variants, inverse-variance recursive bisection still sets the within-cluster allocation. The dendrogram alone changes.

The paper tests daily Binance USDT spot from 2020-02 to 2026-05-18, covering 76 monthly rebalances and 2,299 daily return observations. Each month, the authors rebuild the universe from rolling 30-day quote volume, selecting the top 50 after age and liquidity filters, with entry and exit buffers. In total, 146 symbols appear, and 97 leave the universe before the end. This addresses survivorship in the authors' 2025 conference dataset, which used 50 currently-trading names. Estimation relies on a 365-day rolling window and a 180-day minimum-observation filter. Headline trading costs are a 10bp fee plus 2bp slippage, doubled when exiting positions are forcibly liquidated.

Thirteen variants inside one narrow band

Every HRP-family variant finishes between 0.701 and 0.749 annualised Sharpe. Baseline HRP records 0.741. EWMA leads at 0.749, while the tail-dependence-plus-shrinkage hybrid comes last at 0.701, despite being flagged, following the literature, as the strongest theoretical extension for crypto. All four learned embeddings trail the baseline: path signatures at 0.714, contrastive at 0.710, node2vec at 0.705 and TS2Vec at 0.703.

Then comes the tuning sweep. The authors run a 547-cell grid spanning correlation estimator, distance function, linkage and each variant's parameters, with a complete walk-forward rerun for every cell. None leaves the original band. Ledoit-Wolf shrunk covariance under average linkage produces the grid's best result, 0.765. Its delta is +0.024, with a pairwise Ledoit-Wolf p of 0.16. Across all 547 cells, the combined Hansen SPA returns p_consistent = 0.906. The single-channel path-signature grid contains 108 cells and beats baseline in zero of them.

The authors' explanation is persuasive. A return-only embedding carries information equivalent to the correlation matrix it is meant to improve. Adding channels raises the encoders' mean Sharpe across cells monotonically. The contrastive encoder moves from 0.714 at one channel to 0.725 at three, still below 0.741. Two multi-channel path-signature cells nominally exceed baseline, with the best reaching 0.747. Both drop to 0.735 when mean-fill replaces zero-fill imputation, a correction the authors report themselves.

They disclose a graphical-lasso scale bug as well. It had silently reduced the partial-correlation variant to inverse-variance. An earlier cluster-stability claim is retracted outright. Across all 76 snapshots, Pearson clustering proves more stable, with adjusted Rand index, the cluster-agreement score between consecutive snapshots, of +0.744 versus tail dependence's +0.614. The former claim depended on a single year-pair.

Linkage supplies the grid's one useful engineering result. Average linkage beats López de Prado's single-linkage default by +0.01 to +0.02 across essentially every combination. Under single linkage, all four distance functions are mathematically identical because the method reads only the rank order of distances.

Why allocation moves while clustering stalls

Four allocators beat baseline HRP by a wide margin. MVP, the minimum-variance portfolio, reaches 1.057. CRISP, a correlation-shrinkage solve interpolating between inverse variance and a full minimum-variance solve, reaches 0.989. NCO-CRISP, the hybrid of the two, posts 0.977, while NCO, nested minimum variance within and across clusters, posts 0.933. XGBoost-fed NCOML also finishes ahead at 0.804, though it remains below signal-blind NCO's 0.933.

Concentration explains the ordering. Their effective asset counts are 3.3, 7.1, 9.1 and 5.3, compared with 29.8 for baseline HRP. HRP's largest holding averages 11% of the book, against MVP's 50%. Passive BTC delivers 0.868 and +749% over the same window, as BTC rises from about $7k in early 2020 to about $107k by mid-2026. Ranking the winners by concentration almost exactly recovers their Sharpe order. The authors state the result plainly in the conclusion: a passive 100% BTC hold beats every diversified portfolio, while the constructed strategies that surpass it concentrate back into BTC.

The post-COVID cut is their defence, and it warrants a direct answer. From 2022-01-01, baseline HRP records Sharpe of 0.047 and loses 60% of capital. Every HRP-family variant loses 51% to 64%. CRISP and NCO-CRISP remain pairwise significant against HRP at p = 0.017 and 0.022, each returning +20%. Allocation therefore survives a regime in which the dendrogram variants fail.

Passive BTC still leads over that same period, returning +106% at Sharpe 0.585. MVP follows at 0.482 (+66%), ahead of NCO-CRISP at 0.368 and CRISP at 0.363 (+20% each). NCO reaches 0.264 (-4%). The post-COVID SPA is 0.38.

The abstract makes both concessions the evidence demands: the combined sweep SPA is 0.91, and diversified HRP variants lose 51-64% of capital post-COVID while BTC and BTC-concentrating allocators fare better. The authors describe the surviving edges as "remain statistically fragile under search-adjusted inference". Limitation 1 explains the problem. Six years of data cannot resolve a Sharpe edge below about 0.4 for a strategy that departs substantively from HRP. NCO's +0.19 therefore remains unresolved, neither confirmed nor refuted. Non-rejection leaves refutation unsupported.

Inference turns on the winners

The paper's strongest section argues against its own leading strategies. Hansen SPA across the 27 candidates yields p_consistent = 0.51. Using N = 27, the Deflated Sharpe places the expected best-of-N at 0.37 and clears only MVP at 0.97, aided by skew of 5.0 and kurtosis of 128. Once the sweep counts as part of the search, the hurdle rises to 0.56. Every strategy falls below it, including MVP at 0.92.

Bootstrap Sharpe intervals span about 1.5 units. HRP's interval is [0.02, 1.53], and no pair of strategies has non-overlapping intervals. Quadrupling to weekly rebalancing raises every Sharpe, with HRP moving from 0.74 to 0.82 and NCO from 0.93 to 1.03, yet inference remains unchanged: SPA is 0.57 and the intervals are still 1.5 wide. Across the four cost scenarios, the five strategies included in the cost grid shift by under 0.01 Sharpe. Baseline HRP moves from 0.749 at zero-cost to 0.740 under stress.

The power analysis contains a point worth retaining. Minimum detectable edge increases with a strategy's distance from HRP. For the near-collinear family variants, it ranges from 0.04 to 0.12. MVP, with effective N of 3, requires roughly 0.47, while CRISP requires 0.27. Concentration inflates the variance of the paired difference, so MVP's larger raw edge of +0.314 misses pairwise significance at p = 0.059, whereas CRISP's +0.247 passes at 0.012.

Reproduction failed

We could not reproduce the study. The universe is the binding constraint. Our data covers a smaller crypto set through 2024 and cannot reconstruct the paper's point-in-time Binance USDT listing status or monthly quote-volume ranks. The paper's universe covers 146 ever-included symbols, of which 97 disappear during the window.

We also lack the order-flow channels used by the multi-channel embeddings: trade count, average trade size and taker buy-ratio. Standard-library stand-ins would have to replace the contrastive encoders and path-signature implementation, which would produce different trained models from the authors'. A substituted universe would test a different strategy.

On the paper's own figures, the negative finding is the lasting result, and the authors earn it. Across 547 configurations, four learned embeddings and a documented retraction, clustering is where crypto portfolio effort goes to die. Every reported Sharpe also assumes a zero risk-free rate, which the authors identify as crypto convention, during a period with materially positive USD rates.

Allocation offers more. CRISP and NCO-CRISP beat passive BTC over the headline window, scoring 0.989 and 0.977 against 0.868, with +1693% and +1613% against +749%. They achieve those figures by concentrating back into BTC. Evidence from a window where they outperform a passive BTC hold without that concentration would change my view, provided the sample is long enough to resolve an edge that the authors themselves say six years cannot.