A Hyperliquid maker can identify the counterparty from the tape before the fill arrives, and that identity improves short-horizon forecasts. Daojing Zhai measures the value directly. Features derived from the 231 wallets in the top markout decile raise one-second out-of-sample R2 by 1.43 percentage points across the 100ms grid. The prespecified benchmark uses anonymous quotes, prices and order flow. Those 1.43 points are the result worth debating. Confidential account-level data had already shown that information and performance persist among particular traders. Here, the label is public, permanent and free.
The venue publishes the trader
Hyperliquid uses a conventional price-time priority book whose full input is published by its consensus protocol. Every committed submission, cancellation, rejection and fill includes the persistent pseudonymous wallet address responsible for it, permanently and at no cost. Zhai operated a non-validating node, replayed messages in consensus order and reconstructed the full-depth book, linking every order to its wallet.
The July 2026 record spans the ten most active perpetuals. It contains 17.09bn Level-4 messages, 14.27m aggressive orders from 147,113 wallets and $84.3bn of taker notional. The mean pre-trade quoted spread is 0.92bp. Blocks follow an event clock with a median gap of 67.6ms, while the analysis samples the data on a 100ms grid.
The wallet score is straightforward. Zhai takes the notional-weighted ten-second signed midpoint markout of each wallet's aggressive orders from July 1 to 10, retaining wallets with at least 100 qualifying orders and excluding TWAP and liquidations. The filter leaves 2,314 wallets. Its top decile contains 231 addresses and becomes the "toxic" cohort, a descriptive term for high post-trade markout and nothing more.
That cohort is then frozen. Two forecasters compete on BTC, ETH and SOL. The anonymous block has ten variables: depth imbalance, volatility-scaled near-touch imbalance, quote-update OFI at 1s and 30s, signed prints, signed taker notional at 1s and 30s, one-minute return, one-minute realized vol, spread. The identity block contributes eleven features limited to the frozen decile. Models are fitted from July 11 to 20, followed by one evaluation pass from July 21 to 27.
The money would come from avoiding adverse selection at the touch. Zhai says this explicitly.
Persistence appears in the tail
Wallet scores have a Spearman rank correlation of 0.52 across adjacent ten-day windows. Rescoring odd and even days separately produces 0.52 and 0.48, placing the Spearman-Brown reliability ceiling at 0.68 and 0.65. At this sample length, observed persistence reaches about four fifths of the level available from a perfectly stable trait.
The distribution matters more than the average. Once market, time, order size, volatility and spread are controlled for, validation markouts stay flat through the fifteenth ventile before jumping to 3.11bps in the top ventile. Across the frozen deciles, D10 moves from 2.56bps to 2.20bps out of sample. D1 moves from -1.13bps to +0.27bps. Losers regress toward zero while winners persist. D10 represents 31.0% of discretionary taker notional during scoring and 25.1% during validation, with 91.3% of the 231 wallets trading again.
The markout does not reverse. For the top decile, it rises from 1.25bps at half a second to 2.11bps at ten seconds and remains elevated at five minutes. D1 through D8 stays near zero and follows native TWAP child orders. Pure transient price pressure fits that horizon profile poorly. Repeated same-direction flow and latent metaorders remain possible, as he notes.
Does identity merely reveal order flow faster?
Section 4.3 addresses that question, with a qualified no. Under ridge, one-second R2 increases from 10.88% to 12.31%, a 13.2% relative gain (t=9.2). Gradient-boosted trees using only the anonymous block reach 19.48%, almost twice the ridge result. Nonlinearity in public data therefore recovers most of the naive value assigned to identity. Adding identity takes the trees to 20.65%, a +6.0% gain (t=5.0). At thirty seconds, the tree gain falls to +2.7% with t=1.6.
Three tests carry much of the argument. The first is a placebo built from 200 cohorts of non-toxic wallets, matched on deciles of scoring-window notional crossed with order count. All eleven features are reconstructed for each cohort, producing 8.8bn cohort-observations. At one second, the real increment reaches 1.435pp, compared with 0.877pp for the best draw, 61% of it. The toxic cohort exceeds every draw through ten seconds. At 30 seconds, its 0.107pp increment lies within a placebo range that reaches 0.136pp (p=0.16).
The second test is an embargo. Delaying every feature by 200 and 300ms reduces the one-second ridge gain to 10.2% and 9.1% (t=8.2 both). The anonymous benchmark suffers much more: 10.88% drops to 6.40% under ridge, while 19.48% falls to 9.64% under trees. Wallet history tolerates staleness far better than book state. For practical use, this is the paper's most interesting result.
Arrival tests provide the third check. When evaluation occurs at 3.09m realized maker fills instead of every grid stamp, the one-second increment is 2.47pp (t=10.8). Giving both models the arriving order's realized side removes 39% of the increment and leaves 1.50pp (t=8.1). Another restriction keeps fills followed by no trade from a different parent order. These account for 26.4% of evaluation fills at one second, where the remaining increment is 1.04pp (t=5.5, 816,944 observations). At ten seconds, the same filter leaves 0.04pp (t=0.2) across 20,000 fills. Both exercises condition on realized future information. Zhai describes one as an ex post decomposition and the other as a mechanism diagnostic. They reduce the scope of the order-flow explanation without eliminating it.
The oracle comparison also matters. Top-decile orders arrive on the side of the oracle-midpoint gap only 41.3% of the time, versus 57.9% for D1, and they trade into a mean signed gap of -1.45bp. Even after restricting gaps to at most 1bp, D10 retains 2.12bps of markout against 0.16bps for D1. Cross-venue latency cannot explain the slow channel. Since the oracle updates only every ten seconds, a subsecond channel remains available.
R2 reaches its economic limit
The paper converts the forecast into a payoff, and presents the exercise honestly as a conversion. Always-on half-second midpoint payoff per unit of gross exposure is 0.131bps for the anonymous model and 0.140bps with identity, a +7.0% increase (t=7.8). That compares with the 7.6% mechanically implied by the R2 gain. Gated cells produce 0.25 to 0.76bps, below the venue's 4.5bp median taker fee. Under both threshold rules, the 1% budget cells return -0.2%. The rolling ten-day thresholds have t=-0.1, while the fitting-window thresholds have t=-0.2.
Implementable profit would require a dynamic execution and market-making model incorporating latency, queue priority, fill probabilities, fees and inventory. Zhai places that work beyond scope.
His defence of the fee comparison is persuasive. A taker strategy is the wrong reference for a maker who already quotes and can use identity to avoid adverse selection. I accept that framing. Zhai does not test queue position. We have previously covered a headline that reduces algebraically to a quantity already reported (our note on Sepp and Lucic). The payoff appendix here belongs to the same species, which the author acknowledges.
The statistics deserve two restraints. The main design has ten scoring days, ten fitting days, seven evaluation days and three markets. A large specification grid, covering frozen versus rolling, decile versus quintile, raw versus peer-adjusted, ridge versus trees and two embargoes, is evaluated on those same seven days.
The December 2025 replication answers that concern well. Independently collected by another party, it contains 26.25bn messages, and the design is applied without retuning. Persistence registers rho=0.47. Ridge moves from 12.36% to 14.16% (+14.6%, t=3.9), and trees move from 21.16% to 25.15% (+18.8%, t=5.1). The toxic cohort beats all 200 matched draws at every reported horizon. In that exercise, however, the ridge models are independently tuned and refitted on the 100ms grid.
One finding pushes against the proposed mechanism. Peer-adjusting the markout by removing other wallets' same-coin same-minute performance increases persistence to 0.62, yet halves the forecasting gain to +6.8% (t=5.1). The more stable score predicts less. Some of the raw signal therefore identifies wallets that appear before the market moves, rather than wallets whose trades are right. With that score, the toxic decile pays a 2.97bp median taker fee during the scoring window, compared with 4.50bps for D1 through D9.
We could not test any of this ourselves. The signal depends on venue-specific wallet identity attached to every order, cancellation, rejection and fill in a sub-second book. Our closest substitute is spot crypto minute bars, which contain no addresses and no book state. Substitution would preserve nothing.
A live maker experiment on Hyperliquid would change my view: keep the same frozen decile, widen or pull quotes as toxic depth share rises, then measure realized spread capture after fees and queue loss. Until that experiment exists, this remains a well-defended 1.4 point R2 increment with an unusually good staleness profile.