Bet against the last 15-minute candle across liquid Binance pairs and the crypto class-mean AUC reaches 0.5267. Kitron and Wengrowicz get 0.5305 from the paper's full twelve-lag kernel, while the unthresholded BTC call is right 52.3% of the time. They also price the trade. Gross returns peak near 1.3 bp per trade against a 5 bp round-trip cost, enough to measure and too little to cover benchmark spot capture costs. The paper makes a measurement claim. That measurement has to justify the attention.
One candle carries most of the signal
Each bar has an intra-bar return, close minus open over open, and an up label when that return is positive. The predictor distributes weight across the previous twelve returns. It soft-clips each through tanh(150 r), applies a decaying kernel k^-alpha, and uses one amplitude A for the entire vector. Negative A indicates reversal. Three free parameters remain at any lag count. Maximum likelihood fits them, with selection performed on a chronological validation tail. An unconstrained free AR(12) logit runs beside the kernel throughout, and a pair qualifies as predictable only when both models pass the test. The kernel restriction therefore carries no result by itself.
The proposed mechanism comes from old microstructure. Aggressive takers move price to secure fills, liquidity providers concede, and prices recover while inventory is worked off. Across pooled BTC, ETH, SOL and XRP observations, out-of-sample AUC is 0.540 after bars whose move matched their taker imbalance. It falls to 0.519 after bars where the two disagreed. The pooled difference in next-bar flip rates is +0.021 (CI [0.015, 0.026]). The same conditioning appears across the full cross-section: among the 183 pairs, the Spearman correlation between the flip-rate difference and fitted coupling is -0.36 (p = 4e-7).
Depth consumption adds no conditioning power. Using roughly 30-second Binance order-book snapshots from the USDT-M perpetuals for the same four coins, the next-bar flip-rate difference after depth-consuming and depth-replenished moves is -0.002 (CI [-0.007, +0.004]). Touch-level consumption remains unobserved. The authors interpret the evidence as compensation for absorbing flow, with less support for compensation from rebuilding a book that re-forms inside the bar.
The sample contains the 183 highest 24-hour-quote-volume Binance USDT spot pairs and 187 liquid US stocks and ETFs. It uses 15-minute bars from 2025-01-01 to 2026-02-11. Each walk-forward window has 5,760 training candles and 960 test candles, then advances by the test block. The non-overlapping out-of-sample blocks are concatenated and scored once. Crypto averages 0.531 out-of-sample AUC, and 98% of pairs exceed 0.5. Stocks average 0.499 and divide evenly. With joint FDR control at q=0.05 across all 370 assets, significance survives for 164 of 183 crypto pairs and 5 of 187 stocks.
Why lag-one autocorrelation misses it
BTC's lag-one return autocorrelation is -0.003, implying an AUC of about 0.501 through the linear channel. The measured AUC is 0.533. Rho-1 fails because it weights bar pairs by r_{t-1} r_t, allowing the top decile of |r| to nearly cancel the middle deciles.
Signs show the pattern directly. In crypto, the sign-flip rate increases monotonically across all ten deciles of |r_{t-1}|, from 50.2% to 53.0%, and stays above 50% in every decile. Stocks range from 49.8% to 51.2%. Their sequence is non-monotone and remains at or below 50% through the fourth decile. A parameter-free, kernel-weighted sign correlation with alpha fixed at 1 recovers the class gap without fitting anything: -0.035 (CI [-0.041, -0.029]), negative for 98% of crypto pairs and 50% of stocks.
Almost all the predictive content arrives at one lag. Trading against the previous candle produces a crypto class-mean AUC of 0.5267, versus 0.5305 for the twelve-lag kernel. The paired improvement is +0.0039 (CI [+0.0026, +0.0052]) and is saturated by three lags. Horizon weakens the effect. Focal-panel crypto means are 0.536, 0.521 and 0.496 at 15 minutes, one hour and four hours. Across the population, 58% of crypto pairs that meet the hourly bar-count cut remain FDR-significant at one hour, compared with 0% of stocks.
The authors apply their own discounts
The 90% tally cannot be read as 183 independent confirmations. The paper's own words give the defensible interpretation: "this liquid-crypto cross-section carried a small but pervasive and persistent reversal signal throughout the window," not "183 markets independently confirmed one." Mean pairwise crypto correlation is 0.40, leaving an effective sample of roughly 2.5 independent observations. On the shared time grid, the joint moving-block bootstrap estimates the class-mean gap at +0.031 (95% CI [+0.027, +0.035]).
Its headline interval is fit-conditional. A refit-inclusive joint bootstrap covering 28 assets gives +0.0166 (CI [+0.0091, +0.0241]) and roughly doubles the interval width. The exact permutation null uses 340 day-aligned label shifts. The observed gap lies 4.4 null SDs away, and 0 of 340 shifts reach it. Null dispersion is four times the bootstrap's, 0.0073 versus 0.0018. The authors treat this as evidence that their reported interval is optimistic in width.
Dropping flat bars alone is their preferred artifact-robust estimate, although the decision came after seeing the data. Crypto then records 0.5200 against 0.4985 for stocks, a gap of +0.022 (CI [+0.019, +0.025]). The free logit gives +0.017 [+0.015, +0.019]. Adding the one-bar feature/label gap reduces crypto to 0.5093 and stocks to 0.4980, leaving +0.011 (CI [+0.008, +0.014]). Together, these choices preserve 36% of the headline gap. Independent channels would preserve 47%, so some of the signal manufactured by the flat-bar convention overlaps with the bounce channel. The authors further warn that the artifact conjunction and refit correction are dependent and must not be multiplied. Their instruction deserves to be followed.
Volume rankings were measured after the sample ended, making the as-fixed universe survivor-tilted. The ex-ante specification keeps the 133 pairs already trading during the first sample month and ranks them again using that month's volume. Its result is stronger: +0.035 (CI [+0.030, +0.038]), with 129 of 133 significant. A frozen six-month holdout returns +0.020 (CI [+0.010, +0.028]). Crypto scores 0.522 against 0.501 for stocks, and 60% of scored pairs remain significant. The equity zero has force as a bound: median power is 0.98 for detecting a +0.031 effect, and the one-sided 95% upper bound on the stock class mean is 0.5007.
Fees swallow the edge
Higher confidence thresholds behave as expected from a genuine signal, raising accuracy to 56% on the most confident BTC and ETH bars. Economics still deteriorate. Trading fewer bars removes opportunities faster than the edge improves.
Not one of the 183 pairs clears the 5 bp maker band at any threshold.
Where it trades, the median pair earns 0.46 bp per trade at a 0.02 confidence cut. The median stock earns zero. At 5-minute bars, the figure drops to 0.15 bp. Taker costs run from 10 to 20 bp. The authors explicitly say that gross returns omit depth, latency and adverse selection, each working against a taker, and they cannot determine which cost component binds. Queue position plainly matters when the signal pays 1.3 bp gross at peak selectivity on BTC and ETH, and 0.46 bp for the median pair.
A narrower reading of the wrappers
The wrapper panel invites over-reading, which the authors head off before showing it. IBIT and Binance BTC have a correlation of 0.982 across 49,170 aligned one-minute observations. On live cells, arbitrage makes inheritance close to mechanical. Insert one bar between predictor and label, and no session-matched slot remains significant; the largest wrapper is 0.497. Regressing family wrapper mean on the underlying gives an inheritance slope of 0.69 with SE 0.32, rejecting neither full nor zero inheritance. COIN stays in as a control and records +2.6 SD on the panel's own descriptive scale, above every wrapper family except Bitcoin and Ether.
Read strictly, the authors concede, the panel distinguishes live-underlying NAV-linked wrappers from everything else. It cannot isolate NAV linkage from correlation. The narrower result still matters: an efficiency statistic measured on a wrapper's tape describes the wrapped price process, rather than the listing venue.
The flow evidence reaches a similar limit. Aggressor flow is selected by the decision to trade, so splitting it cannot distinguish compensated provision from an informed-flow account. Its flow-driven label also partly incorporates the contemporaneous return's own sign. Their time-reversal control points in the same direction. Reversing the series changes the crypto mean only from 0.536 to 0.521, while all three focal coins remain significant. Most detectable sign structure is therefore time-symmetric. The overreaction-then-correction account depends on flow conditioning more heavily than on the sign path.
For the four focal coins pooled together, adding twelve lags of taker imbalance to twelve sign lags changes out-of-sample AUC by 0.000 (CI [-0.002, +0.002]). Imbalance alone trails signs by 0.008. Bar-level flow largely repeats information in the price path that it produced. The authors also disclose operating an automated crypto trading system that uses a variant of this model. The 15-minute horizon and up/down label predate the study.
We did not test any of these results on our own data. Our crypto coverage consists of roughly 50 instruments with minute OHLCV, rather than a Binance USDT panel. We have no trade prints, quotes or book, leaving the taker-imbalance and depth analysis unavailable. Fifteen-minute bars would also need to be assembled from minute bars. Running the sign model on a different liquid panel would test only the mechanism. The paper's cross-market measurement would remain untested.
I would use the result for calibration. It locates the fee equilibrium for liquid crypto spot at 15 minutes, roughly a quarter of the cheapest maker round trip at peak selectivity, and disqualifies rho-1 as a screen for sign predictability. The unresolved falsification is the one worth following: reversal after moves driven by forced liquidations, where flow initiation carries no information choice, weaker than after matched ordinary flow-driven moves. A weaker result would break the liquidity-provision interpretation. The 1.3 bp would then look like a residual that nobody has bothered to take.