The headline result offers little that a desk can trade. Bozzetto, Sifat and Nahidi divide the sample on 15 March. For the later period, they write that "a significant linear relationship between OFIs and mid-price changes emerges, yet its predictive power for one-step-ahead changes is negligible." Their proposed next step is to fit autoregressive models to the imbalances. The paper leaves that work undone. The reported results measure contemporaneous fit.
Fees screen the flow
The mechanism rests on trader selection. With free order submission, an uninformed trader can participate without clearing any cost threshold. The resulting book accumulates flow with little signal. Once a commission applies, expected profit must exceed the fee. Surviving orders should therefore line up more closely with subsequent price movement.
Binance shifted BTC/USDT from zero-fee trading to commission-based trading effective 22 March 2023, with minimal advance notice. The authors put Binance at roughly 50% of global Bitcoin spot volume, making this a large shock.
Their OFI measure is the Cont, Kukanov and Stoikov order flow imbalance. It combines bid and ask size changes at level 1, then sums them within each sampling interval. TFI applies the same construction to executed trades. The authors regress mid-price changes on OFI, TFI, and both together at 1 second, 10 seconds, 1 minute, 10 minutes and 1 hour. OFI-only regressions cover both sides of the split. TFI and combined estimates cover only the later period. A Chow test assesses coefficient stability. The XGBoost benchmark uses the same two predictors, along with minute, hour and day-of-week, on a single 70/30 split with Bayesian tuning.
The sample covers one pair on one venue: BTC/USDT on Binance from 31 July 2022 to 31 July 2023. Book states carry nanosecond stamps and 20 levels of depth, although the analysis uses Level 1 alone. Executed trades begin only on 15 March 2023. In July 2022, the average spread was 0.989 USDT.
How 5% becomes 21.55%
The abstract describes the increase in explanatory power as "from 5% to 21.55% R2." The comparison joins separate rows from the same table. The pre-period hourly OFI regression produces 5.03%, while 21.55% belongs to the post-period ten-minute regression.
Keeping the horizon fixed gives a different comparison. Hourly fit rises from 5.03% to 11.15%. At ten minutes, it moves from 4.04% to 21.55%. The paper reports each matched pair: the Table 2 caption gives the hourly figures, and Section 5.1 gives the ten-minute figures. Only the abstract and introduction make the cross-horizon comparison, with the introduction repeating it as "from 5% to over 21%."
The like-for-like increases remain large. At 1 second, fit rises from 1.02% to 14.14%. At 10 seconds, it moves from 2.10% to 26.19%. At 1 minute, the change is from 3.02% to 30.69%. After fees, OFI fit varies unevenly across horizons. It reaches its high at 1 minute, then drops to 11.15% by the hour.
The missing discussion
Fit climbed while price impact per unit of imbalance fell sharply. The hourly OFI slope drops from 49.8971 to 8.7653, a factor of 5.7. At ten minutes, it falls from 48.4102 to 13.5135.
The paper prints both slopes without discussing their direction.
For a univariate regression, R-squared equals the squared slope multiplied by the ratio of regressor variance to dependent variance. Applying the reported values implies that the variance ratio of OFI to mid-price changes increases by roughly seventy times. The same result appears at the hourly horizon and at ten minutes. A much smaller slope alongside far better fit describes a series whose scale, relative to price changes, moved by about an order of magnitude. Queue sizes, event counts or the size feed's units could all generate that arithmetic. The paper cannot distinguish measurement scale from a change in flow composition. The screening mechanism alone also gives no reason for a fivefold decline in impact.
Trade flow dominates after fees
TFI outperforms OFI at every post-fee frequency. At one second, the comparison is 23.45% versus 14.14%. At ten minutes, it is 48.55% versus 21.55%. Hourly, the figures are 51.78% versus 11.15%. The combined regression reaches 56.30% at ten minutes, and both slopes are significant. Aggressive trade flow carries the contemporaneous relationship. TFI fit increases with horizon. OFI instead peaks at 1 minute with 30.69%, before declining to 11.15% by the hour.
Hypothesis 2 says TFI remains above OFI "under both fee structures." The conclusion says this superiority persists "even under zero fees." Yet the trade dataset begins on 15 March 2023, and the TFI table is titled "After March 15, 2023." The paper contains no pre-fee TFI regression. Its claim about that period comes from prior literature rather than this sample.
Did fees cause the break?
The chosen split is 15 March, a week before the fee change. The final week of zero-fee trading therefore enters the post bucket. The data section places the transition on March 22nd, 2023. Every regression and the Chow test instead use March 15th. Executed trade data also starts on 15 March, and our reading is that data availability determined the split.
The Chow F is 1977.61 with k = 2 parameters. With two parameters and a year of high-frequency observations, rejection is close to automatic. The pre-period also includes the FTX collapse of November 2022, an episode marked in the paper's own spread chart.
The authors interpret the test as confirmation that the fee structure fundamentally changed trading behavior and as support for estimating the two periods separately. Their design covers one pair, one venue, no control pair and no control exchange. It remains a before/after comparison. A difference in differences requires a control series. The Chow test identifies a break. Assigning that break to fees rather than the volatility regime requires a comparable series with an unchanged fee schedule.
Nothing tradeable yet
On one split, XGBoost lowers out-of-sample mean absolute error (MAE) from 595.69 to 532.34, about 11%. The authors describe the untuned result as "indicative of overfitting": train MAE is 362.15 against test MAE of 574.29. Bayesian tuning reduces the difference, producing 511.35 in train and 532.34 in test, so a gap of 212 points becomes 21.
TFI receives the highest feature importance at 44.61%, followed by OFI at 35.01%, minute at 12.73% and hour at 7.54%. The prose supplies a different set of importances: 0.45, 0.25, 0.20 and 0.10, with hour ranked above minute. Those values conflict with the figures. Within a plus or minus 5000 USDT band, predictions and realized changes have 61.45% correlation.
No directional accuracy, no turnover, no fee-adjusted PnL, no backtest. The authors acknowledge the unresolved trading question: practitioners must determine whether greater informativeness compensates for higher transaction costs. A one-step-ahead model is needed to answer it, and they concede that they do not have one.
We could not test any of these results. OFI, as defined in the paper, requires event-level bid and ask prices plus queue sizes. TFI requires signed trade prints. Our data consist of 1-minute OHLCV crypto bars, without quotes, prints or aggressor side. A minute bar cannot reconstruct a cancellation.
Table 2 rerun on a pair or venue whose fee schedule did not change from 31 July 2022 to 31 July 2023 would change my mind. If that unchanged series showed no increase in OFI fit while BTC/USDT did, the screening explanation would hold up. The coefficient collapse would then become the paper's interesting puzzle instead of grounds for doubting the measurement.