In the directional test, a memory-free Poisson dealer gives up 12.617 objective units out of 46.20 against the exact solution, a relative regret of 27.312%. The state-feedback rule Barzykin proposes gives up 0.079, or 0.170%. Both figures come from 2x10^4 simulated Hawkes paths, a one-day horizon and common random numbers. Objective units are money in the paper's convention: inventory in millions of notional, price increments in basis points. No transaction, hedging or externalisation cost is modelled, and no realised fill enters anywhere, so 0.170% measures approximation accuracy against a known optimum.
Worth being clear about that before anything else, because the figure is easy to read as a performance claim.
What the dealer learns from an RFQ it loses
The dealer here earns the quoted offset on every request it converts into a fill, and pays a quadratic penalty on running and terminal inventory. Nothing else enters the objective. No hedge cost, no externalisation, no funding.
Classical dealer models treat client arrivals as a point process whose intensity falls as you quote wider. Barzykin splits that single event in two. A request arrives exogenously, with a Hawkes intensity driven by the history of past requests. The dealer's quote then thins that request into a fill with probability f(delta). The memory is driven by requests alone; in the baseline model fills do not feed it, and Barzykin notes the convention can be relaxed to handle public trade prints, last-look rejection feedback, client relationship scores or hedging flow. The consequence is the paper's central idea: an RFQ you lose still changes your forecast of future client demand, that forecast changes the continuation value, and the continuation value changes the shadow price of inventory that sets the skew.
The exact problem has state (t, q, history). Exponential kernels lift to a finite Markov state, but a d-factor mixture across two sides and K sizes needs 2Kd memory variables on top of inventory, and a power-law kernel needs either a large exponential mixture or an infinite-dimensional Volterra state. So Barzykin builds three approximations, the Volterra-Riccati hierarchy of the title. The first freezes the conditional mean forecast curve and feeds it into the quadratic Riccati system of Bergault, Evangelista, Guéant and Vieira. The second adds a functional-delta covariance correction for uncertainty in the future intensity path. The third evaluates the shadow price under the post-request forecast curve, the mean plus the resolvent response to the request just observed. In every case the final quote comes from the exact Hamiltonian optimizer, so the expansion approximates the continuation value rather than imposing a linear quote rule. The hierarchy is validated against an exactly solvable lifted HJB for a scalar exponential kernel, then applied to an 8-factor exponential mixture engineered to mimic a power law.
The empirical motivation is anonymised HSBC spot-FX RFQ arrivals in EURUSD, GBPUSD and USDJPY, three years, more than 5x10^5 qualified events per pair after filtering to RFQs of at least 0.25 million notional and clients with historical fill ratio of at least 1%. Fitted in seasonality-adjusted activity time with a three-exponential mixture at half-lives of 1, 10 and 60 minutes, the branching ratios come out at 0.862 (USDJPY), 0.895 (GBPUSD) and 0.880 (EURUSD). EURUSD puts weight 0.692 on the one-minute factor and still 0.099 on the hour. Requests are predominantly two-way, so what is fitted is unsigned activity. The FX data motivate the kernel; every performance number in the paper comes from simulated paths.
What he reports is 0.170% relative regret for state feedback against 27.312% for the memory-free Poisson dealer in the directional regime, and up to about 20% lower P&L standard deviation in the long-memory experiment.
Did Jusselin already do this?
Jusselin (2021) characterised general Hawkes-driven market-order flow in an electronic market, proving existence and uniqueness of a viscosity solution to the associated HJB equation and developing a consistent numerical approximation. Barzykin says plainly that his approximation does not replace that characterisation and targets a different object: a low-dimensional, interpretable quote rule. So the thing to judge is whether the cheap rule is close enough, and which layer of the hierarchy earns its complexity.
Barzykin's own reading is that the second layer matters where persistence is high: covariance correction, he writes, "becomes more valuable when the Hawkes process is more persistent and intensity uncertainty is larger". His table says the layer stays small either way. In the benign regime the covariance correction moves relative regret from 0.182% to 0.151%. Near criticality it roughly halves it, 0.884% to 0.513%, which is the effect he describes. Simply feeding back the realized memory state gets 0.080% and 0.083% in those same two regimes. In the directional regime mean forecasting alone leaves 2.498% regret, the covariance correction takes it to 2.159%, and state feedback takes it to 0.170%, roughly a 15x reduction. Conditioning on what actually happened dominates the second moment. A desk that builds the covariance layer and skips the state update has bought the wrong half.
The spot-FX fit is unsigned two-way activity
The 27.312% number lives in a specific regime: eta = 0.92, baseline intensities mu_b = 5 and mu_a = 20, and only the bid side self-exciting. One-sided and near-critical. The empirical section fits something else. Barzykin states that, because requests are predominantly two-way, the estimated process is an activity process rather than a signed buy/sell process, and the 0.86 to 0.90 branching ratios describe that unsigned activity.
He concedes the point himself in the conclusion: the empirical motivation uses two-way RFQ activity, "whereas the most economically valuable state variable for market making is directional flow pressure". He calls directionality assessment "a statistical filtering problem" rather than something a dealer can take for granted, and describes the side-marked model as "an idealized control model rather than a claim that signed request flow is directly observed". His answer in the same passage is that dealers can infer likely direction from previous client activity, historical conversion patterns, skewed response behaviour and contemporaneous public market activity. Plausible. No number is attached to that inference anywhere in the paper. What the paper establishes is that unsigned two-way activity is persistent; what the policy monetises is signed pressure. Until someone reports the signed branching ratio and the hit rate of the side filter on the population the paper actually measured, more than 5x10^5 qualified events per pair over three years in three majors, 27.312% reads as the value of conditioning with a perfect side classifier.
The high branching ratios survive both the filtering of price-discovery traffic and the activity-time change, which disposes of the usual seasonality-artefact objection. Barzykin then disclaims the causal reading himself: an omitted exogenous driver absorbed into the kernel would produce the same fit. For a forecasting model that is fine, and he says so.
Quote impact with a Hawkes tail
The long-memory experiment is the intended application, and there is no exact benchmark for it. Eight exponential factors with rates log-spaced on [2, 2500], weights set by nu_mix = 0.65, branching matrix with spectral radius 0.60, stationary intensity 388 requests per day per side, roughly 780 million notional executed per day. After a one-sided burst the conditional imbalance decays in log-log with slope -(1 - nu_mix) = -0.35, and the analytical result is that the quote skew inherits the same exponent, with the sigmoid and inventory feedback affecting only amplitude. Relative P&L standard-deviation reduction against the Poisson dealer rises with burst size, reaching about 20% at the largest burst plotted, 100 RFQs, against a default burst of 50 elsewhere in the experiment. The figure is read off Figure 7, not tabulated. Barzykin states that mean P&L differences are less stable and are not the main message.
Take him at his word on that. The deliverable is inventory and P&L variance control after a directional burst. Two things bound it. He describes his own construction as an intermediate-range power-law approximation, with a short-time cut-off from the fastest factor and the slowest factor eventually dominating at very long lags. And the Poisson comparator in the power-law experiment is described only as using the same RFQ size distribution, win probability and inventory penalty, calibrated to the same stationary side intensities, with no conditioning on realized memory. In the exponential validation regimes the Poisson dealer is solved exactly, as the one-dimensional HJB on the inventory grid, so it is not numerically handicapped there either. About 20%, then, is the value of any conditioning against a dealer using baseline intensities alone, and it tells you nothing about a production desk already skewing on a flow signal.
Build the side filter first
The prerequisite is a live layer mapping requests, client history and public market variables into a conditional forecast of signed flow, and the paper attaches no accuracy figure to any such filter. Then state-dependent win curves, which this model assumes away: f(delta) does not depend on the Hawkes memory, and Barzykin flags that competition, urgency, toxicity and last-look behaviour may all shift the response function in exactly the stressed one-sided moments where the skew matters most. No Sharpe, no drawdown, no realised fill data anywhere, and empirical validation on realised quoting and fill outcomes is named in the conclusion as future work.
We could not test any of this. The state variable is a filtered history of RFQ arrivals with timestamps, size buckets and side marks, and we hold no OTC request data and no spot-FX quotes or fills. Substituting equity or crypto bars would reproduce none of the request-versus-conversion structure the whole paper turns on.
The part I would act on is narrow: if you already condition quotes on flow, put the realized memory state into the shadow price instead of into a variance correction. In the directional regime that is 0.170% regret against 2.159%. Expect the payoff in inventory RMS, with mean spread roughly unchanged. The part I would wait on is the headline. It rests on a directional Hawkes state the spot-FX section does not estimate, and the author says as much.