A frozen 9B language model, shown nothing but an anonymized intraday price ladder, picks the wrong side of the next ten minutes. It does so in all 14 standard configurations Zhang, Jiang, Huang, Li and Chen test, at -45.7 bps per stock-day on the principal panel. The measurement looks real to me.
The paper concedes the harder question before I get to it. The abstract sells the result as "a behavioral signal for studying how another participant could respond to a predictable policy". The conclusion defers actionability to future work "under transaction costs, market impact, and strategic feedback". So the argument below is aimed at readers who will convert -45.7 bps into edge. The authors' own defence deserves an answer in the same breath. Shuffling the saved action sequences attenuates the effect. The abstract reads that as evidence "that alignment between actions and subsequent returns drives the negative score". Section 3.2.2 adds the overlap tests. There the inverted schedule keeps 43.9 of 46.8 bps, 40.4 of 45.7 and 38.9 of 45.7 after linear projection on a GRU price model, a 38-feature GBDT and one-period reversal.
One disclosure before the numbers. We have no Chinese equity data, so any version of this we run uses liquid US listed equities with one-minute bars. The CSI-500 panel, the prompts, the locally cached Qwen3.5-9B multimodal processor and the inference configuration are all unavailable to us. A substitute frozen text model on serialized resampled price paths would not be a replication of their result. The chart-only branch and the VICReg-pretrained GRU with its rank-64 adapter are omitted entirely from our version. We have no figures of our own to print beside theirs yet.
Start with what timing alpha measures
The setup is deliberately impoverished. An anonymized CSI-500 stock-day is laid out on a 239-minute grid. The primary panel scores 23 ten-minute decisions. At each one the model sees a rolling price history, as numbers, as a chart, or both. It also sees whatever endogenous state the condition permits: current position and entry, a trade ledger, an account summary, or memories it wrote to itself on earlier bars. Stock identity, calendar date, absolute price level, news and cross-sectional ranks are concealed. It outputs 1 (long) or 0 (flat). Then the next interval's return is revealed. Temperature 0.7, top-p 0.95, thinking disabled.
Scoring is where the paper earns its keep. Timing alpha measures whether the agent was holding during the better intervals of a stock's own path. Formally it is the exposure-centered within-stock covariance: sum over intervals of (position minus that day's average position) times the subsequent return, aggregated per stock-day, reported in basis points. Center the position and the day's drift falls out. A policy that is always long scores exactly zero. So does one that is always flat.
The paper's decomposition makes the long-only return on a stock-day equal passive exposure to the path plus the centered timing term. Their archived example is the cleanest illustration I have seen of the distinction. A memory-conditioned trace with 16 of 23 intervals long earned +359.4 bps while long, which is a profit. It was flat for the first seven intervals, and those seven carried +1,404.0 bps. Timing alpha: -867.3 bps.
Positive P&L, badly wrong timing.
Across 14 standard configurations, every timing estimate is negative. The principal 10-minute Qwen3.5-9B rows are -45.7 bps per stock-day for numerical text (13,710 stock-days, CI [-49.4, -41.8]), -29.9 for chart-only (14,937), and -48.9 for multimodal (14,438). The 20-minute state rows range from -43.9 to -56.7 bps, and the 5-minute rows from -30.5 to -53.2. Claude Haiku 4.5 keeps the sign at -38.8 bps on 827 stock-days at 5 minutes and -42.9 on 165 stock-days at 10 minutes. That is thin power rather than a small estimate. Granite-4.0-H-Small keeps the sign at smaller magnitudes: -36.3, -25.6, -14.1 and -8.2 bps across price text, position, ledger and account prompts, 600 stock-days each, archived point estimates with no confidence intervals. These panels use different sampling frames, and the authors say so. The repeated sign is the claim. Magnitude comparisons use matched panels.
The shuffle control does the identification work
Permute the saved action sequence while preserving each day's long fraction and the effect collapses. Global shuffle: -3.5 bps. Same-day shuffle: -8.7 bps. Against -45.7 for the intact schedule, on 13,716 stock-days. Whatever produces the number lives in the ordering of the actions relative to the returns.
The signal-overlap tests do real work too. Take the inverted schedule, regress it within-stock on a GRU price model, on a 38-feature GBDT, and on one-period reversal, then score what is left. Residuals of +43.9 of +46.8 bps on the GRU's 2,519-stock-day evaluation frame. Then +40.4 of +45.7 for GBDT-38 and +38.9 of +45.7 for one-period reversal, both on 13,717 stock-days. The paper calls comparisons across those frames descriptive. The pattern is linearly distinct from three conventional signals, with correlations of +0.114, +0.132 and +0.286. Carry the qualifier the authors carry: these are linear within-stock projections onto three specific comparators. Nonlinear or unmodelled predictability is untouched.
The self-conditioning result is the one I would want replicated first. Restrict to stock-days that contain both actions. Show the agent one earlier memory it wrote itself, and timing moves from -62.8 bps (1,194 scored stock-days) to -74.1 (1,054). Position changes fall from 2.5 to 1.7 per stock-day. In the later-period panel the same pair is -68.3 versus -85.6. Consistency buys persistence, and the persistence runs in the wrong direction. The consistency directive was chosen on a broader panel that includes those later dates, which the authors state plainly. Read those rows as a chronological extension.
Where would the money come from?
Invert the schedule: hold the stock when the agent would be flat, stand aside when it would be long. Timing alpha flips sign by construction, because A(1-p, r) = -A(p, r) is an identity in the definition. The paper presents it as a sign diagnostic. Anyone reading -45.7 bps per stock-day as +45.7 bps of harvestable edge has skipped that sentence.
Consider what a live version demands. You need the model's decision before each interval. That means paying inference latency inside a ten-minute bar, then crossing a spread to change position. Between 1.7 and 6.71 position changes per stock-day. One to three round trips. The measured quantity is -45.7 bps per stock-day gross, on research-return labels with no costs, no borrow, no latency, no impact and no price feedback. The authors state all of that as scope. Two things then have to hold at once. The mechanism has to transfer from a Chinese mid-cap index to a universe where you can actually execute. And it has to clear a cost bill roughly proportional to how often the agent switches.
Capacity is the second problem, and the paper's design creates it. Timing alpha is a within-stock quantity: it says when to be in one name. Cross-sectional IC is near zero in nine of ten narrative conditions. The exception is multimodal no-directive at -0.0137, with CI [-0.0261, -0.0021]. Under memory conditioning the correlation between timing and portfolio return is -0.066, against +0.355 for the conventional conditions. The strongly negative quantity fails to translate into a portfolio-return advantage through selection. You would be running a per-name timing overlay. Its size is set by the liquidity in each name.
One structural point about the estimand. Timing is scored only on stock-days containing both actions, because all-long and all-flat days score zero mechanically. In the narrative sweep 211 to 489 constant-action paths per condition are dropped, alongside 233 to 1,690 failed intervals from 34,500 attempts. The reported means describe the switching subset. That is a narrower denominator than the capital a live strategy ties up on the quiet days.
Two inputs that see the future
The fine-resolution charts and the learned encoder features use a one-minute series reconstructed from nested forward-return labels. The authors flag this: the reconstruction depends on later label rows. They then lean on the numerical-text rows for the clean information boundary. That is the correct response, and it is why the -45.7 bps text figure is the number worth arguing about. The latent soft-token arms inherit the same input, so read those ablations as internal comparisons. Within that caveat, one row in the latent-state ablation table is interesting on its own terms. Removing the 92-word lexical constraint and letting the soft tokens be free vectors takes timing to -2.1 bps, with a CI spanning zero. The constrained rows that share the centered outcome run -29.2 to -33.8. The separate-frame trainable and frozen encoders score -38.4 and -39.7.
Other fine print worth carrying. Granite's four rows are archived point estimates recovered after a cache migration, with no confidence intervals. API failures there retained the incoming position, with each failure counted (33, 76, 8 and 0 stock-days of 600). Haiku's 10-minute row is 165 stock-days. The coarse-image row is labeled Qwen3-VL in the archive while the launcher defaults to Qwen3.5-9B, which the authors list as an open verification item. The 20-minute state prompts display 10-minute wording. None of these break the sign result. All of them mean the cross-model and cross-horizon rows are sign checks, as the paper describes them.
The one number that would settle it
The measurement is careful, and the shuffle control carries the identification. My scepticism concerns a single step. Timing alpha is a costless covariance, centered on each day's exposure and conditioned on the switching subset, computed on CSI-500 research labels. Turning that into a position you can hold is what I do not believe yet. One number would change my mind: a net-of-cost, executable-price result for the inverted schedule, on a period the consistency directive was never chosen on. Until then this is a good description of how a frozen language model behaves when handed a serialized price history of previously completed intervals and nothing else. Worth knowing for its own sake.