A 7.7% RMSE improvement across 54 weekly observations cannot establish that Shanghai crude oil futures have grown up. Wang and Yu make no such claim. Their abstract concludes with "improved but still incomplete pricing autonomy," a measured reading supported by careful empirical work. The attached three-stage maturation story comes from a model re-estimated on the full sample. It never faced the holdout.

The model has two parts that lead in different directions.

The model and the test

Brent and WTI provide the reference price for physical crude worldwide, with regional grades trading at differentials. Shanghai's INE contract launched on March 26 2018, denominated in RMB and open to foreign participants. It was intended to give Chinese refiners and importers a hedging instrument tied to Asian supply and demand rather than the North Sea. The empirical question follows directly: how much of INE's return variance comes from a Brent innovation, and has that proportion declined?

Wang and Yu estimate a bivariate weekly system using Brent and INE log returns. Wind supplies 362 observations from March 30 2018 to April 30 2025, aligned on the last common trading day of each week. Their baseline is a recursive SVAR(2), with Brent ordered first. AIC and FPE select the lag length, while SC and HQ both favor p = 1, as shown in the authors' lag table.

A compact neural component sits above the baseline. A single-head linear attention layer processes the two lagged return vectors through a 16-dimensional embedding. Alongside it, a ReLU feed-forward branch contains 32 hidden units, with its output added to the attention forecast. The combined model has roughly 1,157 trainable parameters. The VAR has 10.

Lu and Yang provide the economic motivation. Their work, cited in the paper, shows that structured linear attention can be rewritten as a time-varying VAR. In other words, the attention block reduces to a lag coefficient matrix that varies with the state. The authors therefore decline to interpret raw attention scores as structural coefficients. Wang and Yu instead apply the structural restrictions to the combined forecast residuals. The neural component estimates the conditional mean, while identification remains econometric.

The paper also introduces rho_t, its nonlinear contribution ratio. This diagnostic measures how much of the forecast's squared magnitude comes from the feed-forward branch rather than the attention branch. Zero denotes a purely linear forecast, with the ratio increasing as the nonlinear contribution grows.

Testing follows a strictly chronological split. The final 15% of the 360 lagged pairs, equal to 54 observations, remains untouched as the holdout. Expanding-window validation uses the preceding 85%, and each training window supplies its own standardization moments. Results cover Ten seeds. The authors also run an ablation without the feed-forward branch and examine four identification schemes.

We could not run this ourselves. The mechanism depends specifically on the INE-Brent pair, and we have no INE series. Replacing it with CL futures or a crude ETF would examine price discovery in another market.

A narrow forecasting win

On the test block, the full hybrid records INE RMSE of 0.0398, with a seed standard deviation of 0.0013. VAR(2) produces 0.0431, the source of the reported 7.7% improvement. Brent improves 4.5%, moving from 0.0396 to 0.0378. Against VAR(2), the Diebold-Mariano result is p = 0.031 for INE. The linear-attention-only model reaches p = 0.118, and the MLP reaches p = 0.410.

The table's failed model says the most. Using the same lag inputs, the MLP has 226 parameters and posts INE RMSE of 0.0446, worse than the 10-parameter VAR. Unrestricted nonlinear capacity damages performance on this sample. The linear-attention branch alone contains 931 parameters and reaches 0.0417. The gain therefore comes from the attention layer's imposed autoregressive structure, while additional flexibility by itself makes the forecast worse. Wang and Yu acknowledge this directly. It is their strongest methodological result.

Seed dispersion complicates the incremental hybrid gain. INE RMSE improves from 0.0417 for attention alone to 0.0398 for the full model, about 4.6%. Their seed standard deviations are 0.0010 and 0.0013, so the distributions overlap. Wang and Yu write: "The standard deviations across ten random seeds are small relative to the differences between the full hybrid model and the benchmarks." The statement holds against VAR(2), where the gap is 0.0033, roughly two and a half seed standard deviations. It does not describe the 0.0019 difference between the hybrid and its own attention branch, which falls within a single seed standard deviation of either model.

Table 3 contains only a Diebold-Mariano column comparing each model with VAR(2), and 0.031 is one of the three p-values reported there. The paper gives no DM test between the hybrid and the attention-only branch. Yet that missing comparison is the basis for the 4.6% claim.

Against VAR(2), the move is 0.0431 to 0.0398, or 7.7%, with DM p = 0.031 across 54 weeks. Against the attention branch, it is 0.0417 to 0.0398, or 4.6%, without a test.

Nonlinearity appears in stress

The crisis finding is cleaner than the average forecasting result. A week receives the crisis label when either market's absolute standardized return exceeds 1.5. Mean rho_t reaches 0.184 during those weeks, compared with 0.051 in normal periods, a factor of 3.61. A block-permutation test rejects equality at p = 0.004.

Removing the feed-forward branch raises INE RMSE by 1.9% in normal weeks, from 0.0306 to 0.0312. In crisis weeks, the increase is 9.1%, from 0.0531 to 0.0584. The mean absolute nonlinear adjustment is 0.0067 during crises and 0.0018 otherwise.

The pattern hangs together. The linear path carries almost all the load in ordinary periods. A different component becomes active in March 2020 and early 2022.

Its implementation limits immediate trading use. Although the threshold is fixed beforehand, assigning the crisis label requires that week's realized return. The crisis/normal classification is therefore a conditioning state rather than a switch available on Monday morning. By contrast, rho_t can be calculated one step ahead. The authors' suggestion that it serve as a monitoring indicator is the more practical use of the same result.

Does this establish maturity?

Wang and Yu state the limitation plainly: "They do not establish complete autonomy: Brent remains the largest source of INE forecast-error variance under recursive, sign-restricted, long-run, reverse-ordering, and currency-adjusted specifications." The abstract reaches the same judgment with "improved but still incomplete pricing autonomy." Its positive case rests on the model: "the principal incremental contribution of the hybrid model lies in identifying state-dependent nonlinear adjustments during crisis periods." That crisis-state result comes from the full-sample refit, which creates the central problem.

The hybrid model's unconditional full-sample forecast error variance decomposition stays close to the conventional SVAR. At horizon 8, the SVAR assigns 59.9% of INE variance to Brent and the full hybrid assigns 59.6%. The paper is explicit: "the unconditional 60:40 split is not itself a novel consequence of the Transformer." Alternative identifications cluster nearby. Sign restrictions produce 58.4%, and a Blanchard-Quah long-run restriction produces 57.3%. Converting INE to USD gives 58.8%. At p = 1, the share is 61.0%. Even reverse recursive ordering, the most conservative specification, leaves Brent at 53.2%.

Seven years in, following both the pandemic and Ukraine, the international benchmark explains over half of INE's variance under every specification the authors examine.

Transmission in the reverse direction remains marginal. The nonlinear Granger exclusion test for INE to Brent reports chi-square 5.356 on 2 df, with p = 0.069. The opposite direction has p = 0.000. Flipping the recursive ordering also moves the peak Brent response to an INE shock from 0.005 to 0.014. Wang and Yu call this "consistent with contemporaneous feedback from Chinese demand," using reverse ordering as a conservative sensitivity bound. Under either interpretation, the reverse channel's size depends on the timing assumption. Their own summary is direct: "Reverse transmission is substantially smaller."

The three-stage account divides the sample into cultivation 2018-2019, extreme shocks 2020-2022, and transition 2023-2025. Its support comes from subperiod historical decompositions. Shock ranges contract from about [-0.30, 0.15] in 2020-2021 to about [-0.15, 0.10] in 2023 to April 2025. That describes the data. Lower volatility after a pandemic does not by itself demonstrate a shift in pricing power.

Table 5 is the paper's sole presentation of variance shares, and it has no dated rows. It compares the SVAR, at 59.9% for horizon 8, with the hybrid at 59.6%. It also compares the hybrid's normal state at 58.7% with its crisis state at 60.7%. No FEVD by subperiod is reported.

The estimation sequence matters more. Once model selection and out-of-sample evaluation finish, the chosen specification is re-estimated on the full sample for structural description. Every impulse response, FEVD, and subperiod chart is consequently an in-sample description. The crisis-state decomposition is included. Brent's share there climbs to 65.1% at horizon 1 and 60.7% at horizon 8, compared with 59.8% and 58.7% in the normal state. This is also the paper's main incremental claim, drawn from the full-sample refit rather than the 54-week holdout. Figures 2 through 7 show recursive-SVAR baselines only. Hybrid state-dependent responses appear in the prose as point estimates without intervals.

For the spread book

The impulse responses offer a more useful pattern than the maturity narrative. In the recursive SVAR baseline, INE responds by 0.037 on impact to a one-standard-deviation Brent shock. That equals about 74% of Brent's own 0.05 self-response. INE's response falls to 0.0065 in week 1, reverses to -0.0028 in week 2, and reaches zero by week 3. In the crisis-state hybrid simulation, the week-2 response deepens to roughly -0.005. INE also overshoots after its own shock in the recursive baseline, moving from +0.029 on impact to -0.0077 in week 1 and +0.0031 in week 2.

The implied Brent-INE spread pattern is a one-to-two-week overshoot and reversal, with a deeper move during stress. It can be tested. Figure 2 supplies 95% intervals for the recursive baseline, while the state-dependent hybrid responses remain bare point estimates. The system is weekly and bivariate, and its return magnitudes are small. A trader with INE data and a spread book can treat the pattern as a hypothesis and check whether it survives costs.

Wang and Yu acknowledge most of the constraints. The sample is small. A bivariate model cannot distinguish global supply shocks from global demand or FX shocks. Weekly observations conceal the day/night leadership emphasized in prior literature. The identifying restrictions remain maintained assumptions. The authors give less ground on the burden carried by the maturation stages. Those stages rely entirely on historical decompositions from a model "re-estimated on the complete sample for descriptive historical and structural decomposition." The paper also provides no variance decomposition by subperiod.

For anyone hedging Chinese crude exposure, the operational conclusion remains familiar. Brent is the main risk factor, with pass-through concentrated across two weeks. Correlation strengthens in the states where diversification is most needed. At horizon 1, Brent's crisis-state share is 65.1%, versus 59.8% in the normal state. INE offers a normal-regime diversification benefit that partly fades during stress.

One result would change my view: an INE FEVD share under 45% at horizon 8 from an expanded system containing FX, Chinese import volumes and inventories, estimated out of sample. The evidence here remains far from it.