QUESTrader's clean sweep comes from a single 15-month test window in which all four indices rose. Every figure below belongs to the paper. We ran none of these experiments. The paper trades constituents of DJI, FTSE, Sensex and TAIEX, while anything we built from this mechanism would trade US-listed equities. The published results therefore would not describe our implementation. This is an assessment of someone else's evidence.
On DJI, Orra, Choudhary and Thakur report Sharpe 1.459 for their meta-learned auxiliary tasks, up from 1.076 for plain PPO. Annual return rises to 21.785% from 15.674%. The gain suggests that the auxiliary questions improve how a PPO trading agent learns its state representation. Evidence of a tradable edge would require more than Jan 2024 to Mar 2025, roughly 310 trading days per market.
Inside the two networks
The daily multi-stock MDP is familiar. With n names, the state has 10n+1 dimensions: cash, shares held, the close and eight technical indicators for each stock (SMA30, SMA60, MACD, the two Bollinger bands, RSI, CCI, ADX). Each action is an integer share count per name in {-m,...,m}. Reward equals the change in account value less the transaction cost on the position change.
The contribution sits above that setup. Auxiliary tasks in RL trading are usually hand-picked, such as predicting next-period volatility or the next close change. QUESTrader instead uses General Value Functions whose definitions are learned. A second network, called the question network, reads a short forward slice of states. It emits a cumulant vector and a discount vector, which jointly specify d_q prediction targets.
The main network answers those questions from the current state alongside its policy and value heads. Its auxiliary squared TD error receives weight lambda_aux. Parameters in the question network serve as meta-parameters: K inner PPO updates are followed by one outer update, which differentiates the sum of PPO losses through the unrolled inner steps. The design extends Veeriah et al.'s question/answer split to trading. Its economic premise is simple. A sharper representation could help the policy hold through noise rather than flip positions.
Yahoo daily closes cover 1 January 2010 through 31 March 2025. The study uses four universes of thirty names: all 30 constituents of DJI and Sensex, plus the top 30 of FTSE and TAIEX. Training and validation end on 31 December 2023, with trading beginning on 1 January 2024. Each run starts with one million of capital, charges a flat 0.1% fee on buys and sells, and fills every order at that day's close.
The return gap over plain PPO
DJI gives QUESTrader a 21.785% annual return and Sharpe 1.459, versus 15.674% and 1.076 for plain PPO. On FTSE, the comparison is 19.164% and 1.124 against PPO's 13.018% and 0.630. Sensex records 16.727% and 1.003 against 11.948% and 0.768. TAIEX reaches 30.279% and 1.803 against 21.513% and 1.001.
Across the markets, QUESTrader leads plain PPO by 4.8 to 8.8 points of annual return and 0.23 to 0.80 of Sharpe. The proposed model and the benchmark deep RL models had their hyperparameters tuned separately with Bayesian optimization, which makes the comparison broader than a clean one-factor ablation. Buy-and-hold returned 12.857% on DJI and 15.937% on TAIEX over the same window. The agent's outperformance therefore exceeds a passive long position.
The nearest baselines close much of the apparent PPO gap, including a method outside deep nets. On TAIEX, mean-variance optimisation reaches Sharpe 1.738 against QUESTrader's 1.803. Its maximum drawdown is 8.065%, compared with 13.632% for QUESTrader. MVO sacrifices return, as the paper acknowledges, producing 18.007% against 30.279%. Even so, a Markowitz portfolio matching a meta-learned PPO on risk-adjusted terms while carrying 40% less drawdown demands an explanation.
QUESTrader misses the lowest drawdown on DJI, where it records 10.369% against Adaptive's 7.138%, and on FTSE, where 12.165% trails MVO's 10.062%. Sensex is stronger. QUESTrader has the lowest drawdown among the deep RL methods at 10.584%, narrowly ahead of SRRS at 10.650% and PPO at 10.653%.
SRRS uses an approximate Sharpe ratio to shape its reward. In the FTSE table, it has the worst drawdown at 30.851% and the lowest Sharpe among the modelled strategies at 0.436. Only random trading is lower, at -0.117. The risk-targeted reward degraded risk control. DeepScalper supplies the strongest return baseline in three of four markets, although the authors describe it as a framework for intraday trading and run it here on daily closes.
What does five-seed dispersion tell a trader?
Every deep RL row reports a mean and standard deviation across five independent runs. No significance test appears in either the tables or the text. The DJI result of 21.785% plus or minus 1.42 measures disagreement across seeds. It cannot show how far the 15-month outcome would move under a shifted window. The study has one test period and no rolling refit, leaving seed dispersion untested against the sampling noise in roughly 310 daily returns.
The ablation strengthens that concern. The authors say they retrain the full model for each configuration and evaluate every version on the fixed test window. Their recommended region comes directly from those curves: d_q in [16,64], with K around 10, for maximising Sharpe. The Sharpe curve peaks near 1.4 at d_q=16, falls to about 1.2 at 32, recovers to 1.3-1.35 at 64 and drops again to 1.15 at 128. Total return declines to 24% at d_q=32 and reaches its maximum of 27-28% at 64. Both measures form a double hump on the evaluation period. Sharpe peaks at K=10, while return peaks at K=20. The paper resolves the difference by recommending d_q around 16 to 64 and K around 10 to 20.
The recommendation is in-sample to the window that certifies the method.
The authors keep the wording restrained. They present d_q 16 to 64 with K around 10 as "a safe region if the objective is to maximize Sharpe", offering practical guidance rather than claiming generalisation. That qualification goes some way. Yet the guidance and the headline Sharpe of 1.459 come from the same window, so they are dependent. Readers cannot separate the contribution of the mechanism from the selection of d_q and K. Neither the abstract nor the conclusion identifies the single 15-month window as a limitation. Future work instead covers off-policy GVFs and implicit-gradient meta-updates, online regime detection, and a portfolio extension.
Declared frictions, missing turnover
The paper clearly states three assumptions: zero slippage, negligible market impact and immediate settlement. The candour deserves credit. It also sits awkwardly beside the authors' criticism of supervised models. They argue that supervised approaches fail because they "do not consider the most practical constraints, such as transaction costs, slippage, and liquidity constraints" and consequently "perform well in backtests but often fail in live markets". Their own environment includes only the 0.1% fee, which is never varied.
Reduced turnover is one of the three conclusions the authors draw from the action timeline, and it alone has a direct cost consequence. That timeline compares PPO with QUESTrader for one stock over roughly January to April 2025. QUESTrader's share counts stay mostly below 40 and contain long runs of zeros, while PPO's exceed 60. Smoother markers for one name across one quarter cannot substitute for a cost audit. The flat 0.1% fee remains the sole friction, and the paper reports no turnover figure.
One result would change my assessment. Freeze the architecture at the authors' recommended d_q and K, then run it on non-overlapping windows that include 2022, a year used for training but never for trading. Report turnover with the Sharpe. If the lead over plain PPO survives, the representation story is doing the work.