AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

Xiong's timer tops 50/50, lags 87% static growth

470% annual turnover, six in-sample parameters, and one durable post-2022 drawdown result

2026-08-06 · 7 min read · US ETFs

Reviewing: Continuous Timing Signals for Growth-Defensive Style Allocation: Factor Attribution, Risk Matching, and Out-of-Sample Evidence · Zheli Xiong · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Continuous Macro-Stress Smooth-Score Growth-vs-Defensive ETF Style Allocation

Backtest period 2020-01-01 to 2025-10-08 · hypothetical, net of modelled costs

Why these figures are not the paper's (2)

This is not a replication of the paper (3)

  • Exact VIX index history is not listed as an available data source; implementation should use an explicitly documented substitute such as SPY EOD option implied volatility where available or SPY realized volatility from daily_prices. This tests the same volatility-stress-relief mechanism but is not an exact replication of the paper's VIX-based signal.
  • Exact TNX index data is not listed, but the 10-year Treasury yield can be proxied using the available FRED DGS10 macro_indicators series.
  • Kenneth French daily factor data used for the paper's attribution is not in the provided catalog; the trading policy can be backtested without that attribution, but exact FF5+momentum regression diagnostics would require an external factor dataset.

The figures below measure what we could run, not the paper's own method, so they are not evidence for or against its claim.

Our own audit found this run does not follow the paper faithfully (5)

  • deviation left undescribed by the audit (invalidates: Factor attribution predictions for G-D market beta, SMB beta, HML beta, RMW beta, CMA beta, MOM beta, annualized alpha, Newey-West t-statistic, and adjusted R^2 do not apply to this backtest output.)
  • deviation left undescribed by the audit (invalidates: Fixed parameter validation predictions over 2018-06-28 to 2026-05-15 do not apply to this spec's main strategy results.)
  • deviation left undescribed by the audit (invalidates: Strict incremental Old Best + Bond/Credit main-window, OOS, post-2022, and cost-sensitivity predicted results do not apply.)
  • deviation left undescribed by the audit (invalidates: All paper numeric performance predictions tied to the original 2017-06-28..2026-05-15, 2018-06-28..2026-05-15, and 2022-01-03..2026-05-15 windows are reference diagnostics only and should not be expected to match.)

1 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 111.2%Oct 2025
Sharpe
0.92
Total Return
111.2%
Max Drawdown
-31.6%
CAGR
13.9%
Volatility
21.1%
Beta vs SPY
0.76
Trades
13,932

What the paper reports for its own strategy

  • Selected smooth-score policy, 2017-06-28 to 2026-05-15, 10bp cost: 19.24% CAGR, 19.29% vol, Sharpe 1.01, Sortino 1.22, max drawdown -31.63%, Calmar 0.61, annual turnover 469.67%, final wealth 4.71-4.76
  • Selected policy vs 50/50 G/D: annual excess 1.78%, tracking error 3.74%, info ratio 0.48, max DD improvement 1.95%
  • Selected policy vs SPY: annual excess 3.51%, tracking error 4.32%, info ratio 0.81, max DD improvement 2.08%
  • Walk-forward expanding OOS, 2018-06-28 to 2026-05-15, 10bp: 18.64% CAGR, 19.77% vol, Sharpe 0.96, Sortino 1.16, max DD -32.93%, turnover 332.21%
  • Walk-forward rolling OOS, 2018-06-28 to 2026-05-15, 10bp: 18.02% CAGR, Sharpe 0.93, max DD -32.79%
  • Fixed-parameter OOS, 2018-06-28 to 2026-05-15, 10bp: 17.86% CAGR, Sharpe 0.93, max DD -33.19%, turnover 88.36%

An 87% growth, 13% defensive static book beats Xiong's chosen timing policy. Over his main window, it returned 20.29% a year with a maximum drawdown of -31.65%. The policy produced 19.24% with -31.63%. Its volatility was lower, 19.29% against 22.28%, accounting for the Sharpe advantage. The paper therefore asks whether 0.07 of Sharpe, 1.01 versus 0.94, can justify 469.67% annual turnover and six fitted parameters.

Xiong acknowledges in the abstract that raw CAGR does not exceed 100% G or the strongest high-G static portfolios. The conclusion goes further. High static growth exposure still wins on raw CAGR, turnover remains economically important, and the evidence establishes neither universal alpha nor a production-ready trading system. Xiong argues that the walk-forward and post-2022 checks add support, particularly on drawdown reduction. The case rests there. Of those results, the post-2022 drawdown of -19.89% against 100% G's -33.92% is the one that holds.

The exposure being timed

The trade allocates between two equal-weight ETF baskets. Growth (G) contains QQQ, XLK, VGT, SPYG, VUG. Defensive income (D) contains SCHD, VYM, VTV, FDVV, COWZ. Xiong measures the long-G, short-D spread with a Fama-French five-factor plus momentum regression, using 2,330 daily observations from 2016-12-21 to 2026-03-31. Market beta is 0.273, HML is -0.552, momentum is 0.117, and adjusted R-squared is 0.757. This is mainly a growth-versus-value tilt with a risk-on kicker. Alpha comes to 1.95%, with a t of 0.81. The paper describes the exercise plainly as exposure management rather than discovery.

The timing rule produces a continuous score, without regime labels. It starts with four direction-normalized, z-scored inputs, then standardizes their combined score again using an expanding z-score. The first input is the 21-day change in the 10-year yield, negated so that a higher reading means rate relief. SPY drawdown depth is the second. A VIX stress-relief block combines the 756-day VIX percentile with the 21-day VIX change, while a growth-crowding penalty uses the trailing 126-day G-D return. Softplus smooths the interactions instead of imposing thresholds. The aggregate score passes through a tanh, setting target growth weight at 0.5 plus MaxTilt times tanh. EWMA then smooths the realized weight at eta = 0.05. Signals are formed after the close on day t and used on t+1. Cost is 10bp, charged as 2 times the absolute weight change.

The intended return source is straightforward: add long-duration growth as long rates fall and volatility stress eases, then trim after growth has run during a quiet-rate, low-vol regime. The selected policy averages 45.06% G exposure. Xiong reports single-basket betas of 1.148 for G and 0.874 for D, so the blend implies market beta just below 1.0. It remains a fully invested, long-only equity book.

The comparison that matters

The policy wins against Xiong's preferred benchmarks, returning 19.24% CAGR at a 1.01 Sharpe. A 50/50 G/D allocation delivers 17.12% and 0.91, while SPY delivers 15.25% and 0.85. The matched rate-only score reaches 17.80% and 0.94. Vol-matched 100% G, scaled to 81.95%, reaches 17.66% and 0.94.

Static high-growth allocations win on return. The 100% G portfolio produces 21.34% CAGR. The drawdown-matched 87% G static and the best-Sharpe 89% G static produce 20.29% and 20.46%. Measured against 100% G, the policy has an information ratio of -0.28 on 9.48% tracking error.

What remains is a narrow Sharpe and drawdown case. The 87% static comes within two basis points of the policy's drawdown and trails by 0.07 of Sharpe. Compared with the fixed-structure 50% tilt from the same score family, the chosen configuration adds 0.15% a year with an information ratio of 0.15. We did not find a bootstrap or t-statistic for these policy-level comparisons. Either would help establish whether 0.07 of Sharpe over roughly nine years amounts to anything.

Turnover carries the bill. Raising the charge from 10bp to 20bp cuts the credit overlay's CAGR from 19.80% to 19.31%, a loss of half a point. Taxes are left untouched, and all 4.7 annual turns become short-term realizations in a taxable account.

Can the frozen policy survive?

Xiong reports three validations. Expanding walk-forward produces 18.64% CAGR, a 0.96 Sharpe, -32.93% drawdown and 332.21% turnover. The rolling version returns 18.02% with a 0.93 Sharpe. A fixed-parameter version is selected once from the first training year, then traded unchanged from 2018-06-28. It returns 17.86%, with a 0.93 Sharpe, -33.19% drawdown and only 88.36% turnover. During the same window, 50/50 returns 17.10% with a 0.89 Sharpe, while 100% G returns 21.13% with a 0.91 Sharpe. Freezing the parameters adds 0.76pp of CAGR and 0.04 of Sharpe over passive 50/50, while surrendering 3.27pp of CAGR to buy-and-hold growth.

Both walk-forward variants retune within the candidate grid every 63 days. The grid bounds were specified after the full-sample tilt sensitivity table, in which 50% was both the largest tested value and the best result for CAGR, Sharpe and drawdown. A selection at the grid boundary suggests that the same sample shaped both the support and the answer. Xiong discloses the issue. He cites White, Hansen and Bailey et al. and describes the in-sample grid choice as "candidate discovery."

The parameter shifts reveal the clearest instability. Full-sample selection chooses lambda_c = 0.05, MaxTilt = 50%, tau_w = 0.75, eta = 0.05. Selection on the first year, from the same candidate grid, chooses lambda_c = 0.25, MaxTilt = 20%, tau_w = 1.50, eta = 0.03.

Nearly the opposite configuration.

The post-2022 period gives the policy its strongest evidence. From 2022-01-03, expanding walk-forward returns 15.30% against 15.45% for 100% G. Drawdown is -19.89% versus -33.92%, and Sharpe is 0.90 versus 0.72. The frozen version still limits drawdown to -22.46% while returning 13.99% CAGR, compared with -33.92% for 100% G and -23.78% for 50/50. During 2022, the in-sample selected policy lost -9.53% while 100% G lost -30.61%. In 2023, it gained 29.38% against 48.31%. Xiong added this post-2022 test because the defensive basket broke down in 2020. Over the full window, D's drawdown of -36.71% is worse than the -34.35% recorded by 100% G. Both appear in his Table 7.

Limits of our reproduction

Exact VIX index history was unavailable to us. We therefore constructed the volatility-stress block from a documented substitute, SPY implied or realized volatility, rather than the paper's VIX. We sourced the 10-year yield from FRED DGS10 instead of TNX. Kenneth French daily factors were also unavailable, so we ran no attribution regression and cannot assess the reported 1.95% alpha or 0.81 t. Our window runs from 2020-01-01 to 2025-10-08. Xiong's main window runs from 2017-06-28 to 2026-05-15 and ends beyond the data we can source, preventing an aligned comparison.

The figures printed above this piece are ours. They come from a 2020-2025 run of the main selected policy using alpha 0.50, lambda_s 0.50, lambda_c 0.05, MaxTilt 0.50, tau_w 0.75, eta 0.05 and daily close rebalance. From 2020-01-01 to 2025-10-08, our run returned 111.24% total. Sharpe was 0.92, Sortino 1.20, Calmar 0.44, maximum drawdown -31.61% and volatility 21.08%. For his window, Xiong reports 19.24% CAGR, a 1.01 Sharpe and -31.63% drawdown. Our Sharpe is 0.09 lower. Drawdown is effectively identical, -31.61% against -31.63%, while volatility is 21.08% against 19.29%.

These are different tests, covering a different period and using a substituted volatility input. We charged $0.004 a share and included no slippage. Xiong charges 10bp on every weight change. We did not construct the bond/credit overlay. A weak result from our run is evidence first about our implementation. Before drawing conclusions about the score, I would examine the missing 2017-2019 period and the substituted VIX. As in our earlier FinSMART note (/articles/where-finsmart-s-returns-come-from), this score runs with an average 45.06% growth weight and implied market beta near 1.0.

Credit works better as an overlay

The bond/credit extension supplies the paper's cleanest marginal result, partly because its first version failed. A replacement score combines BAA/10Y credit relief, SPY drawdown depth, the growth-extension signal and an interaction between rate relief and credit stress. It returns 17.73% CAGR with a 0.94 Sharpe and -34.50% drawdown, trailing the original on every measure.

Keeping credit as an overlay works better. Its base is Xiong's Best Local, his name for the configuration selected by the in-sample grid: alpha 0.50, lambda_s 0.50, lambda_c 0.05, MaxTilt 50%, tau_w 0.75, eta 0.05. Adding lambda_credit = 0.10 and lambda_rxcs = 0.50 to that frozen structure lifts CAGR from 19.24% to 19.80% and Sharpe from 1.01 to 1.04. Turnover declines from 469.67% to 410.23%. Drawdown deepens slightly, moving from -31.63% to -31.92%. After absorbing the rate-only component, the rate-relief-by-credit-stress term enters with a residual HAC t of 2.36, compared with a raw t of 1.51.

I respect the discipline. The magnitude remains unconvincing. The overlay emerged from a branch that tested 793 configurations on the same reporting window, yet adds only 0.56pp of CAGR and 0.03 of Sharpe. I would retain the turnover reduction because it follows mechanically from the score rather than depending on a return estimate.

I would change my mind if the frozen configuration maintained its post-2022 drawdown profile on unseen data, using a cost assumption above 20bp. The fixed-parameter version, with 88.36% turnover, is the sole configuration here I would back with money. Across the full out-of-sample window from 2018-06-28 to 2026-05-15, it trails buy-and-hold growth by 3.27 points of CAGR, returning 17.86% against 21.13%. During the post-2022 window, the gap contracts to 1.46 points, with 13.99% against 15.45%.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each trading day t:
  Build equal-weight growth basket G from QQQ, XLK, VGT, SPYG, VUG.
  Build equal-weight defensive basket D from SCHD, VYM, VTV, FDVV, COWZ.

  Compute close-to-close basket returns R_G and R_D.
  Compute G-D trailing 126-day relative return.

  Using data available through t:
    r_t  = - expanding_z(21-day change in 10Y Treasury yield / DGS10)
    d_t  = - expanding_z(SPY drawdown from expanding peak)
    vh_t = expanding_z(756-day VIX percentile)
    vr_t = - expanding_z(21-day change in VIX)

    HighVIX     = softplus(vh_t)
    LowVIX      = softplus(-vh_t)
    VIXRelief   = softplus(vr_t)
    GrowthExt   = softplus(expanding_z(G-D trailing 126-day return))
    RateQuiet   = exp(-0.5 * r_t^2)

    CoreScore    = alpha * r_t + (1 - alpha) * d_t
    StressScore  = 0.5 * z(r_t * vh_t) + 0.5 * z(HighVIX * VIXRelief)
    CrowdedScore = 0.5 * z(GrowthExt * LowVIX) + 0.5 * z(GrowthExt * LowVIX * RateQuiet)
    RawScore     = CoreScore + lambda_s * StressScore - lambda_c * CrowdedScore
    ScoreHat     = expanding_z(RawScore)

  Target growth weight:
    wG_target = 0.5 + MaxTilt * tanh(ScoreHat / tau_w)
    wD_target = 1 - wG_target

  Smooth allocation:
    wG = (1 - eta) * prior_wG + eta * wG_target
    wD = 1 - wG

  Allocate wG/5 to each growth ETF and wD/5 to each defensive ETF.
  Rebalance at the configured close; missing execution closes skip affected trades.