AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

Turnover-limited SPD portfolio selection is worth 0.4 to 2.7 points

Seven indexes, six years, no transaction costs, and the biggest single-index years sit in Hang Seng

2026-09-08 · 9 min read · US-listed equity ETFs and their end-of-day listed option chains.

Reviewing: Market timing and short-term portfolio selection based on state price density · Sebastiano Vitali, Miloš Kopa, Ruth Domínguez et al. · Read it on openalex

Our backtest of this idea

Our automated quick test, not the paper's

Lagged Option-Implied SPD Area-Ratio Long-or-Cash ETF Timing

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (3)

Run on a different market than the paper

The paper trades or times non-dividend-paying stock-market indexes using information extracted from index-option implied-volatility curves. We would trade liquid US equity ETFs (for example SPY, QQQ, IWM and sector ETFs) while estimating state-price densities from their listed ETF options; the mechanism survives because option-implied distributions can be estimated for ETF underlyings, provided the state-price-density procedure is adjusted for ETF dividend yields. Published performance figures for global index options do not transfer to the ETF universe.

The paper's own figures describe its universe and do not carry over to ours.

This is not a replication of the paper

  • The paper's exact non-dividend-paying index formulation cannot be reproduced on dividend-paying ETFs without incorporating dividend yields/carry in the option-pricing and state-price-density calculations. EOD option data support daily or multi-day signals only; the study cannot test intraday option-surface updates or execution. Reliable implementation also requires sufficient strike and maturity coverage after liquidity and stale-quote filtering, which may materially reduce the usable ETF universe.

The figures below measure what we could run, not the paper's own method, so they are not evidence for or against its claim.

Our own audit found this run does not follow the paper faithfully (17)

  • deviation left undescribed by the audit (invalidates: All paper return, risk, timing-frequency, stochastic-dominance, and figure-specific results)
  • deviation left undescribed by the audit (invalidates: Strat3-L annualized returns; Strat3-L second-order stochastic-dominance result; full-rebalancing portfolio results)
  • deviation left undescribed by the audit (invalidates: Strat3-LS annualized returns; Strat3-LS second-order stochastic-dominance result)
  • deviation left undescribed by the audit (invalidates: All Strat2, Strat4-L, Strat4-LS, and Strat4-W results)

13 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 49.0%Jul 2024
Sharpe
0.61
Total Return
49.0%
Max Drawdown
-31.2%
CAGR
9.3%
Volatility
17.7%
Beta vs SPY
0.67
Trades
56

What the paper reports for its own strategy

  • S&P500 timing, 2018: index -6.7% annualized; Strat1-L 3.8%, Strat1-LS 14.3%, Strat2-L 0.8%, Strat2-LS 8.2% (no transaction costs)
  • S&P500 timing, 2023: index 21.6%; Strat1-L 21.9%, Strat1-LS 22.1%, Strat2-L 25.5%, Strat2-LS 29.4% (no transaction costs)
  • S&P500 timing, 2022: index -18.1%; Strat1-L -18.9%, Strat1-LS -19.7%, Strat2-L -19.2%, Strat2-LS -20.3%
  • HangSeng timing, 2021: index -13.6%; Strat1-L 14.3%, Strat1-LS 42.2%, Strat2-L 16.6%, Strat2-LS 46.8%
  • HangSeng timing, 2019: index 8.8%; Strat1-L 26.0%, Strat1-LS 43.2%, Strat2-L 22.5%, Strat2-LS 36.3%
  • EuroStoxx50 timing, 2023: index 17.6%; Strat1-L 30.1%, Strat1-LS 42.7%, Strat2-L 24.9%, Strat2-LS 32.2%

The portfolio result a trader could plausibly use is the authors' more realistic version. Against equal weight, Strat3-W adds between 0.4 and 2.5 percentage points a year. Strat4-W adds 2.7 points in 2023 and trails the benchmark by 0.3 in 2020. These margins come from Table 6, which reports one-year-at-a-time statistics for the 1/N portfolio and the two strategies with limited rebalancing.

The larger figures belong to Strat3 and Strat4 under full rebalancing. Those portfolios may be turned over completely at every close, free of charge. The authors acknowledge the problem before the table: the two strategies "do not consider any turnover constraints, exposing the portfolio to high rebalancing costs", while the full-rebalancing comparison "is not fully realistic because it assumes that the portfolio manager is able to completely rebalance the portfolio allocation every day".

Single-index timing works differently. Turnover there is much lower. S&P500 Strat1 changed position 12 times in 2020 and 56 times in 2023. Hang Seng Strat1 changed 68 times in 2021, when Strat1-LS returned 42.2% against the index's -13.6%. The conclusion describes rebalancing "only few times per year" in some cases and once per week on average in others. The turnover objection belongs with Strat3 and Strat4 rather than Strat1 and Strat2.

Our own version used SPY listed options, a different market from the price indexes traded in the paper. It is separate from any replication of the results here.

Reading direction from the smile

The entire signal reduces to one number: the share of implied density above today's level relative to the share below it.

For each index and trading day, the authors collect closing implied volatilities from every quoted European option at the nearest maturity. They smooth those observations against futures moneyness kappa = K/F using a local quadratic kernel estimator from Benko et al. (2007), with bandwidth fixed at h = 0.12. A constraint keeps the implied state price density non-negative, satisfying the no-arbitrage condition. The density follows from the fitted curve and its first two derivatives.

The paper extracts two readings. Strat1 compares probability mass above moneyness 1 with mass below it, taking a long position when the upper area is larger. Strat2 looks only at the peak, kappa_max, and goes long when it lies above 1. Each has long-only (-L) and long-short (-LS) versions.

Across markets, Strat3 ranks the seven indexes by R_i, the ratio of upper to lower area. Strat4 ranks them by kappa_max. Selected indexes receive equal weights, and performance is compared with a 1/N benchmark where N = 7. Vitali, Kopa, Domínguez, Garofalo and Gianesin also introduce turnover-limited versions (-W). In these, long and short signals become overweight and underweight positions. Every underweighted index receives 10% instead of the 14.28% equal weight, with the balance divided among the remaining indexes.

Option quotes contain investors' paid views about the distribution over the next few weeks. The area split compresses that distribution into a single measure of asymmetry. The authors favor the area measure over the mode.

Datastream supplies seven indexes: S&P500, Hang Seng, EuroStoxx50, FTSE100, DAX, CAC40, Nikkei225. The sample runs from January 2018 to December 2023 and contains 6,540,879 implied volatilities. Some daily cross-sections are sparse. S&P500 averaged 84 quoted IVs per day per maturity in 2018, while CAC40 averaged 13. Fitting a second derivative to 13 points leaves little room for error, and the density depends directly on that derivative.

Most of the money comes from Hong Kong

Hang Seng supplies the standout years. In 2021, the index returned -13.6% annualized while Strat1-LS gained 42.2% and Strat2-LS gained 46.8%. The index lost 12.1% in 2023 as Strat1-LS made 31.7%; 2022 matches an index return of -11.8% with 24.1% from the strategy. Hang Seng also generated the largest number of short signals. Strat1 was short for 130 days out of 256 in 2021, and Strat2 was short for 113.

Only two other index-years resemble those outcomes. EuroStoxx50 2023 combines a 17.6% index return with 42.7% for Strat1-LS. DAX 2022 combines -8.1% with 22.5%.

Nikkei225 provides an accidental control case. Strat1 recorded zero short days in 2019 and 2020, four in 2021 and three in 2022. It therefore followed the index, including both returning 16.8% in 2019. Its eventual signal failed in 2023: the index gained +24.9%, while Strat1-LS returned 0.0%.

All four strategies also lagged the index in four other index-years. EuroStoxx50 2018 returned -14.4%, compared with -27.3% for Strat2-LS. FTSE100 2021 gained 12.6%, while Strat1-LS lost -2.4%. DAX 2018 returned -18.0% against strategy results of -19.7%, -21.4%, -20.3% and -22.7%. CAC40 2020 returned -2.1% against -3.8%, -5.5%, -3.5% and -4.8%. On the S&P500, every version lost more than the index in 2022: -18.1% for the index, -18.9% for Strat1-L, -19.7% for Strat1-LS, -19.2% for Strat2-L and -20.3% for Strat2-LS.

The mechanism supports a narrower interpretation than general return forecasting from the smile. It earns money when density mass remains below current moneyness for months and the market then declines. Persistence carries the result. Long-short beats long-only in most index-years, while the Hang Seng 2021 area signal was close to evenly divided, with 130 short days out of 256 for Strat1.

Portfolio evidence is more persuasive, and the authors subject it to a genuine test. Under full rebalancing, all four strategies beat 1/N in five of the six years. In 2021, 1/N returned 13.0% against 37.4% for Strat4-LS. The corresponding 2020 figures were 3.5% and 28.9%. Strat3-L in 2018 was the lone exception, returning -9.6% versus -8.2%. Over the full period, Strat3-L, Strat3-LS and Strat4-LS second-order stochastically dominate 1/N. The gain comes with higher volatility: 2022 std dev was 16.7% for 1/N and 20.6% for Strat4-L.

Limited rebalancing cuts the advantage sharply. Strat3-W exceeded 1/N by 1.4 points in 2019, returning 19.2% against 17.8%, and by 2.5 in 2023, with 13.3% against 10.8%. Strat4-W added 2.7 in 2023, returning 13.5% against 10.8%. In 2020, 1/N produced 3.5%, Strat3-W 3.9% and Strat4-W 3.2%.

The prose says both -W strategies outperform 1/N in all the considered years, although the paper's own table places Strat4-W below the benchmark in 2020. Its conclusion says transaction costs "would not have too great an impact because the strategies usually mimic the benchmark". It also argues that Strat3-W and Strat4-W rebalancing costs "should be even lower because the positions are never completely closed or reopened". An annual edge of 0.4 to 2.7 points is comparable in scale to costs of 0.4 to 2.7 points a year. No costs are charged anywhere in the study: the transaction costs were not considered.

Several reported figures resist reconciliation. The 2022 benchmark appears as -6.8% in the full-rebalancing table and -7.5% in the limited-rebalancing table, and we could not reconcile the difference. The conclusion says Strat1 always exceeds Strat2 on return and risk, while the Hang Seng table reports 46.8% for Strat2-LS against 42.2% for Strat1-LS in 2021. The portfolio section refers to an equally weighted allocation across the five indexes considered previously, although seven are used and the weight table specifies N=7. The Nikkei table records 305 trading days in 2023, versus 244 to 248 across the five earlier years.

Can a Q-measure density time a P-measure return?

The density is risk-neutral. The paper explicitly declines to transform it to the physical measure, citing both the difficulty of the transformation and its model of a risk-neutral investor. That choice makes the economic reading slippery. Put demand shifts mass below moneyness 1 regardless of whether the index subsequently falls. Persistent left tilt can instead indicate costly downside insurance, a reading about the skew premium rather than expected returns.

A short position based on that signal collects the premium during falling markets and pays it back during rising ones. The Hang Seng rows resemble premium collection more than forecasting. The index declined in 2021, 2022 and 2023, while Strat1 remained short for 130 days out of 256 in 2021 alone. The position was therefore present through much of a three-year fall. Kostakis, Panigirtzoglou and Skiadopoulos (2011), cited by the paper, argue that the Q-to-P transformation matters for this exact portfolio choice. A regression of realized index returns on the area ratio, controlling for variance and skew premia, would distinguish the two explanations.

Implementation brings three concessions and two defenses. Transaction costs are omitted throughout, and position-change counts stand in as a proxy: 56 changes for S&P500 Strat1 in 2023 and 44 in 2022. Closing IVs and closing index prices are assumed simultaneous. The defense is that markets are calm near the close and the fit takes under a second. Fast computation resolves latency while leaving the timing match between the option close and index close unproven. The authors also acknowledge that direct index trading is impossible and offer no defense for it. We did not find a Sharpe ratio or aggregate six-year result. Instead, the paper reports year-by-year means, standard deviations, tail measures and the SSD test.

Our SPY run has narrower scope

Our long-or-cash SPY book returned 48.96% in total from 2020-01-01 to 2024-07-01. Its Sharpe was 0.61, Sortino 0.70, Calmar 0.30 and volatility 17.74%, with a maximum drawdown of -31.23%. This was one automated pass on a substitute market. It does not replicate the paper or test its claim.

The authors use price indexes deliberately, avoiding dividend adjustments in the density. Those indexes cannot be traded directly, so our version uses SPY and its listed American options. Dividend yield and carry consequently enter the forward and the density, unlike the paper's index formulation. American option closes also require a dividend-adjusted Barone-Adesi-Whaley conversion into European-equivalent IVs.

End-of-day data permits daily signals only. We have no intraday surface updates or intraday execution. Filtering for liquidity and stale quotes further reduces strike and maturity coverage. As in the paper, the density uses quoted strikes without synthetic tails. A single ETF cannot test the seven-index ranking, the -W weight table or the stochastic dominance result. Our long-or-cash construction also differs from every -LS result in the paper.

The rule holds long 100% SPY when the lagged density assigns more mass above kappa = 1 than below. Otherwise it holds cash at the effective fed funds rate. We retain the paper's h = 0.12 bandwidth, impose a nonnegative SPD and integrate by the trapezoidal rule on each side of the boundary. A signal from day d executes no earlier than the day d+1 close. Trading costs are 5 bps one way on SPY notional plus four tenths of a cent a share. Modelled slippage is zero, an assumption that favors a daily rule.

Start with the 17.74% volatility and -31.23% drawdown before judging the 48.96% return. A strategy that owns SPY unless the density tilts downward should remain close to SPY. The paper's US counts show the same behavior: Strat1 shorted the S&P500 for 8 days out of 251 in 2020 and changed position 12 times.

The nearest paper comparison is S&P500 Strat1-L, a long-only strategy on the price index with no costs. It returned 21.6% against 20.1% for the index in 2020, 25.8% against 23.2% in 2021, -18.9% against -18.1% in 2022, and 21.9% against 21.6% in 2023. Our figures remain the 48.96% total return and 0.61 Sharpe reported above.

These results measure different objects. The paper gives annualized per-year means for a non-tradable price index without costs. Ours comes from a costed ETF portfolio with a one-day execution lag. A gap in either direction provides no evidence about the authors' claim. Any disappointment in our version first implicates the one-day lag, the American-to-European conversion and the single-name universe, all choices we made.

Evidence that the area ratio predicts realized index returns after controlling for variance and skew premia would change my mind.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

For each SPY option signal date d:
  1. Load point-in-time SPY close, dividends, maturity-matched rates,
     and observed call/put records available by d.
  2. Select the nearest quoted expiration with positive DTE.
  3. Convert valid American option closes to equivalent-European IVs
     using a dividend-adjusted Barone-Adesi-Whaley model.
  4. Compute forward moneyness kappa = K / F, where
     F = S * exp((r - y) * tau).
  5. At each quoted strike, fit a Gaussian-kernel local quadratic IV
     smile with bandwidth 0.12, imposing nonnegative SPD in the fit.
     Skip the signal if conversion or a full-rank constrained fit fails.
  6. Evaluate and normalize the SPD over the finite quoted-strike grid.
     Split any observed trapezoid crossing kappa = 1 at that boundary.
  7. Compute A_above, A_below, and R = A_above / A_below.
  8. Mark SPY favorable when A_above > A_below, equivalently R > 1.

At the eligible execution close at least one trading day later:
  - If favorable, target a 100% long SPY weight.
  - Otherwise, target 0% SPY and allocate capital to cash.
  - Apply trading costs to absolute SPY notional traded.
  - Skip an order if its observed execution close is unavailable.

Daily is the primary executed variant. The weekly variant evaluates on
Friday using the most recent signal at least one trading day old.
Record the SPD mode and expected-upside ranking as diagnostics rather
than as controls of the primary one-ETF portfolio.