A 0.99 pooled correlation says Sepp and Lucic have the accounting of trend Sharpe almost exactly right. It does not yet make the formula a forecasting tool. Across 84 futures contracts and filter spans from 5 to 520 days, the regression slope between predicted and realized Sharpe is 0.96. The inputs are two sample moments: the autocorrelation function and the drift of volatility-normalized returns.
The 0.99 comes from the European single-filter panel. Applying the same closed form to the other two designs produces 0.92 and 0.89, discussed below. The decomposition therefore captures what determines the risk-adjusted return of a trend program, with the arithmetic closing almost perfectly.
It closes in sample.
The authors state this plainly. A practitioner still has to decide what use the identity has once that limitation is accepted.
Where the returns come from
The paper divides trend systems into three families.
The European system is their term for continuous EWMA-filtered position sizing of the kind used by several large European CTAs. Its signal is an exponentially weighted moving average of volatility-normalized returns z_t = r_t / sigma_{t-1}. The position is proportional to that signal and a volatility-target weight. The American system takes its name from the North American turtle-trading lineage. It trades a fast EWMA against a slow one, adding ATR-scaled entry buffers, binary positions and trailing stops. TSMOM uses the normalized sum of the signs of z_t across M periods of L days.
For the European system, daily P&L equals the lagged signal multiplied by the current normalized return. Summing over T days yields an exact sample-path identity. Cumulative P&L becomes a geometrically weighted sum of the sample autocorrelations of z, minus one, plus the squared sample mean and a boundary term. The boundary term is of order span/T, and the authors omit it when T is much larger than the span.
With zero drift, profits require a positive weighted sum of positive-lag autocorrelations. In frequency space, the same sum is a Poisson-kernel average of the spectrum of z, while the span controls the kernel bandwidth. The authors describe trend alpha as "excess spectral mass at low frequencies". In trading terms, the lookback determines which persistence horizon the system seeks to monetize. The drift contribution rises with the square root of the span and adds at most 0.08 to median predicted Sharpe.
An extension of the Isserlis theorem then gives gross Sharpe in closed form. Excess kurtosis appears through one loading. The paper also derives a signal-turnover proxy and a leading-order net Sharpe. In Monte Carlo, 1000 paths of 50 years keep the analytics within 0.05 under white noise, AR-1 and ARFIMA for the LS(250,20) filter inside the full traded system.
The empirical universe contains 84 futures: 21 equities, 16 bonds, 4 STIR, 10 FX, 7 energy, 6 metals and 20 agriculture. Grid backtests run from 1 January 1997 to 10 July 2026. Each system has 64 configurations, and volume-based costs come from Hurst et al (2017).
After transaction costs and 2%/20% fees, the European, American and TSMOM systems produce 0.47, 0.50 and 0.55 from 31 December 1999 to 30 June 2026. The SG Trend Index returns 0.47 over the comparison period. It tracks the 10 largest CTAs and rebalances annually. Ledoit-Wolf p-values are 0.96, 0.82 and 0.62. The authors explicitly caution that failure to reject equality does not establish equivalence.
The 0.99 stays inside the sample
Both the sample autocorrelation function and sample drift are estimated over the same observations used for the backtest. The 0.99 therefore tests whether the moment decomposition remains arithmetically faithful after filter initialization, implementation effects and the omitted boundary term. It does.
The cross-sectional medians carry more information for a trader. At the 5-day span, the autocorrelation channel alone predicts median Sharpe of 0.55. It falls to 0.33 at two years. The drift channel grows with span and contributes up to 0.08. Realized European medians range from 0.48 to 0.39, leaving them 4% to 14% below the total prediction.
Forward use runs into the drift estimate. Squared sample drift has an upward bias of a/T. With their six-year minimum history, the bias is roughly 0.17 in squared annualized drift. The authors recommend debiasing or shrinking the estimate before using it prospectively, and they leave out-of-sample application for future research.
This is an attribution instrument. Span selection requires the user to supply a forecast of the autocorrelation function. The authors' own estimates show the difficulty: universe lag-1 autocorrelation of z declined from about 0.04 in the 1990s to about 0.01 after 2010.
Costs bind
Under AR-1 with phi = 0.05, proxy break-even cost ranges only from 37bp to 41bp as spans move from one week to two years. Gross Sharpe and cost drag each decline as one over the square root of the span, leaving break-even cost nearly unchanged. Realistic costs are 40bp to 60bp.
The paper says in both its introduction and conclusion that observed first-lag autocorrelations make short-memory predictability unexploitable after frictions. Positive net performance at one-to-three-month spans instead comes from the long-memory component.
The break-even calculation applies to the single filter. The traded setup uses the long-short filter LS(250,20). Under AR-1, its Sharpe remains near zero for either sign of phi: analytic 0.001 at +0.05 and 0.000 at -0.05. Both legs have the same first-lag loading, so the effects cancel.
Turnover creates the sharper problem. The proxy underlying net Sharpe stays within 8.4% of signal-level Monte Carlo turnover. For the long-short filter, however, it understates turnover in the full traded system by a factor of 1.6 to 2.3. Full-system turnover for the single filter is about 4% above the proxy. The largest miss therefore lands on the configuration actually traded.
The proxy estimates 393% per year for a single 250-day filter and 88% for LS(250,20). On the selected European configuration, realized costs reach 1.7% p.a. Absolute annualized turnover is near 2000%, driven mainly by low-volatility bond and rates contracts. The reported net Sharpe divides by gross volatility, making it a cost-adjusted gross-risk ratio rather than the Sharpe of the net return series. The authors flag this treatment. They also flag the exclusion of market impact.
One result remains unreconciled in the text. With d = 0.02 and 20bp, theory places the continuous-span optimum at 37 days. Net Sharpe is hump-shaped and peaks at 0.19 over one-to-three-month spans. Yet the grid chooses a 250-day long leg.
The 37-day result belongs to a single-filter exhibit under the pure long-memory process ARFIMA(0,d,0), leaving the difference in filter type insufficient to resolve the gap. The grid gives a reasonable separate rationale for 250/20: long-term performance at moderate turnover, with realized costs of 1.7%.
Different wrappers, similar exposure
All three systems have an average correlation of 80% with the SG Trend Index. Correlation between European and American reaches 95%. Above one-month spans, the European closed form ranks realized American Sharpe ratios at correlation 0.92 and slope 0.61. TSMOM records 0.89 and 0.73. The latter slope is close to the Gaussian sign-filter benchmark of sqrt(2/pi) = 0.80.
To a first approximation, binary sizing and trailing stops leave the breakout system carrying the same signal with an implementation discount.
Turnover supplies the main separation. The American system runs at 125% to 200% volatility-normalized turnover, compared with 300% to 400% for European and TSMOM. Its rolling cost is about half as large. Since costs bind, discretization becomes the design choice that pays.
The comparison carries one caveat. Cross-system span matching was calibrated through a parameter search on the same sample. The authors describe the exercise as in-sample calibration rather than independent validation.
Right skew can arrive without alpha
Under zero-drift white noise, closed-form skewness for T-day aggregated returns is positive at every horizon beyond one day. It reaches its maximum near half the filter span even though expected return is exactly zero.
The empirical results line up. For a 100-day single filter across the 84 contracts, median cross-sectional skewness peaks at 2.33 at a horizon of 55 days. The closed form gives 2.35. Its interquartile range remains above zero at every horizon.
Multiplying a lagged signal by the current return generates positive skew by construction. Anyone presenting convexity as evidence of edge should read Proposition D.2. The framework also accounts for weaker skew at portfolio level. Under white noise, quarterly skewness for LS(250,20) is 1.7, compared with 2.3 for the single filter. Realized quarterly skewness is 0.5 in the post-2000 sample and 1.8 over the full history.
Our adaptation failed badly
We built one automated adaptation and tested it on 32 liquid US ETFs from 1 January 2015 to 1 January 2025. Signal construction follows the paper, using volatility-normalized returns with an EWMA span of 33. Candidate variance-preserving filters cover spans from 5 to 500 and include LS(250,20). For each instrument, span selection uses the analytic net-Sharpe diagnostic after removing the a/T drift bias.
Sizing targets 15% instrument volatility with a 10% per-name cap. The portfolio rebalances weekly at the close. We charged four tenths of a cent a share in commissions and assumed no slippage.
Sharpe was 0.10. Total return reached 8.43% over the decade, with a maximum drawdown of -30.4%, realized volatility of 8.14% and a Calmar of 0.03. The paper reports 0.47 for its European system across 84 global futures from 31 December 1999 to 30 June 2026, after costs and 2%/20% fees. Our result is about a fifth of theirs.
Those figures cover different objects: a different universe and decade, a long-only implementation, and lighter costs in our run because we charged no fees or slippage.
Shorting is disabled in our version. A positive selected signal produces a long position; otherwise the strategy holds cash. This discards roughly half the signal set and leaves a residual long-risk book. A -30.4% drawdown at 8.14% volatility fits that exposure. The payoff shape reversed as well. Our long/cash implementation won 76.85% of 4,356 trades and posted a profit factor of 1.08. The paper's long-short trend filter has the opposite, structurally right-skewed profile.
Universe and period account for more of the difference. Equity sleeves make up 46% of our universe, including 11 S&P sector slices that are collinear with SPY. The paper includes 21 equity contracts among 84, alongside genuine FX, short-rate and agricultural breadth. Our period excludes 2000-2002 and 2008.
Span selection introduces another mismatch. It ranks candidates using the continuous-European net Sharpe formula, while the traded position is sign-discretized long/cash. The diagnostic therefore optimizes a payoff our implementation never realizes. Sizing probably detracted too. BIL and SHV call for roughly 30x inverse-volatility weight before being clipped at 10%. As a result, the low-volatility sleeve carries far less than its target risk while equity legs remain near full size.
Those decisions are ours, and their likely effects run in the expected direction. We cannot fully account for the gap through them, especially the depth of the drawdown. This single automated pass provides no evidence against the paper's futures result and should not be read as a verdict on the authors' work.
What would change my mind?
A rolling-window attribution would. Estimate the autocorrelation function and a shrunk drift using data available through date t, select the span, then trade it forward. The selected span should beat a fixed 250/20 after realized turnover, rather than proxy turnover.
The paper recommends exactly this exercise and leaves it for future work. Until it is run, the closed form reproduces realized European Sharpe ratios at 0.99 in sample and remains untested forward.