Research
A continuously updated feed of research papers that pass our automated relevance screening for systematic trading — plus every paper we have published a review of, whatever it scored. Particular focus on alpha hypotheses that can be formalised and tested. The Radar also covers portfolio construction, market risk and execution where the research is directly relevant to systematic investment processes. Follow new entries by RSS.
15,697 papers screened · 250 on the radar · 50 shown
Frequent market instability and the lack of rigorously validated forecasting frameworks pose significant challenges for predicting market stress in the Dhaka Stock Exchange (DSE).
PAPER REPORTS · Random Forest crash-gated risk-off strategy, pooled equal-weighted, 2019-2022 test period: total return 67.26% (15.01%… · Annualized volatility 16.39% (strategy) vs. 18.11% (buy-and-hold); maximum drawdown -32.21% vs.
OUR BACKTEST · Sharpe 0.47 · Return +25.2% · Max DD -21.6%
Artificial intelligence (AI) now supports investment workflows from data and prediction through research, portfolios, execution, and tool use. Technical capability, however, is not evidence of investment profitability.
Let a finite population of n labelled examples carry a class-weighted loss, with pi*n in a rare positive class weighted by N0/N1. We study estimation of total risk from a subsample K << n under designs allocating K0 and K1 draws to the two strata.
Abstract Value-at-Risk (VaR), the most widely used measure of market risk, is typically evaluated through backtesting of point forecasts. Such procedures, however, say little about the uncertainty of the estimated quantile.
This paper develops uniform inference and certified capacity decisions for an estimated financial stability boundary. Conditional risk, temporary cross-impact, and effective risk-bearing capacity are jointly estimated from dependent observations.
PAPER REPORTS · Capacity/regret under the baseline loss convention (simulated 60-cell design, no market data): Projected safe - planned… · Under the high convention-loss calibration: projected-safe mean regret 0.107, below pointwise delta 0.120 and plug-in…
Abstract Sentiment indicators are widely used in digital asset markets, but their economic meaning remains ambiguous.
PAPER REPORTS · Expanding-window out-of-sample R-squared vs historical-mean benchmark, 2018-2026 sample, no transaction costs… · Out-of-sample directional hit rate of the ridge sentiment model: 49.8% (1d), 51.5% (7d), 48.1% (30d), no cost…
Recent advances in LLM agents enable a new paradigm for asset pricing, which we call Agentic Empirical Asset Pricing (AEAP): systems that autonomously conduct the scientific discovery process itself. We define AEAP and identify its core building blocks.
PAPER REPORTS · SEADS mean per-factor OOS Sharpe 0.25 on Panel A (JKP) and 0.16 on Panel B (CRSP/Compustat), OOS windows 2020\u20132025… · SEADS productivity 14.0 (Panel A) / 13.8 (Panel B) admissions out of a 300-candidate budget
This study investigates the forecasting performance of machine learning models and traditional econometric volatility models in predicting daily stock price volatility across selected Southern African Development Community (SADC) markets from 02 January 2015…
This study comparatively examined the forecasting performance of machine learning and traditional volatility models in predicting daily exchange rate volatility across selected economies from 01 January 2015 to 08 May 2026.
Foundation models promise accurate forecasts with little or no task-specific training, but whether they can replace models designed specifically for electricity price forecasting remains unclear.
PAPER REPORTS · Total BESS arbitrage profit over 2021-2025 (1,826 days, 1 MWh battery, 25 EUR round-trip operating cost deducted per… · Total profit Poland: unlimited-bid best 106,685 EUR (TabPFN-3) vs Oracle 123,439 EUR, 86.4% of perfect foresight; Spain…
Small-cap-inclusive equity universes contain recently listed and intermittently traded securities, so enforcing a common look-back discards a substantial fraction of the available information.
PAPER REPORTS · Annualized five-session volatility 11.17%, January 2000 - December 2025, net of all modeled execution costs… · Sharpe 0.814, same period, net of execution costs
This paper investigates the dynamic response of Shanghai crude oil futures (INE) to international benchmark price shocks and evaluates the evolution of market maturity from its inception to early 2025.
In financial markets, a sequential policy that reacts systematically to price movements may become predictable to other market participants.
PAPER REPORTS · Exposure-matched within-stock timing alpha, principal 10-minute Qwen3.5 price text: -45.7 bps per stock-day, 95% CI… · Chart-only 10m: -29.9 bps, CI [-33.7,-26.2], 14,937 stock-days; multimodal 10m: -48.9 bps, CI [-52.7,-45.0], 14,438…
Large language models (LLMs) are increasingly used to discover trading strategies, and much of the resulting literature shares a methodological weakness: many candidate strategies are generated, the best is reported, and neither look-ahead bias nor the…
PAPER REPORTS · Best gpt-4.1 discovery (E3, RSI x volume, 453-stock universe): design Sharpe 1.69 (2017-2021), evaluation Sharpe 0.18… · Best claude-sonnet-5 discovery: design Sharpe 0.44 (2017-2021), evaluation Sharpe -0.33 / -29% (2022-2025), DSR 0.18,…
OUR BACKTEST · Sharpe -0.10 · Return -10.8% · Max DD -47.8%
This paper develops a unified framework for assessing systemic risk and identifying contagion channels in the global banking system using a Temporal Heterogeneous Multiplex Graph Neural Network.
PAPER REPORTS · MSE 0.0309 on one-quarter-ahead change in log(1+CDS), out-of-sample, N=336 bank-quarter forecasts (sample 1998-2025,… · MAE 0.1342, out-of-sample
PAPER REPORTS · Out-of-sample RMSE (70-30 split, 2021–2024, no transaction costs): LLF BTC 0.541, ETH 0.256, USDT 0.244, BNB 0.676, BCH… · Out-of-sample MAE: LLF BTC 0.377, ETH 0.199, BNB 0.446, BCH 0.567, LTC 0.458, ICP 0.519, MATIC 0.612, USDT 0.137 (RF…
PAPER REPORTS · Baseline HRP: annualised Sharpe 0.741, total return +417%, max DD -85%, 76 monthly rebalances 2020-02 to 2026-05,… · HRP-family variants span Sharpe 0.701-0.749 over the same window and cost assumption (best HRP_Dynamic_94 0.749,…
PAPER REPORTS · Monthly PCI-based mean-variance portfolio (1871:02-2023:12), leverage/risk-aversion setting 6: return 2.3331,… · Monthly PCI-based portfolio, setting 8: return 2.9252, volatility 0.0260, Sharpe 16.6522 vs RV benchmark 0.8883; no…
This work complements our previous paper, which studies borrower-side strategies in decentralized lending markets, by focusing on lender-side capital allocation.
PAPER REPORTS · core-inspired strategy: 5.5% APY, $100k budget, Morpho USDC markets on Ethereum, Jan 1 2026 - Apr 1 2026, daily… · prime-inspired strategy: 3.3% APY, $100k budget, same period and cost assumption
Lead-lag relationships are widely used in financial time series, and many clustering algorithms based on them have been developed. The traditional DTW-KMedoids algorithm performs well both on the synthetic dataset and the real financial dataset.
PAPER REPORTS · Sharpe 0.866, annual return 6.21%, annual volatility 7.17%, max drawdown -63.908, hit rate 0.520, profit-loss ratio… · Sharpe 0.808 / 0.790 (KShape mod / med), lead strategy, 679 assets, same period; drawdowns -67.604 / -69.418
OUR BACKTEST · Sharpe 0.39 · Return +27.5% · Max DD -39.3%
This paper develops the first end-to-end application of cross-sectional learning-to-rank to the S&P 500 weekly options (SPXW) zero-day-to-expiration surface, integrated with margin-aware position sizing, an abstention rule driven by model uncertainty, and a…
PAPER REPORTS · Out-of-time 2025 annualized Sharpe 4.308 to 5.761 across seven sizing methods, net of Reg-T margin, tiered IBKR fees,… · Headline Edge Allocation OOT 2025: Sharpe 5.7612, Sortino 7.0291, annualized return 10.48% (excess of risk-free),…
Backtests of trading strategies are often selected after many parameter trials. A strong historical result can therefore reflect search luck rather than a persistent signal.
PAPER REPORTS · Genuine-edge discrimination AUROC 0.9890 in synthetic ground truth at headline difficulty (n = 2000, T = 1260 daily… · OOS-survival (Sharpe_OOS > 0) AUROC 0.863 and Spearman 0.611 vs realized OOS Sharpe, same synthetic headline cell
Large language models (LLMs) are increasingly used in investment decision-making, yet prior work shows that they exhibit systematic, model-specific investment preferences.
PAPER REPORTS · Weekly rebalanced equal-weight long-only top-100 portfolio from Qwen3-8B scores, 427 S&P 500 stocks, 29 signal weeks /… · Realized Sharpe of the top-100 portfolios is shown only graphically (Figure 5c, axis range roughly 2–4); no point…
OUR BACKTEST · Sharpe 0.52 · Return +57.3% · Max DD -40.6%
At 15-minute horizons, directional mean reversion is far stronger and more pervasive in cryptocurrency markets than in US equities: scored under one matched, strictly out-of-sample protocol, 90% of 183 Binance pairs carry significant directional reversal…
PAPER REPORTS · Gross edge per trade peaks near 1.3 bp of notional (BTC/ETH, 15m, 2025-01 to 2026-02) against a 5 bp cheapest maker… · Directional accuracy rises monotonically with confidence threshold, clearing 56% on the most confident bars (BTC/ETH);…
Costly LLM features matter only if calibration lets them affect the forecast. We document a failure of this link in a next-day risk study of two broad-market funds. Full-history scoring preceded the 2022 calibration.
PAPER REPORTS · Prespecified LLM importance feature: zero improvement, 95% interval [0,0], on all four endpoints (SPY/VIX binary and… · Signed LLM repair, SPY/VIX continuous variance: improvement -0.007452, 95% interval [-0.015888, -0.001066], Bonferroni…
We present in this article a non-parametric value-at-risk (VaR+CVaR) algorithm that remains accurate for an arbitrarily large number of underlying positions. The algorithm solves the two inherent problems of VaR estimation.
PAPER REPORTS · Median 99% daily VaR breach rate 1.0 +/- 0.1% across 9 parameter settings, 500 random portfolios, no transaction costs… · Per-configuration medians (no added delay): 0.97%, 1.06%, 1.09% (1260-day window, 14/30/45-day vol); 0.90%, 0.99%,…
OUR BACKTEST · Sharpe 1.58 · Return +17.6% · Max DD -3.4%
We develop parametric Entropic Value-at-Risk (EVaR) portfolio optimization for tempered stable Lévy returns.
PAPER REPORTS · ICA+NTS minimum-EVaR (EVaR_95): gross annualized Sharpe 0.616, CAGR 8.60%, annualized vol 15.28%, cumulative return… · ICA+NTS minimum-EVaR net Sharpe: 0.608 at 5bp, 0.599 at 10bp, 0.573 at 25bp; net cumulative return 630.95% at 25bp
OUR BACKTEST · Sharpe 0.24 · Return +20.5% · Max DD -37.6%
Reinforcement learning has gained increasing attention as a data-driven approach for stock trading. However, learning a policy that is both profitable and stable remains challenging due to non-stationary market behaviour and noisy reward signals.
PAPER REPORTS · DJI (test 1 Jan 2024 - 31 Mar 2025, 0.1% transaction fee both sides): annual return 21.785%±1.42, cumulative return… · FTSE (same period and costs): annual return 19.164%±1.36, cumulative return 24.596%±1.79, Sharpe 1.124±0.08, max…
Cryptocurrency time-series forecasting is a challenging task because market data usually exhibit high noise, strong volatility, non-stationarity, nonlinear dynamics, and long-range dependencies.
PAPER REPORTS · Bitcoin h=96: MSE 0.175 ± 0.009, MAE 0.314 ± 0.008 (5 seeds, most recent 20% of data as test) · Dogecoin h=96: MSE 0.413 ± 0.012, MAE 0.374 ± 0.010
We develop a geometric theory of arbitrage-free implied variance surface dynamics.
PAPER REPORTS · Out-of-sample RMSE(delta a2) improvement of full (beta,eta,psi) model over SSR-only: 17-21% at 3M-6M (215.7 vs 272.8 at… · Out-of-sample RMSE(delta a1) improvement of adding eta: 1-4% versus SSR-only at 1M-6M, essentially flat at 12M
OUR BACKTEST · Sharpe -0.84 · Return -0.1% · Max DD -0.1%
Shariah-compliant equity screening provides a transparent setting in which institutional rules determine who may own a stock.
PAPER REPORTS · SC Malaysia inclusions, 295 continuously listed liquidity-qualified events, Nov 2013-Nov 2025: matched… · 410 continuously listed inclusions (no turnover floor): +0.896pp [0,10] (p_date=0.127; p_wild=0.134) and +1.458pp…
Financial volatility is regime dependent, yet incorporating regime information into neural networks can also destabilize training. This paper asks where such information should enter a neural cross-sectional volatility forecasting model.
PAPER REPORTS · U.S. panel, 30 walk-forward windows Apr 2018 - Oct 2025, mean over 30 seeds: IC 0.5469 +/- 0.0012, ICIR 6.14 +/- 0.06,… · IC advantage over capacity-matched MLP-L: +0.0048 full sample (p<1e-4); +0.0207 top market-vol decile; +0.0322 COVID…
Financial forecasting models are typically developed in full precision, yet production deployment often requires low-precision inference to reduce memory and computational cost. Post-training quantization (PTQ) enables such deployment without retraining.
PAPER REPORTS · FP32 mean daily IC (cross-sectional Spearman, averaged over test dates, 2018-2025 walk-forward): TSMixer 0.490±0.043,… · Baseline mean daily IC over the same folds: pooled HAR 0.408, persistence (trailing 5-day realized vol) 0.315
OUR BACKTEST · Sharpe 0.41 · Return +41.7% · Max DD -37.2%
Specialist training beats generalist scale when forecasting financial statements. To our knowledge, no prior work jointly forecasts complete financial statements beyond one year, yet in a discounted-cash-flow valuation most firm value sits past that window.
PAPER REPORTS · Forma change-space R^2 0.289 on full test sample, test period 2010-2024, no transaction costs applicable (forecasting… · Forma per-horizon R^2: 39.0% at h=1 falling to 22.5% at h=20 (full sample); 38.0% at h=1 to 23.7% at h=20 (LLM sample)
OUR BACKTEST · Sharpe 0.57 · Return +17.5% · Max DD -10.5%
We calibrate credit default swaps and index tranches with elastically stopped Lévy processes: each firm defaults when the running supremum of a latent, spectrally positive distress process crosses an independent exponential barrier.
The paper considers the problem of variable selection for forecasting electricity spot prices.
PAPER REPORTS · BMT hourly rMAE averaged across six areas: 0.501 (2022-2025 out-of-sample, no transaction-cost concept; rMAE < 1 means… · BMT daily baseload rMAE_t averaged across six areas: 0.408 (2022-2025)
Cryptocurrency exchange-traded products (ETPs) listed on European exchanges provide a regulated environment for studying intraday market anomalies.
PAPER REPORTS · AUC-ROC up to 0.823 for one-bar-ahead DPOT prediction (LR, cumulative features, VIRBTC.ST), out-of-sample last 20% of… · AUC-ROC 0.821 for DPOT, LR, cumulative, VBTC.XE; 0.812 session-based
How much capital a trading strategy can absorb before its edge disappears is a causal question about how much is deployed, but it is answered with observational proxies that rest on incompatible assumptions.
Predicting financial asset returns remains one of the most difficult challenges in empirical finance, driven by the low signal-to-noise ratio and the semi-strong form of market efficiency.
PAPER REPORTS · Sector LSTM (main 3-layer, k=10, 1995-2024): mean daily long-short return 0.100% before costs, Newey-West t = 7.81,… · Sector LSTM appendix figures (after 2bp per half turn): mean daily return 0.053%, t-statistic 4.106, annualized return…
OUR BACKTEST · Sharpe -0.03 · Return -9.6% · Max DD -140.3%
Risk-aware Q-learning (RaQL) provides a model-free, two-timescale estimator for dynamic risk objectives, but its finite-budget behavior remains fragile: fixed inner-loop hyperparameters can produce unstable value estimates, persistent Bellman residuals, and…
PAPER REPORTS · Scheme 6 out-of-sample (918 daily obs, chronological test set, after 5bp turnover costs, mean over 20 seeds): Sharpe… · Scheme 0 fixed-parameter baseline out-of-sample (same test set, after 5bp costs, 20 seeds): Sharpe 0.5628 (sd 0.2281),…
OUR BACKTEST · Sharpe 0.34 · Return +16.6% · Max DD -10.4%
Informed traders are supposed to need anonymity: they profit by hiding among the uninformed. A decentralized exchange now publishes the counterparty. Every committed order, cancellation, rejection, and fill carries a persistent pseudonymous wallet address.
PAPER REPORTS · One-second out-of-sample R2: 10.88% anonymous vs 12.31% with identity, +13.2% relative (t=9.2), ridge, evaluation July… · Gradient-boosted trees, one second: 19.48% -> 20.65%, +6.0% (t=5.0), same evaluation window, no costs
Conformal prediction has traditionally been used to quantify prediction uncertainty.
PAPER REPORTS · DEV 2016-2021 (1,511 days), Config A: 28.45% annualised net log growth, Sharpe 1.336, max drawdown 27.68%, Calmar… · DEV 2016-2021, Config B: 25.84% annualised net log growth, Sharpe 1.386, max drawdown 20.26%, Calmar 1.376, annualised…
OUR BACKTEST · Sharpe 0.41 · Return +45.1% · Max DD -40.4%
I revisit the exchange rate disconnect puzzle, first documented by Meese and Rogoff (1983), using generative artificial intelligence (AI) to forecast currency returns based on economic fundamentals.
PAPER REPORTS · Annualized Sharpe ratio 0.594 at 48-month lookback (cross-sectional long top-2 / short bottom-2 of 9 currencies,… · Annualized Sharpe ratios 0.577 / 0.604 / 0.594 / 0.491 / 0.467 for 36 / 42 / 48 / 54 / 60-month lookbacks (2001-2024,…
Implied volatility surface forecasting is essential for option valuation, hedging,and risk management, but remains difficult because future surfaces are stochastic while pricing inputs must satisfy static no-arbitrage shape restrictions.
Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research.
PAPER REPORTS · FinanceHarness (Qwen3.6-27B backbone): overall rubric score 32.4%, pre-cutoff 45.7%, post-cutoff 11.8%, bootstrap SE… · FinanceHarness + GRPO (RFT): overall 32.8%, pre-cutoff 46.2%, post-cutoff 12.1% (same backbone and benchmark)
Recent advances in Generative AI have substantially improved financial sentiment analysis through post-trained financial large language models (LLMs).
PAPER REPORTS · FinSMART (static): cumulative return 264.9%, annualized return 91.5%, Sharpe 1.97, Sortino 2.40, Calmar 4.23, RankIC… · FinSMART (periodically retrained every 6 months): cumulative return 406.2%, annualized return 125.7%, Sharpe 2.41,…
OUR BACKTEST · Sharpe -0.46 · Return -48.2% · Max DD -75.2%
Parent-order execution is a core problem in algorithmic trading, where the goal is to split a large order into smaller orders while reducing execution costs.
PAPER REPORTS · Aggressive setting, April 2026, DS-v4-f: value-weighted price performance -2.26 bps vs TWAP benchmark price, i.e. · Passive setting, April 2026, DS-v4-f: -3.92 bps wbp, +1.07 bps vs TWAP and +0.71 bps vs strongest baseline; 100%…
Production forecasting systems retrain models regularly, but a retrained candidate does not necessarily outperform a continuously maintained incumbent that has continued to learn.
PAPER REPORTS · Metric is negative log-likelihood of a 3-class 300s direction forecast, not trading P&L; no Sharpe, return, alpha or… · Pooled 48 weeks (4 Aug 2025 - 5 Jul 2026, Binance USD-M + COIN-M, 8 underlyings, 3 seeds): SBS relative NLL reduction…
We audit whether candle-based machine-learning models can turn predictions of cryptocurrency extrema or short-horizon outcomes into positive Binance Spot paper policies after assumed costs.
PAPER REPORTS · Frozen mandatory-daily selector: -6.72% compounded, 2026-07-01 to 07-19, 19 cycles, 3 wins/16 losses, 31 bps… · Frozen mandatory-daily selector cost stress, same period: -4.74% at 20 bps, -6.72% at 31 bps, -10.21% at 51 bps…
OUR BACKTEST · Sharpe 0.00 · Return +0.0% · Max DD 0.0%
The distribution of a normal mean-variance mixture depends on the law of its positive mixing variable. We compare six parametric mixing laws with a grid nonparametric maximum likelihood estimator under the same determinant identification constraint.
PAPER REPORTS · Model-based lower-envelope CPT value at the robust optimum: -0.04478 at 5% annual reference return; -0.09245 at 10%… · Empirical holdout CPT value at the robust weights: -0.04867 at 5% reference; -0.10049 at 10% reference (holdout 483…
OUR BACKTEST · Sharpe 0.13 · Return +0.1% · Max DD -0.2%