Research
A continuously updated feed of research papers that pass our automated relevance screening for systematic trading — plus every paper we have published a review of, whatever it scored. Particular focus on alpha hypotheses that can be formalised and tested. The Radar also covers portfolio construction, market risk and execution where the research is directly relevant to systematic investment processes. Follow new entries by RSS.
15,697 papers screened · 250 on the radar · 21 shown
Artificial intelligence (AI) now supports investment workflows from data and prediction through research, portfolios, execution, and tool use. Technical capability, however, is not evidence of investment profitability.
Large Language Models (LLMs) increasingly use user context such as memory, profiles, and role prompts to personalize their responses.
PAPER REPORTS · Q5-Q1 one-year (252 trading day) cumulative abnormal return relative to the S&P 500, sorting n~3,393 filings by… · Per-persona 12-model panel means: Value +5.63%, LongOnly +3.84%, ESG +3.44%, SellSide +3.30%, Credit +3.29%,…
Recent advances in LLM agents enable a new paradigm for asset pricing, which we call Agentic Empirical Asset Pricing (AEAP): systems that autonomously conduct the scientific discovery process itself. We define AEAP and identify its core building blocks.
PAPER REPORTS · SEADS mean per-factor OOS Sharpe 0.25 on Panel A (JKP) and 0.16 on Panel B (CRSP/Compustat), OOS windows 2020\u20132025… · SEADS productivity 14.0 (Panel A) / 13.8 (Panel B) admissions out of a 300-candidate budget
Quantitative trading is moving from isolated predictive models toward agentic workflows that combine reasoning, tool use, memory, and feedback.
Portfolio risk assessment ordinarily relies on reliable estimates of cross-asset return covariances, which are difficult to obtain in short, high-dimensional panels.
PAPER REPORTS · In-sample standardized variance percentile of the news-only allocation: 0.69%-1.33% across four prespecified capped… · Standardized in-sample variance 0.357, 8.3% below the equal-risk (inverse-volatility) benchmark and 35.6% above the…
OUR BACKTEST · Sharpe 0.47 · Return +44.2% · Max DD -29.9%
In financial markets, a sequential policy that reacts systematically to price movements may become predictable to other market participants.
PAPER REPORTS · Exposure-matched within-stock timing alpha, principal 10-minute Qwen3.5 price text: -45.7 bps per stock-day, 95% CI… · Chart-only 10m: -29.9 bps, CI [-33.7,-26.2], 14,937 stock-days; multimodal 10m: -48.9 bps, CI [-52.7,-45.0], 14,438…
Large language models (LLMs) are increasingly used to discover trading strategies, and much of the resulting literature shares a methodological weakness: many candidate strategies are generated, the best is reported, and neither look-ahead bias nor the…
PAPER REPORTS · Best gpt-4.1 discovery (E3, RSI x volume, 453-stock universe): design Sharpe 1.69 (2017-2021), evaluation Sharpe 0.18… · Best claude-sonnet-5 discovery: design Sharpe 0.44 (2017-2021), evaluation Sharpe -0.33 / -29% (2022-2025), DSR 0.18,…
OUR BACKTEST · Sharpe -0.10 · Return -10.8% · Max DD -47.8%
Large language models (LLMs) are increasingly used in investment decision-making, yet prior work shows that they exhibit systematic, model-specific investment preferences.
PAPER REPORTS · Weekly rebalanced equal-weight long-only top-100 portfolio from Qwen3-8B scores, 427 S&P 500 stocks, 29 signal weeks /… · Realized Sharpe of the top-100 portfolios is shown only graphically (Figure 5c, axis range roughly 2–4); no point…
OUR BACKTEST · Sharpe 0.52 · Return +57.3% · Max DD -40.6%
Costly LLM features matter only if calibration lets them affect the forecast. We document a failure of this link in a next-day risk study of two broad-market funds. Full-history scoring preceded the 2022 calibration.
PAPER REPORTS · Prespecified LLM importance feature: zero improvement, 95% interval [0,0], on all four endpoints (SPY/VIX binary and… · Signed LLM repair, SPY/VIX continuous variance: improvement -0.007452, 95% interval [-0.015888, -0.001066], Bonferroni…
Two old market sayings hold that news is already priced in by the time it is published, and that the rumor is bought while the news is sold. Both place the price move associated with a piece of news before and at publication rather than after it.
PAPER REPORTS · Fade small-cap launch/partnership news (short after positive, buy after negative; enter close of day +5, exit close of… · Short any covered small cap (sentiment-ignoring benchmark, 260,472 events, same windows, 2023-2026): 15.9% annualized,…
OUR BACKTEST · Sharpe -0.23 · Return -18.9% · Max DD -51.9%
Large language models can extract richer signals from financial news than fixed sentiment lexicons, and recent work has explored feeding such signals into portfolio construction.
PAPER REPORTS · Sharpe 2.33, annualized return 95.9%, cumulative net 100.1%, max drawdown -18.3% — pure beta, GPT-4o mini, Student-t,… · Sharpe 2.44, annualized 100.0%, net 106.9%, max DD -18.3% — same configuration at 50 bps (2025)
This paper investigates whether textual tone derived from large language models (LLMs) can predict future stock returns. Using Korean news articles, we employ five LLMs to extract textual tones: BERT (KrFinBERT), DistilBERT, RoBERTa, ELECTRA, and Llama3.
PAPER REPORTS · BERT (KrFinBERT) equal-weighted quintile High-Low: 0.242% per day, t=3.772 (5.324% per month), May 2023-May 2024, gross… · BERT High-Low Fama-French 3-factor alpha 0.232% per day (t=3.582); 5-factor alpha 0.232% per day (t=3.566), same…
We study how differences in AI-generated financial recommendations are transmitted into individual portfolio choices.
A new class of software systems is transforming investment analysis. Large language model agents assembled into collaborative team structures including analysts, researchers, and risk managers are increasingly deployed across financial markets.
PAPER REPORTS · Author's own prior human practice (not the AI framework): 127.17% return vs 50.67% benchmark within sixteen months,… · Ranked first among 276 comparable peer funds nationally through the 2020 market-stress period (peer average negative);…
This paper documents an applied natural-language-processing framework for measuring the tone of Brazilian Monetary Policy Committee (Copom) statements.
We ask a representative sample to write prompts seeking spending and investing advice from LLMs, then simulate the lifetime effects of following the advice under realistic asset and labor market conditions.
Agentic AI is gaining acceptance in asset management, but governance has not kept pace: 88% of surveyed finance professionals report no operational governance framework for agentic AI despite universal awareness of its deployment, and only 24 of 75 large U.S.
OUR BACKTEST · Sharpe -1.47 · Return -45.1% · Max DD -49.7%
I revisit the exchange rate disconnect puzzle, first documented by Meese and Rogoff (1983), using generative artificial intelligence (AI) to forecast currency returns based on economic fundamentals.
PAPER REPORTS · Annualized Sharpe ratio 0.594 at 48-month lookback (cross-sectional long top-2 / short bottom-2 of 9 currencies,… · Annualized Sharpe ratios 0.577 / 0.604 / 0.594 / 0.491 / 0.467 for 36 / 42 / 48 / 54 / 60-month lookbacks (2001-2024,…
Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research.
PAPER REPORTS · FinanceHarness (Qwen3.6-27B backbone): overall rubric score 32.4%, pre-cutoff 45.7%, post-cutoff 11.8%, bootstrap SE… · FinanceHarness + GRPO (RFT): overall 32.8%, pre-cutoff 46.2%, post-cutoff 12.1% (same backbone and benchmark)
Recent advances in Generative AI have substantially improved financial sentiment analysis through post-trained financial large language models (LLMs).
PAPER REPORTS · FinSMART (static): cumulative return 264.9%, annualized return 91.5%, Sharpe 1.97, Sortino 2.40, Calmar 4.23, RankIC… · FinSMART (periodically retrained every 6 months): cumulative return 406.2%, annualized return 125.7%, Sharpe 2.41,…
OUR BACKTEST · Sharpe -0.46 · Return -48.2% · Max DD -75.2%
Parent-order execution is a core problem in algorithmic trading, where the goal is to split a large order into smaller orders while reducing execution costs.
PAPER REPORTS · Aggressive setting, April 2026, DS-v4-f: value-weighted price performance -2.26 bps vs TWAP benchmark price, i.e. · Passive setting, April 2026, DS-v4-f: -3.92 bps wbp, +1.07 bps vs TWAP and +0.71 bps vs strongest baseline; 100%…