A rolling four-year tally of an LLM's macro calls beats carry across G10 currencies. The four years drive the result.
We ran no test of this paper, and every figure that follows belongs to Izadyar. Two constraints stopped us. We lack the G10 spot and one-month forward series needed to trade the paper's market. Currency ETFs would substitute a different market for one built on actual currency excess returns, forward-discount and carry controls, and country-specific macro releases. We also could not reproduce the classification step. The particular GPT-4o, GPT-3.5 and DeepSeek-V3 versions are unavailable, while model drift changes the AIFX signals themselves.
Counting the model's votes
Izadyar feeds GPT-4o each of 174,820 economic calendar entries from January 1996 through October 2024. The prompt includes the release title, actual value, consensus forecast, previous value and currency name. It omits the date. GPT-4o supplies a brief analysis and assigns one of three labels: STRENGTHEN, WEAKEN, or INSIGNIFICANT OR UNCERTAIN. With four numerical fields as inputs, the task is numerical reasoning in prompt form, rather than sentiment extraction from text.
Investing.com provides the calendar. It contains 544 unique indicators for the G10, with a heavy US tilt: 72,060 releases versus 4,864 for NOK.
AIFX measures the share of STRENGTHEN labels minus the share of WEAKEN labels across a trailing window of one to sixty months. At every month end, it ranks nine exchange rates, each quoted in USD per foreign unit. The portfolio buys the top two, sells the bottom two and rebalances monthly. Bloomberg supplies London-close spot and one-month forwards.
The economic case runs through expected policy-rate differentials. If GPT-4o applies Taylor-rule arithmetic, reading inflation above forecast as expected tightening and currency appreciation, AIFX accumulates those judgments. The trade requires FX to absorb them slowly. Example 3 shows GPT-4o following exactly this chain on a UK CPI release.
The paper reports 286 months from January 2001 to October 2024, gross of costs. At a 48-month lookback, annual return is 4.354% on 7.335% volatility, for a Sharpe of 0.594. Carry produces 0.418 over the same window, Value 0.328, Dollar Carry 0.255, Dollar 0.030 and one-month Momentum minus 0.054.
Against those five factors, annual alpha is 3.24%, with a Newey-West standard error of 1.32. The alpha equals 74% of mean return, and R-squared is 24.2%. The abstract promotes a Sharpe "exceeding 0.7". Table 3 reports performance across lookback periods and reaches 0.604 at 42 months. Its five listed windows are 36, 42, 48, 54 and 60, leaving any Sharpe above 0.7 at an unreported tabular lookback. Figure 1 places the 0.7 on its curve.
The paper also constructs a Weighted AIFX index from GPT-assigned importance scores of 1 to 100. Izadyar writes that it "enhances the Sharpe ratios for nearly all lookback windows". A higher Sharpe could appear there, although Figure 6 gives the levels only as a curve. The paper describes turnover as moderate and concludes that trading costs should have little effect. We did not find a turnover figure or spread assumption anywhere.
Why does it need four years?
Table 3 covers only 36 to 60 months, with Sharpes of 0.577, 0.604, 0.594, 0.491 and 0.467. Shorter-lookback results appear solely as a curve in Figure 1, while the paper locates predictive power at 36 to 60 months. This looks like an unusually slow macro-fundamental momentum factor. Izadyar connects it to Dahlquist and Hasseltoft (2020), who examine whether past trends in major macroeconomic indicators forecast currency returns.
Three interpretations remain. FX may need years to incorporate accumulated fundamentals. Persistent monetary regimes could make the four-year count a slow policy-stance measure. The grid may simply have selected the window. Publishing all sixty lookbacks weakens the last interpretation. The mechanism section revives it by focusing "for lookback periods ranging from 36 to 60 months, where the predictive power of the strategy is strongest," in the author's own words.
Within that band, removing Inflation and Price Indices cuts the Sharpe by 39.79%. Removing Employment costs 5.07%, while Broad Economic Activity costs 3.92%. Each of the other five categories has a negative marginal contribution. Interest Rates and Monetary Policy is the largest, at minus 12.23%.
Limiting AIFX to the top three categories raises Sharpe ratios by roughly 40% on average. Full-sample leave-one-out results determined those categories, making the gain in-sample by construction. Inflation supplies most of their contribution, with employment adding five points. The interest-rate result sounds contradictory until the bucket's contents are examined. The paper's category table includes bank lending, M2 money stock, the Fed's balance sheet and the deposit facility rate, a collection extending well beyond policy decisions.
Good news carries the forecast
The Strength ratio by itself exceeds 0.5 Sharpe at most lookbacks. Beyond 16 months, the Weakness ratio becomes slightly negative. Its panel-regression coefficient is insignificant in nearly every specification, while the Strength coefficient is significant. Counts offer no explanation: 55,129 positive labels, 58,575 negative and 61,116 neutral.
Izadyar argues that negative news prompts a larger, faster exchange-rate response, using up its information over the short run. Release-day returns support the immediate-response portion. Positive days gain plus 1.92bp, neutral days lose minus 0.28bp and negative days lose minus 4.99bp from 2008 to 2024. Drift continues the next day: negative-news currencies lose another 0.73bp, while positive-news currencies add 0.30bp. The monthly result alone carries the exhaustion account.
Memorization under four tests
When asked to date its own inputs, GPT-4o averages 5.68% precision, 4.17% recall and 2.36% F1 across years. Annual F1 has a minus 0.36 correlation with annual strategy return. The distribution of mistakes matters. Table A.4 separates the year-guessing exercise by year, guesses, correct guesses, precision, recall and F1. A total of 73,860 guesses fall in 2023, producing 75.90% recall and 5.13% precision. For 2024, the model makes 2,566 guesses even though 3,946 releases occurred then. Its answers cluster around the cutoff era. The low accuracy leaves that 2023 anchoring short of evidence for look-ahead.
The cutoff exercise compares GPT-3.5, dated September 2021, with GPT-4o, dated October 2023. Monthly correlations between their indices range from 0.67 to 0.94 and persist across subperiods. On 5,958 observations, the difference-in-differences interaction is 0.01 with a standard error of 0.02.
The dependent variable is the AIFX index, computed "using a one-month lookback window" in the author's description. This test therefore concerns the one-month index. The reported Sharpes come from the 36 to 60-month signal. The exercise shows that the two vintages classify releases similarly, and Izadyar concedes that the two-year gap may restrict its power.
The fourth test has the most force, in both directions. A portfolio based entirely on GPT-4o's recollection of realized monthly currency direction records a Sharpe of 0.914 from January 1996 to October 2023. Its skewness is 2.52 and excess kurtosis is 18.61, with gains concentrated from 2008 to 2023. The model remembers.
Regressing AIFX on that portfolio leaves annual alpha between 5.52% and 6.06%, with every estimate significant at 1%. Betas range from minus 0.36 to minus 0.44, and information ratios from 0.75 to 0.85. Yet orthogonality to a memory factor leaves the central design exposed. A model whose knowledge reaches October 2023 scored the entire 1996-2024 history in one pass. The sample contains Twelve post-cutoff months, and the cumulative-return chart marks the date. We did not find a separate return or Sharpe for them.
The conclusion nevertheless says these exercises "rule out look-ahead bias and demonstrate that the AI model's performance stems from genuine reasoning rather than memorized information". That wording is too strong when the author's memory portfolio earns 0.914 and Table A.4 records 75.90% recall for 2023. Mitigated is the honest word.
DeepSeek-V3 offers the paper's strongest supporting evidence. In Figure 5, its Sharpe curve closely follows GPT-4o across the 1 to 60-month range, though neither model has tabulated levels there.
One result would change my view: Twenty-four months of post-cutoff signal, produced by a pinned model from release vintages as they appeared. Run it at the 48-month window and retain all eight categories, rather than selecting the three that won the leave-one-out.