AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

ESGU's Monthly Correlation With SPY Rounds To 1.00

Chen pairs each ESG fund with its regional benchmark, and the window that reverses the verdict does not reconcile.

2026-09-08 · 7 min read · US-listed equity ETFs

Reviewing: Risk-Adjusted Performance of ESG ETFs: Benchmark, Geographic, and Market Exposure Analysis · Victor Chen · Read it on openalex

Our backtest of this idea

Our automated quick test, not the paper's

Monthly Benchmark-Relative Tactical ESG Allocation

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (1)

Our own audit found this run does not follow the paper faithfully (8)

  • deviation left undescribed by the audit (invalidates: All paper-predicted full-period metrics; all 2016-2019 metrics; all complete 2024-2026 metrics; direct reproduction of the paper's full correlation matrix and full-period regressions)
  • Formula-consistent annualization of full-period mean monthly returns: The implementation reports the paper's stated arithmetic and geometric annualization formulas rather than treating the internally inconsistent Table 3 annualized-return values as executable targets. (invalidates: Direct reproduction of Table 3's full-period annualized-return values for ESGU, SPY, QQQ, VICEX, ESGD, and VEA from the table's reported mean monthly returns.)
  • Common DGS3MO risk-free rate in Sharpe calculations: The implementation applies one point-in-time DGS3MO series on a common annualized time base to all six instruments instead of reproducing the mutually inconsistent implied risk-free rates in the paper's reported Sharpe ratios. (invalidates: Direct reproduction of the exact Sharpe ratios in Table 3 and use of those six values as a joint validation target for a common-DGS3MO implementation.)
  • Trailing tactical allocation procedure: The 12-month positive relative-performance gate, beta bounds, one-month signal lag, conditional switching, and implementation costs are a tactical extension rather than procedures performed in the source paper. (invalidates: Any claim that the paper's Table 3, Appendix D, correlation matrix, or regressions establish the returns, turnover, costs, or efficacy of the conditional tactical strategy.)

4 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 53.4%Jul 2024
Sharpe
0.56
Total Return
53.4%
Max Drawdown
-37.6%
CAGR
10.0%
Volatility
21.1%
Beta vs SPY
0.93
Trades
125

A U.S. ESG ETF whose monthly returns correlate 1.00 with SPY is a fee decision wearing the costume of a portfolio decision. Chen prints that number in the full-sample correlation matrix, and it is the most useful thing in the paper.

Six funds and one Excel file

Chen compares six funds on monthly closing prices from December 2016 to May 2026. The window holds 114 monthly closes, so 113 monthly returns. Two ESG funds: ESGU (iShares ESG Aware MSCI USA) and ESGD (iShares ESG Aware MSCI EAFE). Two benchmarks paired to them by region: SPY for ESGU, VEA for ESGD. Then QQQ as a growth comparison and VICEX, the USA Mutuals Vice Fund, as a non-ESG comparison.

Chen does all of it in Excel, and all of it is descriptive. Monthly mean and standard deviation, annualized return and volatility, Sharpe using DGS3MO converted to a monthly rate. Cumulative growth of $100, and maximum drawdown as the worst peak-to-trough decline of that $100 series. A six-by-six correlation matrix, plus scatterplot trendline regressions of each ESG fund on its benchmark. Both of those are computed on the full period only. Five metrics get recomputed for each of four sub-periods: mean return, annualized return, annualized volatility, cumulative return and Sharpe. The windows are 2016-2019, 2020-2021, 2022-2023 and 2024-2026.

The paper is honest about where an ESG premium would have to come from. On one side, Friede and co-authors aggregate more than 2,000 empirical studies, and Whelan and co-authors aggregate 13 corporate meta-analyses covering 1,272 studies plus 2 investor meta-analyses covering 107 studies. Chen adds the cost-of-capital argument: heavy ESG inflows make these firms cheaper to finance. On the other side, the greenium. Demand pushes prices up and expected returns down. Hartzmark and Sussman find no evidence that high-sustainability funds beat low-sustainability ones. So the ESG trade is either a slow-burning risk premium or a valuation penalty already paid.

What Chen finds over the full sample: ESGU 1.06% a month, 11.99% annualized, 15.99% annualized volatility, Sharpe 0.60, max drawdown -26.40%. SPY 1.07%, 12.12%, 15.64%, Sharpe 0.62, max drawdown -24.80%. ESGD 0.59%, 5.95%, 15.51%, Sharpe 0.38, max drawdown -30.79%. VEA sits at 0.61%, 6.20%, 15.96%, Sharpe 0.39, max drawdown -30.70%. QQQ was the sample's best fund at 18.49% annualized, Sharpe 0.85, drawdown -33.07%. VICEX was the only loser: -2.09% annualized, Sharpe -0.26, drawdown -39.98%.

Chen reaches my spine himself. The abstract says their returns "appear largely explained by benchmark exposure, geography, and sector composition." Investors are also told not to expect "distinct, superior performance solely from ESG screening, as performance is largely based on other factors affecting the fund's actual holdings." I agree with that. My complaint sits elsewhere. ESGU is an ESG-screened version of an MSCI USA parent universe, so the near-identity with SPY is close to mechanical. And the paper reports no tracking error, no standard errors and no numeric beta to make it more than that.

The pairing is the contribution

Friede and co-authors aggregate corporate ESG-performance studies. Hartzmark and Sussman study fund flows and sustainability rankings. Chen's own step is the pairing discipline: each ESG fund gets measured against a benchmark in its own geography.

That matters for the comparison an investor actually makes. ESGU beat ESGD by 604 basis points annualized (11.99% versus 5.95%). Read alone, that looks like evidence that U.S. ESG screening works better than international ESG screening. SPY beat VEA by 592 basis points (12.12% versus 6.20%).

Almost the entire gap is region.

Chen backs this with holdings: ESGU's index puts over 30% into U.S. technology-prominent companies while ESGD sits heavy in Financials and Industrials, per BlackRock. Two funds with the same label, two very different absolute outcomes, and the label explains roughly a tenth of a percentage point of it.

Appendix C puts the full-sample ESGU-SPY correlation at 1.00 and ESGD-VEA at 0.99. Both funds are ESG-screened versions of MSCI parent universes, which is why the near-unit correlation is close to mechanical rather than a discovery. Chen writes that the comparisons indicated high R-squared values and that both beta values were close to 1. The two pair regressions are printed as Figures 4 and 5, with additional charts in Appendix B. No numerical alpha, beta or R-squared appears anywhere in the text. I wanted those values. No tracking error is reported, and no numeric beta either. So the 0.60 versus 0.62 Sharpe difference cannot be told from zero. The paper prints no standard error on any Sharpe or beta difference.

Which subperiod would you trade on?

The ESG edge changes sign across windows. In 2016-2019, ESGU ran a 1.03 Sharpe against SPY's 0.99, and ESGD 0.56 against VEA's 0.48. In 2020-2021, ESGU 1.25 against SPY 1.18. ESGD 0.47 and VEA 0.48 are the same fund for practical purposes. 2022-2023 was weak for all four paired funds: ESGU -0.15 against SPY -0.09, ESGD -0.20 against VEA -0.23. QQQ printed 0.04 in that window, VICEX -0.76. Then 2024-2026 flips the order: SPY 0.99 against ESGU 0.50, VEA 0.80 against ESGD 0.61.

Chen states the implication directly. Outperformance "was not consistent across funds, time windows, or metrics." He also flags that the windows were drawn around known regimes, which makes them an in-sample partition. The sample itself runs about 9.5 years, though the text calls it ten.

The 2024-2026 reversal is the one I would not build on. That Sharpe gap is a volatility gap. ESGU returned 14.91% annualized against SPY's 15.63% in that window. ESGU's volatility is printed at 20.64% against SPY's 11.10%. A fund whose full-sample monthly correlation with SPY rounds to 1.00 cannot run at 1.86x SPY's volatility. The paper reports no sub-period correlations. So the contradiction sits between the full-sample matrix in Appendix C and one row of Appendix D. It also sits against Chen's own sentence that both beta values were close to 1.

The 20.64% is the exact figure printed for ESGU in the 2022-2023 window as well. ESGD and VEA volatility repeat too: 11.14% and 11.45% appear in both 2016-2019 and 2020-2021. And in 2020-2021 ESGU is shown at 25.37% annualized with a cumulative return of 18.07% over 24 months. Cells in Appendix D do not reconcile. Chen's headline conclusion survives that, because the conclusion is that the pairs are near-identical. The specific claim that benchmarks have pulled ahead since 2024 does not.

One more measurement issue. Returns come from closing prices, and the paper never states whether they are dividend-adjusted. Within a pair the bias points the same way for both legs, so ESGU against SPY holds up. Across funds with different payout policies it does not. VICEX is the sample's only mutual fund, and a mutual fund's NAV drops on distributions. That matters for the fund the table ranks last on return (-2.09% annualized), worst on drawdown (-39.98%) and alone in negative Sharpe (-0.26), with the second-highest volatility at 17.30%. I would have expected a vice fund to hold up better than that.

We ran it as a switching book

We built a monthly switching version of Chen's pairing logic and backtested it from 2020-01-01 to 2024-07-01. Each month-end, for each pair, we form 12 monthly returns from real unadjusted closes. We also average the DGS3MO observations available in that month. Hold the ESG fund if its information ratio against the benchmark is positive. Hold it only if its Sharpe advantage is also positive. And only if its rolling beta sits between 0.90 and 1.10. Otherwise hold the benchmark. Half the book in the U.S. sleeve, half in the developed-international sleeve, executed at the next trading day's close. We charged 5 bps per unit of one-way turnover plus $0.004 a share with a $1 minimum, and modeled zero slippage.

The book returned 53.36% total over those 4.5 years. Sharpe 0.56, annualized volatility 21.10%, maximum drawdown -37.56%. Sortino 0.74, Calmar 0.27.

Our 0.56 is below the paper's 0.62 for SPY and 0.60 for ESGU. The setups differ: a switching two-sleeve book over 4.5 years against single funds over December 2016 to May 2026. I am not reading a winner out of that difference.

Two honest weaknesses in our version. The benchmark is the default holding whenever the three conditions do not all fire or history is short. Our sizing rules also conflicted: 50% sleeves against a 16.67% per-position cap. A 16.67% cap makes 50% sleeves infeasible, so the outcome depends on which one binds. We had no usable VICEX price history for 2020 through 2024, so none of the paper's VICEX statistics describe our run.

One automated pass is evidence about our implementation before it is evidence about the idea. The mechanism is where I would put the blame, and Chen's correlation matrix points at it. A 12-month information ratio computed on the difference between two series whose full-sample correlation is 1.00 is an estimate on almost pure noise. Every switch pays 5 bps of one-way turnover, with no slippage modeled on top.

A total-return series with a printed tracking error, and standard errors on the ESGU-minus-SPY difference, would change how I read this. If the 9.5 point volatility gap in the 2024-2026 row is real, ESGU deserves a fresh look. That gap and a full-sample correlation of 1.00 with SPY describe different funds. I do not believe the cell.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

At each completed month-end:
  For each pair (ESGU, SPY) and (ESGD, VEA):
    1. Form monthly returns from real, unadjusted month-end closes.
    2. Average point-in-time DGS3MO observations available during the month.
    3. Convert the annual-percent yield to a monthly risk-free return.
    4. Using the trailing 12 monthly observations, calculate:
         - compounded relative return
         - annualized tracking error and information ratio
         - ESG and benchmark Sharpe ratios
         - rolling beta, alpha, correlation, R-squared, and drawdowns
    5. Select the ESG ETF if:
         information_ratio > 0
         AND ESG_Sharpe - benchmark_Sharpe > 0
         AND 0.90 <= beta <= 1.10
       Otherwise select the natural benchmark.
    6. If history or risk-free data are insufficient, select the benchmark.

On the next trading day's close:
  - Allocate 50% to the selected US instrument.
  - Allocate 50% to the selected developed-international instrument.
  - Skip an order lacking a real execution price and retain the filled position.
  - Compute one-way turnover as 0.5 × Σ|target weight - prior weight|.
  - Deduct applicable transaction costs from net performance.

QQQ and VICEX are comparison instruments only and are not eligible tactical holdings.