A Balancer-style pool charging a 3% swap fee gives a 60/40 investor something unusually concrete: mandate compliance visible in the reserves. Verification needs neither a monthly factsheet nor trust. The paper's lasting contribution lies there. Its contest with VBIAX, EQL and EDOW rests on shakier ground. The authors find that the G3M outperforms incumbent funds on both metrics for certain fee ranges across these historical case studies, using arbitrage-only order flow. The winning ranges depend on three measurement conventions selected after the simulation.
Begin with the mechanism. A geometric mean market maker (G3M) holds N assets, reserves x and target weights w, while enforcing the invariant L = prod x_i^{w_i}. Its internal marginal prices obey p_ij = w_i x_j / (w_j x_i). An external price move therefore leaves the pool quoting stale relative prices. Arbitrageurs trade until its internal price meets the external price. Feinstein, Florescu and O'Leary use the resulting property: without a fee, arbitrage always returns the pool exactly to its target weights. The rebalancing agent remains outside the fund and receives whatever value he extracts from liquidity providers.
A zero-fee G3M tracks perfectly while surrendering the entire value of each rebalance to the first trader who arrives. The paper contributes a fee design that preserves clean accounting, then puts the resulting portfolios against conventional funds. The authors describe this as the first direct comparison of an AMM-held portfolio with traditional mutual funds and ETFs. They also say previous work has not quantified tracking error for an AMM-generated position.
Every reserve update is split. The largest common proportional scaling becomes fee-free minting or burning, while the fee applies to the residual change in composition. The liquidity update is g(rho) = (1 + alpha(rho))^gamma times prod (1 + rho_i)^{(1-gamma)w_i}, which interpolates geometrically between the fee-free G3M and a purely proportional update. From it comes the bid-ask matrix pi_ij = [gamma + (1-gamma)w_j] / [(1-gamma)w_i] times x_i/x_j. Whenever gamma is positive, pi_ij times pi_jk exceeds pi_ik. Routing through a third asset cannot avoid the charge. Breaking a trade into many pieces remains free only when each piece shares the same minimum-scaling asset; every other split costs more.
What does gamma = 3% buy?
Theorem 5 and its corollary deliver the main result. With competitive, value-maximizing arbitrage, profitable arbitrage is absent if and only if every realized weight satisfies hat w_i >= (1-gamma) w_i. Each weight must remain within [(1-gamma)w_i, gamma + (1-gamma)w_i]. Total L1 drift stays below 2 gamma (1 - min_i w_i). Competitive arbitrage is required for both bounds. Without it, the ex ante guarantee disappears.
The band has a direct price. For a 60/40 mandate with gamma = 3%, equity remains within [58.2%, 61.2%], bonds occupy the mirror range, and total drift cannot exceed 3.6 percentage points. Anyone with the reserves and contemporaneous prices can verify those limits immediately. No discretion enters the inspection. While weights remain inside the band, the pool drifts and liquidity providers retain part of each rebalance instead of giving away the full proceeds. The paper states the bargain plainly: "Positive fees permit temporary mis-weighting while allowing LPs to retain a portion of the proceeds from any rebalancing."
Small target weights expose the loose side of the design. Gamma alone determines the width; w_i does not. At gamma = 3.32%, the lower edge of EDOW's dominating range against NAV under the economic mandate, a Dow constituent targeting 3.33% can range from [3.22%, 6.54%]. Across the 30-name portfolio, permitted L1 drift reaches 6.4 percentage points.
The simulations and their winners
The empirical section builds G3M pools from the underlying assets of three actual mandates. Arbitrage supplies all order flow. Those agents trade frictionlessly in external markets, no more than once per day at the close. Bloomberg total return data are used, with income reinvested in the asset that paid it. Gamma is swept across [0%, 10%]. Each run lands on a frontier of CAGR versus annualized tracking error. Dominance requires higher CAGR and lower TE than the incumbent.
VBIAX, with over $50bn in AUM, forms the first case. The authors simulate one G3M using VTSAX/VBTLX and another using VTI/BND from January 2, 2014 to June 30, 2026. EQL is compared with a pool of the 11 Sector SPDRs from June 19, 2018 to May 29, 2026. EDOW is set against a pool holding the 30 DJIA components from November 11, 2024 to June 29, 2026. According to the paper, choosing that window removes concerns about changes in the components.
EQL and EDOW each receive two tracking-error benchmarks. The first is the fund's legal index. The second rebalances daily to the positions implied by its economic description. VBIAX has one target because it combines two separate indices. Against NAV and the economic mandate, EQL dominates for gamma in [3.22%, 7.09%], while EDOW dominates in [3.32%, 9.90%]. VBIAX wins in [2.73%, 3.90%] when TE is measured monthly. Using market value, the EQL range begins at 3.56% and the EDOW range at 3.44%. Each continues through 10%, the sampled grid's upper limit.
Three of the ten summary-table rows have no dominating region.
The goalposts move in every winning row
All seven successful rows make one consequential choice. They measure the incumbent at market price instead of NAV, or substitute the continuously rebalanced economic allocation for the legal index. Under NAV and the legal mandate, both EQL and EDOW produce empty regions. The authors explicitly call NAV the closer counterpart to the G3M's NAV-based measurement.
VBIAX turns on the remaining convention. Daily TE gives an empty result; monthly TE produces dominance in [2.73%, 3.90%]. The paper attributes this reversal to negative autocorrelation in daily active returns, which overstates TE for every fund. The explanation is plausible and symmetric. Even so, the judgment on a $50bn incumbent changes with sampling frequency.
The benchmark problem is acknowledged directly. The authors call the comparison "of course, not immediate" and treat the legal index and continuously maintained target-weight portfolio as separate readings of one mandate. An EQL investor bought the legal index. Against that mandate, the pool beats EQL at market price while failing to beat its NAV. Those results have to travel together.
Fee size creates another obstacle. Dominance appears between 2.73% and 10%. By my own arithmetic, those fees are one to two orders of magnitude above current AMM quotes. The paper confines its claim to saying that the successful ranges are high relative to fees typically offered by AMMs, then treats them as rebalancing-band widths instead of exchange spreads. I accept that interpretation. Deployability becomes harder to claim. The ranges also come from the same sweep used to judge performance, and I did not find an out-of-sample or cross-fund check.
The differences under dispute are small. In the VBIAX monthly panel, CAGR spans 8.7% to 9.2% while TE runs from 0 to 50bp. CAGR and TE are the paper's only two metrics. It reports no Sharpe, volatility, drawdown or confidence interval. The legal-index panel for EDOW covers 10.0% to 13.5% of CAGR over 1.6 years. Its economic-mandate panel extends roughly from 10% to 14%.
The most useful concession concerns the market-making leg, defined as fees minus loss-versus-rebalancing. LVR is the value removed from the pool by arbitrageurs. Because the simulation contains only arbitrage flow, this leg is "non-positive by construction". Uninformed flow would add fee income. A footnote says realized LVR net of fees, within the dominating ranges, annualizes to the same order of magnitude as the incumbent fund's expense ratio. It supplies no figure. Reading LVR as the execution cost of mandate maintenance, in addition to adverse selection, is the paper's real conceptual move.
Our ETF adaptation
We cannot hold VTSAX, VBTLX, or the CRSPTMT and LBUFTRUU index series used in the paper, so our test substitutes tradable instruments. The 60/40 leg uses VTI and BND against a daily-rebalanced VTI/BND basket. This exercise adapts the band rule; it does not replicate the paper, and every figure below is ours.
At each close, the rule trades only after a weight breaches [(1-gamma)w_i, gamma + (1-gamma)w_i], then restores that weight to the nearest edge. Gamma runs from 0 to 0.100 in 0.001 steps. The sample covers January 2015 to June 2025 and contains 2,638 daily returns. We charged four tenths of a cent per share and assumed no slippage.
Our annual return was 8.42%, a little below the paper's VBIAX-panel range of 8.7% to 9.2%. The comparison is not like-for-like. That panel contains two simulated pools, one based on the VTSAX/VBTLX mutual funds and the other on the VTI/BND ETFs. Both use Bloomberg total-return data from January 2014 to June 2026. The ETF pool is closer to our setup, and the authors say it consistently tracks worse than the mutual-fund pool.
Our sample uses ETF closes over a shorter period. Its discrete restore-to-band-edge rule also allocates rebalancing proceeds differently from the paper's competitive-arbitrage simulation. The paper reinvests all income into the paying asset. Our BND series may contain prices only, which would depress the bond sleeve and pull our CAGR below the authors' result.
Our risk results fail a basic plausibility check. Annualized volatility reached 18.0%, and maximum drawdown was 42.1%. A long-only 60/40 VTI/BND portfolio did not behave that way from 2015 to 2025. An 8.42% annual return with 18.0% volatility gives a Sharpe of 0.71. The Calmar ratio against that drawdown is 0.20. Those figures indicate that the book did not remain a 60/40 portfolio.
We can identify two possible causes. The traded-symbol list includes a third ticker, EQL, beside VTI and BND, and we have not established whether it entered executed holdings. If it did, equity exposure would rise above 60%, which is the direction required to produce 18.0% volatility. (This EQL is a stray ticker in our implementation and has nothing to do with the paper's EQL case study.) The reported total return is also the simple sum of daily returns, 133.03%. Annualization may therefore overstate dispersion relative to a compounded curve.
The run records 5,278 trades across 2,638 days. It reports neither weight turnover nor tracking error by gamma. Consequently, our test leaves the paper's headline claim, joint dominance on CAGR and TE within a fee band, untested in both directions. We cannot fully account for the discrepancy using the information available. The causes we can identify belong to our implementation, not theirs.
The authors did compare NAV with the legal index for both EQL and EDOW. Each produced an empty dominance region. Evidence of dominance there, using a fund whose fee range had not been fitted through the same sweep, would change my mind about the fund comparison.
For now, the theorem carries the paper. A publicly verifiable ex ante band for mis-weighting is a genuine product design, and it pays the rebalancing agent from the pool instead of an expense ratio. Beating a large balanced index fund remains a different proposition. Across the ten rows of the summary table, the paper's answer changes with the choice among three measurement conventions.