I Asked Gemini, ChatGPT, and Claude to Backtest Three Simple Investment Strategies. The Differences Are…Eye-Opening.
Most people compare portfolios.
I decided to compare AIs.
Specifically, I asked Google Gemini 3, ChatGPT 5.1, and Claude 4.5 the same question:
**"Backtest three investment strategies over the last 30 years with monthly contributions:
One fund: Vanguard LifeStrategy Growth
Two funds: VTSAX + VBTLX
Four funds: VTSAX + VTIAX + VBTLX + DJP (commodities)"**I expected small differences in wording.
Instead, I got three completely different personalities, workflows, and results.
Here’s how each AI handled the exact same task — and what it reveals about how you should use them for your own financial research.
1. Gemini 3 — Polished, Confident… and a Little Too Confident
Gemini behaves like Google: clean UX, pretty charts, and instant answers.
Here’s the interactive tool Gemini gave me:
Gemini delivered:
- Automatic CAGR tables
- A beautiful “Portfolio Growth Simulator”
- Allocation summaries
- Simple pros and cons
But here’s the problem:
Gemini never explains where the numbers came from.
No citations.
No assumptions.
No acknowledgment that many of these funds didn’t exist 30 years ago.
No discussion about synthetic data vs real data.
Gemini simply acts like it has the data — then presents results with extreme polish and confidence.
It’s great for beginners.
But if you need accuracy, transparency, or actual methodology?
You need something more rigorous.
2. ChatGPT 5.1 — The Analyst Who Tells You the Truth
ChatGPT took a completely different approach.
Instead of pretending everything was straightforward, it began with the most honest line of all:
“We can’t literally backtest all 3 for 30 years because several funds didn’t exist yet.”
This is what an analyst would say — not a marketer.
ChatGPT then built a model using:
- VASGX since-inception return
- Long-run S&P 500 return as proxy for VTSAX
- Bloomberg Agg for VBTLX
- MSCI ACWI ex-US returns for VTIAX
- DJP’s actual return since 2006
Then it walked me through:
- The formulas
- The assumptions
- The derived weighted returns
- The future-value calculation for monthly investing
This is what its response looked like (formatted as a screenshot-style block):
ChatGPT’s Key Output
---------------------Final Values (assuming $1,000/month for 30 years):• LifeStrategy Growth (≈8.5%/yr): ~$1.65M
• Two-Fund 60/40 (≈7.7%/yr): ~$1.40M
• Four-Fund with Commodities (≈7.2%/yr): ~$1.27M Interpretation:
• Commodities dragged long-term returns
• International stocks underperformed the US
• LifeStrategy’s higher stock allocation explains its higher outcome
ChatGPT gives you a thinking partner — not just an answer.
If you’re running a finance channel or writing an analysis, ChatGPT is the most trustworthy for method transparency.
3. Claude 4.5 — The Quant Intern Who Shows You Every Step
Claude takes things to another level.
It didn’t just answer the question.
It tried to run an actual backtest.
Literally.
Claude wrote Python code, attempted to download historical data, hit network restrictions, and showed me the errors.
Then it pivoted — pulled historical performance numbers manually — and built a 30-year synthetic backtest with:
- CAGR
- Crisis-year performance
- Expense ratios
- Drawdowns
- Final values
- A CSV export
- A Markdown report
- A one-page summary
- And this full chart:
📊 Claude’s Backtest Chart
Claude’s results were:
StrategyFinal ValueCAGRTwo-Fund (70/30)$738,9024.82%LifeStrategy Growth$628,9654.26%Four-Fund (with commodities)$522,1253.61%
And Claude even added human-like commentary:
“Strategy 2 outperfored by $109,937 vs Strategy 1 and $216,776 vs Strategy 3.”
Claude is the AI to use if you want research reports, automation, code, and deliverables — not just an answer.
What the Three AIs Agree On
Despite the different workflows, tones, and confidence levels, all three models agreed on one big takeaway:
Simple, low-cost U.S.-heavy portfolios dominated from 1995–2024.
Specifically:
- The 4-fund portfolio underperformed because commodities and international stocks lagged
- The “extra diversification” didn’t compensate for lower long-term returns
- The real differences came from:
- U.S. market dominance
- Expense ratios
- Stock/bond mix
- Staying invested through crises
This is a great lesson for investors age 50–65:
Complexity rarely beats simplicity.
Costs matter.
Behavior matters more.
Which AI Should You Trust for Financial Research?
Here’s the simple answer:
Use Gemini for…
- Quick summaries
- Visual explanations
- Beginner-level investing content
Use ChatGPT for…
- Transparent methodology
- Balanced financial analysis
- Explaining tradeoffs
Use Claude for…
- Deep research
- Coding
- Simulations
- Full reports with charts & CSVs
- Anything that feels like “quant work”
Final Thoughts
The real value of AI in personal finance isn’t to “predict the future.”
It’s to help you understand your decisions better today.
All three AIs gave different answers —
but the lesson was the same:
Low cost beats complexity.
Simplicity beats fancy.
Consistency beats timing.
