Saturn tested the most popular offerings from OpenAI’s ChatGPT, Anthropic’s Claude, Microsoft’s Copilot, xAI’s Grok and Google’s Gemini and found accurate responses only 43% of the time on average Read more: https://t.co/779jcPa7lT
Why now
Thesis: AI reliability gap challenges hype-driven valuations
Catalyst to watch: Follow-up benchmarks or vendor rebuttals
Main risk: Single test may not reflect real-world use
Why it matters
Low accuracy across major AI models questions near-term enterprise reliability and monetization.
Details
Saturn tested the most popular offerings from OpenAI’s ChatGPT, Anthropic’s Claude, Microsoft’s Copilot, xAI’s Grok and Google’s Gemini and found accurate responses only 43% of the time on average Read more: https://t.co/779jcPa7lT
Sources
- x · 2026-09-25 03:57 UTC