Google Staff Say Gemini 4 Aces Benchmarks, Stumbles on Coding
TL;DR
- Bloomberg reports Gemini 4 performs well on industry benchmarks but employees say it stumbles on real coding, especially front-end design.
- Some Google staff believe Anthropic's Fable and OpenAI's Astra are improving faster than Gemini; others say Gemini 4 has caught up.
- Google disputes the characterization; Alphabet shares pared gains from up over 2% to up 0.5% after the report.
Google's newest flagship AI model performs well on the industry benchmarks used to gauge model efficacy, but does less well when the company's own engineers put it to work on real tasks, Bloomberg reported, citing people with direct access to the effort.
The specific weakness employees flagged: Gemini "isn't particularly adept at front-end design, which shapes how apps and websites look and feel." Some staff also believe Anthropic's Fable and OpenAI's Astra "are improving at a faster rate than Gemini," and that even at its best Gemini 4 will lag those rivals in some areas. Other people inside Google told Bloomberg the opposite, saying the coming version has caught up with the leading AI labs.
Google pushed back. The company said it would be "inaccurate to say that Gemini 4 is underperforming in areas such as coding," and a Google employee familiar with model development told Bloomberg there is "large consensus" internally that "Gemini 4 is at the frontier." Koray Kavukcuoglu, head of Google DeepMind, was cited saying: "I have the utmost trust in the team. In my mind, it's a certainty that we are always gonna be at the frontier."
Investors did not wait for the argument to settle. Alphabet shares pared their gains from up more than 2% to up 0.5% after the story landed.
The money at stake is not small. Training runs for models of this class can cost as much as $400 million, per Bloomberg. Google previously scrapped a planned Gemini 3.5 Pro release that had been announced at its May I/O conference, which Bloomberg said likely cost the company "dearly in time and money." The report lands the same day as our coverage of Google rolling Gemini 4 Argon out to cyber defenders first, one thread inside a run of 197 Google stories we've tracked over the last 90 days.
Shared on Bluesky by 1 AI expert
Originally reported by Bloomberg
Read the original article →Original headline: Bloomberg: Google Employees Say Gemini 4 Aces Benchmarks but Struggles With Real Coding, Alphabet Stock Slips