Google is facing internal scepticism over the performance of its latest AI model, Gemini 4 Argon, with some employees questioning whether its benchmark results accurately reflect how well it performs in real-world use, particularly on coding tasks.
People familiar with the model’s internal evaluations told Bloomberg that while Gemini 4 has performed strongly on widely-used benchmarks, some employees have found it inconsistent when handling practical coding tasks. Concerns have also been raised about its capabilities in front-end development.
The views inside Google are not uniform. Some employees believe competing models from Anthropic and OpenAI are improving at a faster pace and could continue to outperform Gemini in certain areas. Others believe Gemini 4 has caught up with the leading models and is operating at the frontier of AI development.
A Google employee familiar with the model’s development reportedly said there was broad internal agreement that Gemini 4 was at the frontier and rejected suggestions that it struggled with complex, real-world coding tasks. The employee said Google had conducted extensive testing of the model.
The disagreement centres partly on the gap between benchmark performance and practical usability. People familiar with the model said some of the internal frustration stems from concerns that models can be optimised to perform well on standardised tests without necessarily delivering the same results when employees use them for more complex, less structured tasks.
The issue is particularly significant for Google employees working on AI development as the company seeks to compete with OpenAI and Anthropic. The company had previously planned to release Gemini 3.5 Pro but abandoned the model after delaying its planned launch, according to people familiar with the matter.
Google has disputed the suggestion that Gemini 4 is underperforming and pointed to positive assessments from its AI leadership.
The differing assessments highlight an internal debate over how AI models should be evaluated — whether benchmark scores are sufficient indicators of progress or whether performance in everyday, real-world applications should carry greater weight.

