Key highlights:
- Google started releasing Gemini 4 Argon to a small group of trusted cybersecurity partners with plans to allow paid subscribers access later.
- This new model came after Google dropped its planned Gemini 3.5 Pro which was supposed to launch in June.
- Google expects Gemini 4 to work with its main products, while competitors like OpenAI and Anthropic are quickly building their own AI tools.
- The new model is reportedly very large, which has some people wondering how much it will cost to run on a big scale.
Google just reached another stage in the race to create top AI models. The company rolled out Gemini 4 Argon to a limited group of trusted partners in cybersecurity on Wednesday. Once the product is thoroughly tested, the company will open access to more users, and paid subscribers should be next. This move comes as Google looks to compete more strongly with Anthropic and OpenAI.
According to the announcement, Google has introduced Gemini 4 Argon for complex, deep reasoning and long-horizon workflows across software engineering, enterprise knowledge work, and cybersecurity defense.
Moreover, Google has also been revising its AI plans. They ended up abandoning Gemini 3.5 Pro, which they originally said would launch in June. Instead, Gemini 4 is now set to handle much of Google’s AI strategy and power its biggest products.
Strong Benchmarks but Questions about Regular Use
The real issue around Gemini 4 is the gap between test scores and actual performance in day to day tasks. Several people with knowledge of the model’s internal evaluations told Bloomberg its coding skills are inconsistent. They said it can have a hard time handling some real coding jobs, even though it does well on standard tests. This concern has started a discussion about “benchmaxxing”. That is when AI developers focus too much on getting good test results, instead of making the system better for real work. Edwin Chen, who started the AI company Surge AI, compared it to a student who gets a top SAT score but does not always do well when faced with real life problems.
This issue is especially noticeable with coding. A model can be trained to pass a specific programming test but it might still run into problems building or changing an entire app. Bloomberg’s sources said Gemini 4 has weaknesses with front end design, an area that is growing more important as Google competes with other companies making AI coding tools. Still, not everyone at Google is worried about these issues. One employee with knowledge of the project told Bloomberg that most people at the company believe Gemini 4 is a leader. This person also said the model does not have trouble with complex, real world coding.
Gemini 4 Argon is our next era of frontier intelligence.
— Google (@Google) September 30, 2026
It shows significant improvements across benchmarks, setting a new state of the art for real-world long-horizon software engineering tasks. pic.twitter.com/wQzNpp7Dx4
Google’s Bigger Gemini Challenge
The debate over Gemini 4 comes after a difficult period for Google’s AI development. The company announced Gemini 3.5 Pro at its I/O conference in May and said it would release it in June, but that never happened. People familiar with the situation said Google ended up abandoning the model. Training advanced AI is also extremely expensive. Bloomberg Intelligence analyst Mandeep Singh estimated just training a model like this once costs up to $400 million and highly paid researchers make it even pricier.
Google now depends heavily on Gemini 4 for its main products. Gemini powers AI features in Search, Maps, Gmail, and Chrome, all with over a billion users. The company said its chatbot app for consumers and the AI Mode in Search have each crossed 1 billion users as well. Meanwhile, Google faces growing competition from OpenAI and Anthropic, who are both building their own products instead of just selling AI models. These competitors are also making coding agents and systems that can handle tasks on their own.
Despite the criticism, Gemini 4 does have some strong points. One person familiar with the model said it is especially good at understanding more than just text like pulling metadata from video. The same source highlighted Gemini 4’s safety, cybersecurity features, and natural communication style. Google states that Gemini 4 supports an output limit of up to 1 million tokens at once. The company describes it as being built for tough, lengthy work in software engineering, finance, law, and cybersecurity.
The rollout starts with a small group of trusted cybersecurity partners, wider access coming later after more testing. Paid subscribers are expected to get access as Google expands availability. Inside Google, opinions are still split. Some employees think Anthropic’s Fable and OpenAI’s Astra are improving faster, and that Gemini 4 could still fall behind them in some areas. Others believe Gemini 4 has caught up with the top AI labs. Publicly, Google says the model is doing well but inside the company, the main question is whether good test scores actually translate into real world usefulness.









