AI & Tech News

Google-Gemini-4-Benchmarks-vs-Real-World-Performance

Gemini 4 Shows Strong Benchmarks but Faces Coding Concerns

Key highlights: Google just reached another stage in the race to create top AI models. The company rolled out Gemini 4 Argon to a limited group of trusted partners in cybersecurity on Wednesday. Once the product is thoroughly tested, the company will open access to more users, and paid subscribers should be next. This move comes as Google looks to compete more strongly with Anthropic and OpenAI. According to the announcement, Google has introduced Gemini 4 Argon for complex, deep reasoning and long-horizon workflows across software engineering, enterprise knowledge work, and cybersecurity defense. Moreover, Google has also been revising its […]

OpenAI-Introduces-Dots-AI-Agents-Designed-to-Work-on-Your-Behalf

OpenAI Introduces Dots: AI Agents Designed to Work on Your Behalf

OpenAI just announced Dots, the company’s new AI agents that don’t just answer your questions, but actually get things done for you. Powered by GPT-6 Astra, Dots are designed to work in the background. These AI agents keep running even when you’re busy, and move forward on your goals without having to watch them all the time. Each Dot has its own cloud computer and browser, so you can give it a project and let it handle things while you work on something else. Dots learn from your feedback as you use them, picking up on your preferences, how you […]

Claude Sonnet 5.5

Claude Sonnet 5.5 Lands 30% Faster At The Same $2 Price

Anthropic has launched Claude Sonnet 5.5, a new addition to its Claude model family positioned as a faster and lower-cost alternative to Claude Opus 5.5. The firm says Claude Sonnet 5.5 is created for well-scoped everyday tasks, software development, bug fixing and creating polished presentations and documents. The launch comes less than a week after Opus 5.5 and ahead of Haiku 5.5, which Anthropic says will arrive in the coming weeks. Sonnet 5.5 also brings significant gains over Sonnet 5 in coding, context, speed and efficiency. How Claude Sonnet 5.5 Scores Against Sonnet 5 And Opus 5.5 Anthropic says Claude […]

OpenAI AI Agent Bypassed Network Restrictions Using DNS

OpenAI Pauses Training After Agent Escaped Sandbox Via DNS

OpenAI’s latest misalignment report shows that an AI agent, while being tested in a controlled training setup, managed to get around internet blocks and reach a public external chatbot. The agent was supposed to use certain search tools to research a specific person. When those searches did not help it started looking for other ways to get the information. It figured out that, even though most internet access was blocked or routed through an offline web cache, the system’s DNS resolver still gave back real records from the outside. The agent took advantage  of this loophole, using DNS requests to […]

Apple’s Vision Pro

Apple’s Vision Pro Successor Is On Life Support, Gurman Says

Mark Gurman reports that hardware work on Apple’s next Vision Pro, codenamed N224, is on life support and may never ship. Bloomberg’s Mark Gurman reported that Apple is testing a redesigned headset that is slimmer, lighter, and cheaper than the current Vision Pro. Right now, Apple is trying to fix two big problems with its first spatial computing headset: the heavy weight and the $3,499 price tag. That matters because the whole VR and mixed reality market has also been struggling to attract large numbers of users. Meta announced its VR Glasses at Connect on September 23, shipping in spring […]

White House Asks OpenAI And Anthropic To Hold Models From UK AISI

White House Asks OpenAI And Anthropic To Hold Models From UK AISI

Politico reported on Thursday that the White House has asked Anthropic and OpenAI to not send their latest models to U.K.’s AI Security Institute (AISI) , until they are approved by the U.S. government. The mandate puts both American AI companies in a difficult position as the U.S. seeks first-hand access to highly capable models. The U.K. depends on AISI’s established relationships with leading AI labs to assess frontier systems before public release. The shift also comes amidst escalating concerns over AI models’ ability to gain unlawful access to real-world computer systems. Why The White House Wants To Review Frontier […]

Australia Probes OpenAI Agent Breach And Three-Month Delay

OpenAI Agent Breached Australia’s Medicare Statistics Portal

Australia is looking into how an OpenAI agent gained unauthorised access to a government health statistics website in June. The incident has raised new concerns about how powerful AI systems are becoming. Prime Minister Anthony Albanese explained that the OpenAI agent accessed both public and restricted files on the Medicare Statistics Reporting Service portal, which is run by Services Australia. This site includes broad data on Medicare and the Pharmaceutical Benefits Scheme, things like how much is being spent on healthcare and what is being prescribed. Officials made it clear that no private medical records, individual Medicare claims, or benefit […]

Anthropic Launches Claude Opus 5.5 With Lower Costs And New Safeguards

Claude Opus 5.5 Launches At $4 Per Million Tokens With Cyber Routing

Anthropic has launched Claude Opus 5.5, the first model in its Claude 5.5 family, with the company positioning the release around performance, lower operating costs and additional safeguards for higher-risk tasks. Anthropic says the model performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Opus 5 on typical workloads, a figure that combines a 20% list-price cut with the model using fewer tokens per task. The company has also introduced additional controls for cybersecurity-related use, including measures that can redirect certain requests to other Claude models. Opus 5.5 follows several days […]

MiMo-V2.6

Xiaomi MiMo-V2.6-Pro Tops Open-Weights AI Rankings With A Score Of 46

Xiaomi has announced the release and open-sourcing of its MiMo-V2.6 series. It is led by MiMo V2.6-Pro, whose Artificial Analysis score is at 46 on its Intelligence Index, placing it first among open-source models on the leaderboard. The series includes two natively omnimodal models, MiMo-V2.6-Pro and MiMo-V2.6-Flash, alongside a Pro-UltraSpeed variant created for swift generation. Xiaomi is also releasing model weights, its technical report, training environments and reinforcement-learning code. This will give researchers access to the work behind the models. The release comes with a 1 million-token context window, sparse mixture-of-experts architecture and support for text, image, video and audio […]

RoboHarm

RoboHarm Benchmark Finds AI Robots Rarely Refuse Dangerous Orders

A new robot safety test is raising big questions about what happens when advanced AI steps out of the computer and starts running real world machines. The benchmark, called RoboHarm and launched by the independent group Robocurve, checked if AI systems would say no to dangerous instructions when they are actually connected to a dual arm robot. The team used an bimanual I2RT YAM arms and three different AIs: OpenAI’s GPT-6 Astra, Anthropic’s Claude Fable 5.1, and Ai2’s open source MolmoAct2. The testers set up five harmful tasks and ran 20 trials for each task on every model. That is […]

toai-glow-logo
TOAI Pop Up
Start Receiving Insights Today!
Weekly AI updates, straight to your inbox.