Microsoft has introduced two new in-house AI models: MAI-Image-2.5-Pro and MAI-Voice-2-Flash. This move shows the company’s effort to strengthen its proprietary AI ecosystem while improving performance, reducing infrastructure costs, and expanding the use of its own models across Microsoft products. The latest models build on the MAI lineup introduced earlier this year and are already being deployed across services such as Bing Image Creator, PowerPoint, OneDrive, Dynamics 365 Contact Center, and Azure. According to Microsoft, the new releases focus on delivering better quality, faster performance, and improved efficiency for both image generation and voice applications.
What MAI-Image-2.5-Pro and MAI-Voice-2-Flash Do
Microsoft has positioned MAI-Image-2.5-Pro as its highest quality image generation model to date. The model is designed to handle both image creation and image editing while offering greater accuracy and creative control.
MAI-Voice-2-Flash launches today! Flash is 2x faster than MAI-Voice-2 and 32% cheaper, at $15 per 1M characters.
— Mustafa Suleyman (@mustafasuleyman) July 23, 2026
MAI-Voice-2-Flash is also in public preview and powers Dynamics 365 Contact Center, our enterprise platform for call center agents, and reduces GPU costs up to 89%.…
Some of the key capabilities of MAI-Image-2.5-Pro include:
- High quality text to image generation
- More precise image editing using natural language prompts
- Improved rendering of text within images
- Identity preserving face editing
- Support for detailed and design ready visuals
The company has already integrated the model into several of its products. Bing Image Creator is now powered entirely by MAI-Image-2.5, making it Microsoft’s fully in-house image generation solution. PowerPoint uses the model for image to image editing, reducing GPU costs by up to 84% compared with GPT-Image-2. In OneDrive, MAI-Image-2.5 has become the default model for key image editing workflows. Microsoft says it has increased save rates by 26%, reduced P95 latency by approximately 25%, and delivered 2.5x greater efficiency under medium utilization production workloads.
Along with the image model, Microsoft also introduced MAI-Voice-2-Flash, which is a lightweight version of its voice generation model aimed at enterprise use cases which require low latency and lower operational costs. According to Microsoft, MAI-Voice-2-Flash offers:
- Low latency voice generation
- Natural sounding speech
- Lower infrastructure costs
- Faster inference for real time conversations
- Support for enterprise voice applications
It powers Dynamics 365 Contact Center and Azure Voice Live, which helps businesses build AI voice agents while cutting GPU costs up to 89%. Microsoft also points out that MAI-Voice-2-Flash matches the natural quality of MAI-Voice-2 but runs twice as fast and at 32% lower cost.
What Has Improved Compared With Earlier MAI Models?
The latest releases are an evolution of Microsoft’s earlier MAI models, including MAI-Image-2.5, MAI-Image-2.5-Flash, and MAI-Code-1-Flash, with more focus on production deployment and operational efficiency.
Better image generation and editing:
The earlier MAI-Image-2.5 model stood out for faster, quality image generation and better creative support. It featured more than 2x faster image generation than the previous generation, natural lighting and accurate skin tones, better texture quality, clear in image text, and competitive pricing for developers. With MAI-Image-2.5-Pro, Microsoft has sharpened edit accuracy, prompt understanding, identity aware changes, and text rendering.
Greater production efficiency:
Compared to the first MAI-Image-2.5 rollout, Microsoft now shares new performance data. There is a 26% increase in OneDrive save rates, approximately 25% reduction in P95 latency, 2.5x greater production efficiency, and GPU cost reduction of up to 84% in PowerPoint image editing. These improvements reflect Microsoft’s emphasis on lowering infrastructure costs while maintaining image quality across its own services.
Faster and more affordable voice generation:
MAI-Voice-2 already brought high quality, multi language speech generation, voice cloning from short samples, and security checks for safe use. The new MAI-Voice-2-Flash builds on that foundation by prioritizing speed and cost efficiency. According to Microsoft, the Flash variant runs twice as fast, costs 32% less, delivers the same natural speech quality as MAI-Voice-2, and reduces GPU costs by up to 89% in production deployments. These improvements are aimed at businesses deploying AI powered customer support and conversational voice agents at scale.
Continued expansion of Microsoft’s in-house AI ecosystem:
These launches follow the recent debut of MAI-Code-1-Flash, which is a coding model fine tuned for GitHub Copilot and Visual Studio Code. Paired with MAI-Image, MAI-Voice, and MAI-Transcribe, Microsoft is building out a suite of specialized AI tools for different jobs, rather than trying to stretch one general model everywhere. With MAI-Image-2.5-Pro and MAI-Voice-2-Flash now entering public preview, Microsoft is further integrating its proprietary AI technology across consumer and enterprise products while focusing on improving the quality, speed, and operational efficiency.









