Models & Releases
New models, versions, and capabilities shipping across the frontier and open-model labs.
New models, versions, and capabilities shipping across the frontier and open-model labs.
OpenAI is retiring Priority Processing in favor of a new API tier called Fast mode, announced in the same July 30 update that cut prices on GPT-5.6 Luna and Terra. For GPT-5.6 Sol, the company's top…
Midjourney released its V8.2 image model on July 24, 2026, focused on aesthetic quality, reducing low-quality outputs, and sharpening how the tool learns individual users' taste. What's new…
xAI has expanded Grok Imagine Video 1.5 with reference-based video generation, letting users lock a specific face, product, or location into a generated clip, alongside a jump to native 1080p output…
Google has added direct Chrome integration to Gemini Spark, its autonomous personal AI agent, letting the agent use a person's logged-in accounts and saved passwords to complete web-based errands on…
DeepSeek moved DeepSeek-V4-Flash out of preview and into an official public beta on July 31, re-post-training the same architecture to deliver what the company describes as significantly enhanced…
NVIDIA expanded its Agent Toolkit on July 26, 2026 to include NVIDIA PhysicsNeMo and CUDA-X libraries as agent-callable tools, letting AI agents invoke physics simulations and accelerated solvers…
Google DeepMind introduced Gemini Robotics 2 on July 30, 2026, describing it as "the intelligence layer powering the next generation of truly adaptable robots." The release pairs a new…
xAI has released Grok Voice Think Fast 2.0, the successor to its real-time voice model, and will automatically migrate all traffic pointed at the grok-voice-latest alias to the new version starting…
Google DeepMind has released Lyria 3.5, the latest version of its music-generation model, rolling it out today inside Google Flow Music. The update focuses on four areas of the music-creation…
Google has added a new voice-driven natural-language mode to the Gemini app for macOS, letting users dictate, summarize, and rewrite text in any desktop application by holding a single key. What's…
AWS announced on July 28, 2026 that AgentCore Gateway now supports the 2026-07-28 revision of the Model Context Protocol (MCP), which the post describes as the protocol's biggest change since it…
xAI introduced Build Mode on July 28, 2026, a new feature inside Grok that turns a plain-language description into a working, publishable app, game, website, or dashboard without the user writing or…
OpenAI added two new speech-to-text models to its API on July 28, 2026: GPT Transcribe for file-based transcription and GPT Live Transcribe for low-latency streaming transcription, rounding out a…
Google has moved Gemini 3.6 Flash and Gemini 3.5 Flash-Lite out of preview and into general availability on the Gemini API as of July 21, 2026, while simultaneously deprecating the classic sampling…
Cursor has launched Router, a feature that automatically picks which AI model handles each coding request instead of leaving that choice to the developer, aiming to cut costs without giving up…
Black Forest Labs introduced FLUX 3, a multimodal foundation model the company says jointly learns from images, video, and audio inside one architecture rather than stitching together separate…
Anthropic released Claude Opus 5 on July 24, 2026, the newest version of its mid-tier flagship reasoning model, available immediately at the same price as its predecessor. What's new According to…
xAI updated its Speech to Text API with an adjustable voice-activity-detection threshold, giving developers finer control over how the model distinguishes speech from silence or background noise.…
OpenAI has begun rolling out Health in ChatGPT to logged-in US users 18 and older on web and iOS, across the Free, Go, Plus, and Pro plans. The feature lets people securely connect Apple Health and…
Anthropic launched a Claude connector for its Economic Index on July 22, 2026, letting anyone query the organization's dataset on real-world AI usage directly through conversation instead of digging…
AWS has published details on Agentic Retrieval, a new capability for Amazon Bedrock Managed Knowledge Base that lets a foundation model decompose complex, multi-part questions into sub-queries,…
Anthropic shipped a batch of Claude Managed Agents API updates on July 22, giving developers finer control over agent cost, faster session startup, and webhook-driven visibility into agent…
AWS has published a detailed account of how monday.com runs "AI Teammates" — agentic AI built on Amazon Bedrock — in production across a decade-old codebase, reporting that nine in ten of its…
Google has deprecated the temperature, topp, and topk sampling parameters on its newest Gemini models, telling developers to strip them from every request now or risk outright failures once the…
Google shipped two new members of its Gemini Flash line on July 21, 2026 — Gemini 3.6 Flash and Gemini 3.5 Flash-Lite — alongside a specialized cybersecurity variant, Gemini 3.5 Flash Cyber,…
Anthropic published a case-study blog post on July 16, 2026 walking through how large codebase migrations can be run end-to-end with Claude Code and agentic loops, headlined by a roughly…
Anthropic has shipped mid-conversation system messages for the Claude API, a feature that lets developers inject new operator-level instructions partway through a conversation without invalidating…
NVIDIA has introduced Cosmos 3 Edge, a version of its Cosmos world-model family built to run reasoning directly on edge hardware rather than in the cloud, alongside a wider push into Japan's…
Google's next flagship model, Gemini 3.5 Pro, is running months behind its internal schedule, according to Bloomberg, which cited people familiar with the matter. The company has been holding the…
Google added two new capabilities to Google Vids on July 16, 2026: Gemini Omni, a natural-language video editing tool, and personal avatars, which let users generate talking-head video clips of…
1Password and Anthropic have shipped 1Password for Claude, a new integration that lets Claude fill in stored logins during browser tasks without the password or one-time code ever entering the…
Together AI is now serving Inkling, a new multimodal mixture-of-experts model from Thinking Machines Lab, on its inference platform starting the day of the model's release. The launch gives…
Google has begun rolling out Gemini in Chrome to desktop users in the United Kingdom, extending the AI browsing assistant beyond its earlier North America, Asia-Pacific, Latin America, Africa, and…
Google rolled out a set of Gemini-powered features to Waze on July 13, letting drivers search for destinations and report road conditions by talking naturally to the app instead of typing or picking…
Snowflake announced Cortex Sense on June 30, 2026, a system designed to let enterprise AI agents accurately query business data that was never formally modeled into semantic views, with private…
Suno shipped natural-language lyric editing on its web app on July 9, 2026, letting users revise AI-generated song lyrics by describing the change they want rather than manually rewriting text.…
Mistral has added version control for prompts and agent skills inside Studio, aiming to fix a governance gap it says most enterprises don't realize they have. What's new Mistral frames the problem…
OpenAI has shipped the GPT-5.6 model family, moving the line from preview to general availability with three tiers built for different cost and latency budgets: GPT-5.6 Sol for frontier capability,…
Mistral AI has introduced Robostral Navigate, an 8-billion-parameter model that lets robots navigate physical spaces using a single ordinary camera and plain-language instructions, with no depth…
xAI has made Grok 4.5 generally available on the xAI API, positioning the model as its new option for coding, agentic tasks, and knowledge work, according to the company's own release notes. What's…
Meta released Muse Image on July 7, the first publicly shipped model from Meta Superintelligence Labs, rolling it out inside the Meta AI app, Instagram Stories in the US, and WhatsApp in limited…
xAI has released 21 new flagship voices for Grok Voice, more than quadrupling its lineup, and paired the launch with a new no-code Voice Agent Builder for assembling custom voice agents. What's new…
Google has moved its Interactions API to general availability, positioning it as the new default way to build with Gemini models and agents in place of the long-standing generateContent endpoint.…
OpenAI shipped two new realtime models on July 6, 2026 — GPT-Realtime-2.1 and GPT-Realtime-2.1 mini — both available immediately through the v1/realtime endpoint, according to the official API…
Tencent has released Hy3, the general-availability version of its Hunyuan large language model, following up on the Hy3 Preview it launched on OpenRouter and Hugging Face in April. The model card…
X has launched a hosted Model Context Protocol (MCP) server, letting AI assistants and coding tools connect directly to X's platform and API without developers having to build or maintain their own…
Google will discontinue the consumer version of Gemini Code Assist on GitHub on July 17, 2026, according to its own documentation, redirecting developers to a separate enterprise version installed…
Anthropic has rolled out a set of admin tools for Claude Enterprise aimed at giving IT and finance teams more visibility into how their organizations use Claude and more control over what it costs.…
AWS has published details on a new detection approach in Amazon Bedrock designed to catch phishing emails written by generative AI, which increasingly slip past filters built around old-fashioned…
Google has moved gemini-3.1-flash-lite-image — nicknamed Nano Banana Lite — out of preview and into general availability, positioning it as its fastest, cheapest image generation and editing model in…
xAI has added a /goal mode to Grok Build, its agentic coding CLI, letting developers hand off a single objective and have the agent plan, execute, and verify the work autonomously instead of managing…
Mistral AI has shipped Leanstral 1.5, an updated version of its specialized model for automated theorem proving and autoformalization in Lean 4, the interactive proof assistant used for mechanically…
Anthropic released five platform updates to Claude Managed Agents on June 30, 2026, covering real-time event streaming, per-session configuration flexibility, and a more complete webhook system that…
Anthropic released Claude Sonnet 5 on June 30, 2026, as the next generation of its mid-tier Sonnet model family. Available immediately across all plans and via the Claude API under the model ID…
xAI shipped a set of API capability upgrades to its Grok inference platform in June 2026, including a new scheduling priority tier, expanded Files API integration with permanent URL generation, and…
Google released two new generative AI models on June 30, 2026: Gemini Omni Flash, a multimodal video generation and editing model priced at $0.10 per second of video output, and Nano Banana 2 Lite…
Cohere released North Mini Code 1.0 on June 9, 2026 — its first model purpose-built for agentic coding workflows. The model is a 30-billion parameter Mixture of Experts (MoE) architecture with only 3…
Google's Veo 2.0 and Veo 3.0 video generation models reached end-of-life on June 30, 2026, with access to the deprecated model IDs now cut off. Developers who have not yet migrated away from these…
Anthropic quietly completed the removal of fast mode support for Claude Opus 4.6 on June 29, 2026, following through on a deprecation notice it issued when Claude Opus 4.8 launched last month. The…
Google has announced the deprecation of three Imagen 4.0 image generation models in the Gemini API, with the shutdown scheduled for August 17, 2026. Developers using these model IDs in production…
OpenAI released the Safety Usage Dashboard to the OpenAI API platform on June 23, 2026, giving developers a new tool to observe and analyze blocked Responses API requests tied to specific end-user…
Google has scheduled the shutdown of its Veo 2.0 and Veo 3.0 video generation model IDs for June 30, 2026 — three days away — with users directed to migrate to Veo 3.1 preview or GA models. The…
Runway launched Seedance 2.0 Mini (seedance2mini) on its API on June 26, 2026, adding a new lighter-weight tier to its Seedance 2.0 model lineup. The mini variant targets fast generation from text,…
Anthropic announced on June 25, 2026 that fast mode for Claude Opus 4.7 is deprecated, with a hard removal date of July 24, 2026. Developers using the claude-opus-4-7 model with speed: "fast" have…
OpenAI announced a preview of GPT-5.6 Sol on June 26, 2026, describing it as a next-generation model with stronger capabilities in coding, science, and cybersecurity — and pairing it with what OpenAI…
Notion published a case study on June 25 showing how it embedded Cursor's agentic coding engine directly into its productivity workspace using the Cursor SDK. Users can now tag Cursor in Notion…
Mistral AI has released a round of enterprise governance features for its Connectors system — the layer that lets agents and workflows connect to external platforms like Google Drive, Slack, and…
Google on June 24 announced that computer use — the ability for an AI model to see, reason about, and take action across digital interfaces — is now a built-in capability in Gemini 3.5 Flash.…
Anthropic released Claude for Foundation Models in beta on June 9, 2026—a Swift package that integrates Claude into Apple's Foundation Models framework for developers building on iOS 27, macOS 27,…
Mistral AI released OCR 4 (mistral-ocr-4-0) on June 23, 2026, its latest document understanding model. The update adds structured spatial output — bounding boxes, content-type labels, and per-word…
Google on June 17 enabled streaming support for the gemini-3.1-flash-tts-preview model, allowing developers to receive speech audio as it is generated rather than waiting for a complete synthesis…
OpenAI on Monday expanded its Daybreak cybersecurity platform with three simultaneous additions: an updated Codex Security plugin, the full launch of the GPT-5.5-Cyber model for vetted defenders, and…
Meta announced Muse Spark on April 8, 2026, the first model from Meta Superintelligence Labs — a new AI research and product organization built with former Scale AI co-founder Alexandr Wang, in whom…
OpenAI has shipped a cluster of enterprise-facing Codex features across June 2026, highlighted by the general rollout of Computer Use for enterprise customers and a new Record & Replay capability for…
OpenAI launched three new real-time audio models on May 7, 2026 — GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper — while simultaneously moving the Realtime API from research preview…
Cohere on May 20, 2026 released Command A+ (model ID: command-a-plus-05-2026), a 218-billion-parameter mixture-of-experts model under the Apache 2.0 license, targeting production enterprise…
Amazon Web Services has launched Web Search on Amazon Bedrock AgentCore as a generally available feature, giving AI agents access to a continuously updated proprietary web index without routing…
OpenAI announced on June 18 that GPT-5.5 Instant — the model now powering health responses for all free ChatGPT users — outperformed physician-written answers on a set of 3,500 real-world health…
Google announced on June 15 the deprecation of six media generation models across its Gemini API — all three Imagen 4 image generation variants and all three Veo video generation models in the 2.0…
xAI's Grok 4.3 became available on Amazon Bedrock on June 15, 2026, making the reasoning-first model accessible to enterprise developers through AWS's managed AI infrastructure. The launch gives Grok…
Z.ai released GLM-5.2 on June 16, 2026, an open-source model with a 1M-token lossless context window and benchmark scores that place it at the front of available open-source coding models. The model…
Amazon has updated SageMaker AI Async Inference to accept raw inference payloads directly in the request body, removing the long-standing requirement to stage input data in Amazon S3 before…
Google updated the Gemini API on June 17, 2026 to support streaming for text-to-speech output via the gemini-3.1-flash-tts-preview model. Developers can now receive audio chunks as they are generated…
Google announced on June 15, 2026 the deprecation of its Imagen 4 image generation family and Veo 2/3 video generation models from the Gemini API. Veo users face a June 30 deadline — less than two…
Amazon Web Services has added container caching to Amazon SageMaker AI, removing the container image download step from the scale-out critical path and reducing model startup latency by 51% in tested…
Amazon Web Services has made P-EAGLE, a parallelized speculative decoding technique, available through Amazon SageMaker JumpStart. The method accelerates large language model inference by generating…
Meta introduced AI Mode on Facebook on June 16, 2026, a new search capability that replaces traditional link-based results with AI-generated answers drawn from public conversations happening across…
Cursor shipped a significant update to Bugbot on June 10, 2026, making its automated code review agent more than three times faster, 22% cheaper to run, and 10% more effective at catching bugs. The…
xAI has shipped a series of updates to its Grok Imagine API in June 2026, adding deep integration with the Files API, public URL sharing for stored assets, and a new image-to-video model — Grok…
Databricks completed its native AI SQL function suite in June 2026, promoting aiextract, aiclassify, and aiquery to general availability in a series of releases between June 11 and June 15 — giving…
Google announced on June 15 that it is deprecating all three Imagen 4 image generation models and three Veo video generation models from the Gemini API. The Imagen 4 family shuts down August 17,…
Anthropic retired two original Claude 4 models — Claude Sonnet 4 and Claude Opus 4 — effective June 15, 2026. All API requests to the model IDs claude-sonnet-4-20250514 and claude-opus-4-20250514 now…
Google released Gemini 3.5 Flash as a generally available model on May 19, 2026, and simultaneously launched Managed Agents in the Gemini API in public preview — a hosted agent execution service that…
Anthropic has shipped a /fork command to Claude Code, enabling developers to branch an AI coding session into a parallel subagent that inherits the full context of the existing conversation — a…
Midjourney set V8.1 as its default model on June 11, 2026, replacing V7 after user testing and feedback. The transition marks V8.1's general availability as the standard experience for all Midjourney…
Google formally ended API access to all four Gemini 2.0 Flash models on June 1, 2026. Requests to gemini-2.0-flash, gemini-2.0-flash-001, gemini-2.0-flash-lite, and gemini-2.0-flash-lite-001 now…
xAI updated its developer platform on June 10, 2026, closing the loop between the Files API and the Imagine image and video generation endpoints. Three interlocking changes let applications…
Alongside the June 9, 2026, launch of Claude Fable 5 and Claude Mythos 5, Anthropic shipped a new beta API parameter called fallbacks that lets developers automatically reroute refused requests to a…
xAI launched Grok Build in May 2026, entering the agentic coding tool space alongside Anthropic's Claude Code, OpenAI's Codex CLI, and GitHub Copilot Workspace. The release packages two connected…
Google is integrating Gemini directly with Google Business Profile, giving small business owners an AI assistant that has access to their real operational data — customer reviews, search impressions,…
Anthropic announced on June 5, 2026 that Claude Sonnet 4 (claude-sonnet-4-20250514) and Claude Opus 4 (claude-opus-4-20250514) will be retired from the Claude API on June 15, 2026—five days from…
Anthropic has shipped three developer-facing additions to its Claude Managed Agents platform alongside the June 9 Claude Fable 5 launch. The most significant is scheduled deployments: developers can…
Anthropic on June 9, 2026, launched Claude Fable 5 and Claude Mythos 5, its most capable publicly released models to date. Fable 5 is immediately available to all users across Claude.ai and the API;…
OpenAI updated its Responses API on June 9, 2026, enabling web search to return image results alongside text. The change is live in the v1/responses endpoint and targets applications where a…
Google has launched Gemini 3.5 Live Translate, a speech-to-speech translation model that handles more than 70 languages and 2,000 language pair combinations in real time. The model enters public…
At London Tech Week on June 7, NVIDIA founder Jensen Huang and UK Prime Minister Keir Starmer declared that "the U.K. would be an AI maker, not an AI taker" — a commitment backed by a cluster of…
OpenAI on June 4, 2026 added moderation scoring to both the Responses API and the Chat Completions API, enabling developers to check the safety of prompts and model outputs in a single generation…
NVIDIA and LG Group announced on June 7, 2026 a broad collaboration to build an AI factory designed to accelerate LG's businesses across robotics, autonomous driving, data center infrastructure, and…
NVIDIA and Doosan Group announced an expanded collaboration on June 7, 2026, covering physical AI systems, industrial robotics, power infrastructure for AI factories, and substrate materials for AI…
Anthropic shipped two API changes on June 2, 2026 that reduce billing costs and improve control for developers running inference at scale. What's New No billing for zero-output refusals: On the…
Google has moved two of its native Gemini image generation models from preview to general availability on the Gemini API: gemini-3.1-flash-image and gemini-3-pro-image. The GA release also introduces…
ElevenLabs on May 28, 2026, released Dubbing v2, an AI dubbing model designed to preserve the emotional performance and vocal characteristics of the original speaker across more than 90 languages.…
ElevenLabs on May 26, 2026, released Music v2, the next version of its AI music generation model. The update adds genre-transition support, embedded sound effects, an improved multilingual engine,…
OpenAI on June 7, 2026 introduced Trusted Contact — an optional, opt-in safety feature in ChatGPT that can notify a designated contact if automated systems and trained human reviewers determine a…
OpenAI on June 7, 2026 published a detailed overview of three new realtime voice models now available in the API: GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper. The models advance…
Google moved two native visual generation models to general availability on May 28, 2026: Gemini 3.1 Flash Image (internally branded "Nano Banana 2") and Gemini 3 Pro Image ("Nano Banana Pro"). The…
Google retired four Gemini 2.0 API models on June 1, 2026, ending support for the Flash and Flash Lite variants that anchored the 2.0 generation. Developers still hitting these endpoints receive…
Amazon Web Services on June 4, 2026 added NVIDIA Nemotron 3 Ultra to SageMaker JumpStart, giving cloud developers one-click access to a 550-billion-parameter reasoning model designed specifically for…
Anthropic on June 2 shipped two developer-facing improvements to the Claude API: a new parameter for capping the advisor tool's output per call, and a billing change that eliminates charges on API…
Google used its annual I/O developer conference on May 19, 2026 to announce Gemini 3.5 Flash — a frontier-class model designed for fast, low-cost agentic inference — and Gemini Spark, a 24/7 personal…