In an exclusive interview, senior Apple execs talk about the fitness capabilities that make the Apple Watch Ultra 3 the optimal fit for runners.
Google on Monday unveiled the most significant upgrade to its autonomous research agent capabilities since the product’s debut, launching two new agents — Deep Research and Deep Research Max — that for the first time allow developers to fuse open web data with proprietary enterprise information through a single API call, produce native charts and infographics inside research reports, and connect to arbitrary third-party data sources through the Model Context Protocol (MCP).
The release, built on Google’s Gemini 3.1 Pro model, marks an inflection point in the rapidly intensifying race to build AI systems that can autonomously conduct the kind of exhaustive, multi-source research that has traditionally consumed hours or days of human analyst time. It also represents Google’s clearest bid yet to position its AI infrastructure as the backbone for enterprise research workflows in finance, life sciences, and market intelligence — industries where the stakes of getting information wrong are extraordinarily high.
“We are launching two powerful updates to Deep Research in the Gemini API, now with better quality, MCP support, and native chart/infographics generation,” Google CEO Sundar Pichai wrote on X. “Use Deep Research when you want speed and efficiency, and use Max when you want the highest quality context gathering & synthesis using extended test-time compute — achieving 93.3% on DeepSearchQA and 54.6% on HLE.”
Both agents are available starting today in public preview via paid tiers of the Gemini API, accessible through the Interactions API that Google first introduced in December 2025.
The launch introduces a tiered architecture that reflects a fundamental tension in AI agent design: the tradeoff between speed and thoroughness.
Deep Research, the standard tier, replaces the preview agent Google released in December and is optimized for low-latency, interactive use cases. It delivers what Google describes as significantly reduced latency and cost at higher quality levels compared to its predecessor. The company positions it as ideal for applications where a developer wants to embed research capabilities directly into a user-facing interface — think a financial dashboard that can answer complex analytical questions in near-real time.
Deep Research Max occupies the opposite end of the spectrum. It leverages extended test-time compute — a technique where the model spends more computational cycles iteratively reasoning, searching, and refining its output before delivering a final report. Google designed it for asynchronous, background workflows: the kind of task where an analyst team kicks off a batch of due diligence reports before leaving the office and expects exhaustive, fully sourced analyses waiting for them the next morning.
The Google DeepMind team framed the distinction on X: “Deep Research: Optimized for speed and efficiency. Perfect for interactive apps needing quicker responses. Deep Research Max: It uses extra time to search and reason. Ideal for exhaustive context gathering and tasks happening in the background.”
“Deep Research was our first hosted agent in the API and has gained a ton of traction over the last 3 months, very excited for folks to test out the new agents and all the improvements, this is just the start of our agents journey,” Logan Kilpatrick, who leads developer relations for Google’s AI efforts, wrote on X.
Perhaps the most consequential feature in today’s release is the addition of Model Context Protocol support, which transforms Deep Research from a sophisticated web research tool into something more closely resembling a universal data analyst.
MCP , an emerging open standard for connecting AI models to external data sources, allows Deep Research to securely query private databases, internal document repositories, and specialized third-party data services — all without requiring sensitive information to leave its source environment. In practical terms, this means a hedge fund could point Deep Research at its internal deal-flow database and a financial data terminal simultaneously, then ask the agent to synthesize insights from both alongside publicly available information from the web.
Google disclosed that it is actively collaborating with FactSet, S&P, and PitchBook on their MCP server designs, a signal that the company is pursuing deep integration with the data providers that Wall Street and the broader financial services industry already rely on daily. The goal, according to the blog post authored by Google DeepMind product managers Lukas Haas and Srinivas Tadepalli, is to “let shared customers integrate financial data offerings into workflows powered by Deep Research, and to enable them to realize a leap in productivity by gathering context using their exhaustive data universes at lightning speed.”
This addresses one of the most persistent pain points in enterprise AI adoption: the gap between what a model can find on the open internet and what an organization actually needs to make decisions. Until now, bridging that gap required significant custom engineering. MCP support, combined with Deep Research’s autonomous browsing and reasoning capabilities, collapses much of that complexity into a configuration step. Developers can now run Deep Research with Google Search, remote MCP servers, URL Context, Code Execution, and File Search simultaneously — or turn off web access entirely to search exclusively over custom data. The system also accepts multimodal inputs including PDFs, CSVs, images, audio, and video as grounding context.
The second headline feature — native chart and infographic generation — may sound incremental, but it addresses a practical limitation that has constrained the usefulness of AI-generated research outputs in professional settings.
Previous versions of Deep Research produced text-only reports. Users who needed visualizations had to export the data and build charts themselves, a friction point that undermined the promise of end-to-end automation. The new agents generate high-quality charts and infographics inline within their reports, rendered in HTML or Google’s Nano Banana format, dynamically visualizing complex datasets as part of the analytical narrative.
“The agent generates HTML charts and infographics inline with the report. Not screenshots. Not suggestions to ‘visualize this data.’ Actual rendered charts inside the markdown output,” noted AI commentator Shruti Mishra on X, capturing the practical significance of the change.
For enterprise users — particularly those in finance and consulting who need to produce stakeholder-ready deliverables — this transforms Deep Research from a tool that accelerates the research phase into one that can potentially produce near-final analytical products. Combined with a new collaborative planning feature that lets users review, guide, and refine the agent’s research plan before execution, and real-time streaming of intermediate reasoning steps, the system gives developers granular control over the investigation’s scope while maintaining the transparency that regulated industries demand.
Today’s release crystallizes a strategic narrative Google has been building for months: Deep Research is not merely a consumer feature but a piece of infrastructure that powers multiple Google products and is now being offered to external developers as a platform.
The blog post explicitly notes that when developers build with the Deep Research agent, they tap into “the same autonomous research infrastructure that powers research capabilities within some of Google’s most popular products like Gemini App, NotebookLM, Google Search and Google Finance.” This suggests that the agent available through the API is not a stripped-down version of what Google uses internally but the same system, offered at platform scale.
The journey to this point has been remarkably rapid. Google first introduced Deep Research as a consumer feature in the Gemini app in December 2024, initially powered by Gemini 1.5 Pro. At the time, the company described it as a personal AI research assistant that could save users hours by synthesizing web information in minutes. By March 2025, Google upgraded Deep Research with Gemini 2.0 Flash Thinking Experimental and made it available for anyone to try. Then came the upgrade to Gemini 2.5 Pro Experimental, where Google reported that raters preferred its reports over competing deep research providers by more than a 2-to-1 margin. The December 2025 release was the pivot to developer access, when Google launched the Interactions API and made Deep Research available programmatically for the first time, powered by Gemini 3 Pro and accompanied by the open-source DeepSearchQA benchmark.
The underlying model driving today’s improvements is Gemini 3.1 Pro, which Google released on February 19, 2026. That model represented a significant leap in core reasoning: on ARC-AGI-2, a benchmark evaluating a model’s ability to solve novel logic patterns, 3.1 Pro scored 77.1% — more than double the performance of Gemini 3 Pro. Deep Research Max inherits that reasoning foundation and layers autonomous research behaviors on top of it, achieving 93.3% on DeepSearchQA (up from 66.1% in December) and 54.6% on Humanity’s Last Exam (up from 46.4%).
Google is not operating in a vacuum. The launch arrives amid intensifying competition in the autonomous research agent space. OpenAI has been developing its own agent capabilities within ChatGPT under the codename Hermes, which includes an agent builder, templates, scheduling, and Slack integration, according to reports circulating on social media. Perplexity has built its business around AI-powered research. And a growing ecosystem of startups is attacking various slices of the automated research workflow.
What distinguishes Google’s approach is the combination of its search infrastructure — which gives Deep Research access to the broadest and most current index of web information available — with the MCP-based connectivity to enterprise data sources. No other company currently offers a research agent that can simultaneously query the open web at Google Search’s scale and navigate proprietary data repositories through a standardized protocol. The pricing structure also signals Google’s intent to drive adoption: according to Sim.ai, which tracks model pricing, the Deep Research agent in the December preview was priced at $2 per million input tokens and $2 per million output tokens with a 1 million token context window — positioning it as cost-competitive for the volume of research output it generates.
Not everyone greeted the announcement with unalloyed enthusiasm, however. Several users on X noted that the new agents are available only through the API, not in the Gemini consumer app. “Not on Gemini app,” observed TestingCatalog News, while another user wrote, “Google keeps punishing Gemini App Pro subscribers for some reason.” Others raised concerns about the presentation of benchmark results, with one user arguing that Google’s charts could be “misleading” in how they represent percentage improvements. These complaints point to a broader tension in Google’s AI strategy: the company is increasingly directing its most advanced capabilities toward developers and enterprise customers who access them through APIs, while consumer-facing products sometimes lag behind.
The practical implications of today’s launch are most immediately felt in industries that depend on exhaustive, multi-source research as a core business function. In financial services, where analysts routinely spend hours assembling due diligence reports from scattered sources — SEC filings, earnings transcripts, market data terminals, internal deal memos — Deep Research Max offers the possibility of automating the initial research phase entirely. The FactSet, S&P, and PitchBook partnerships suggest Google is serious about making this work with the data infrastructure that financial professionals already use.
In life sciences, the blog post notes that Google has collaborated with Axiom Bio, which builds AI systems to predict drug toxicity, and found that Deep Research unlocked new levels of initial research depth across biomedical literature. In market research and consulting, the ability to produce stakeholder-ready reports with embedded visualizations and granular citations could compress project timelines from days to hours.
The key question is whether the quality and reliability of these automated outputs will meet the standards that professionals in these fields demand. Google’s benchmark numbers are impressive, but benchmarks measure performance on standardized tasks — real-world research is messier, more ambiguous, and often requires the kind of judgment that remains difficult to automate. Deep Research and Deep Research Max are available now in public preview via paid tiers of the Gemini API, with availability on Google Cloud for startups and enterprises coming soon.
Eighteen months ago, Deep Research was a feature that helped grad students avoid drowning in browser tabs. Today, Google is betting it can replace the first shift at an investment bank. The distance between those two ambitions — and whether the technology can actually close it — will define whether autonomous research agents become a transformative category of enterprise software or just another AI demo that dazzles on benchmarks and disappoints in the conference room.
It’s been only a few months since OpenAI released its last big improvement to AI image generations in ChatGPT and through its application programming interface (API) — namely, a new image generation model known as GPT-Image-1.5, released in December 2025, which brought about improved instruction following, colors, and lighting.
Now, after weeks of testing, the company that kicked off the generative AI boom is unveiling a far more dramatic and even more impressive update: ChatGPT Images 2.0, which has been available not-so-secretly for several weeks on LM Arena AI, a third-party testing platform used by OpenAI and other major AI model providers to get early feedback, under the name “duct tape.”
Throughout that time, it’s already blown early users’ minds with its capacity to generate long blocks of text or disparate text panels within the same image, its insanely realistic generation of user interfaces and screenshots from popular websites and platforms, its reproduction of real life figures like OpenAI co-founder and CEO Sam Altman, and its ability to perform web research and put the results into the image itself.
Now today, it’s officially rolling out to ChatGPT users on all tiers, and OpenAI confirms it can also produce floor plans, image grids and sets of many smaller images, and character models from multiple angles, and apply almost all of these features to user-uploaded imagery as well.
The update, which encompasses the new gpt-image-2 model for API users and a suite of “Thinking” features for ChatGPT subscribers, represents a fundamental shift in how the company views visual media. As the official release notes state, “Images are a language, not decoration. A good image does what a good sentence does—it selects, arranges, and reveals”.
OpenAI did not release benchmarks to us ahead of time on ChatGPT Images 2.0, but it is safe to say the model is performing at the “state-of-the-art” based on all the outputs I’ve seen.
The move comes as the AI image model space has seen increasing competition, especially with the release of Google’s Nano Banana 2 image generation model (also known as Gemini 3 Pro Image or Gemini 3.1 Pro Image) in February 2026, which also offered dense text options “baked into” images similar to ChatGPT Images 2.0. But the latter’s fidelity in reproducing user interfaces, screenshots, and multiple image packs at once seem to exceed even Google’s latest image model’s capabilities in my brief testing and anecdotal usage and observation of other users’ images.
OpenAI spokespersons and researchers re-iterated the company’s commitments to safety and tagging its image outputs with metadata as AI generated in the face of rising reports — including one recently from The New York Times — on AI user-generated characters (AI UGC) being used as the seed for realistic AI videos posted en masse on social media as part of political influence campaigns, including showing support for historically unpopular U.S. President Donald J. Trump with an army of fictitious people masquerading as “real Americans.”
When VentureBeat asked in a closed press briefing directly about this story and GPT Images 2.0’s potential for usage in deceptive campaigning or advertising/influence campaigns Adele Li, OpenAI’s Product Lead for ChatGPT Images, responded:
“We take safety and security incredibly seriously. That includes anything when it comes to political or election interference. And so while other platforms and companies may not have those safeguards, ChatGPT does, and we take monitoring and protection of our users, as well as the influence that our photos as they are created, incredibly seriously..in the last couple years, we’ve seen a lot more new entrants into the image generation space with different standards and philosophies as ChatGPT, but we’ve stayed steady through all that, and we’re really proud of releasing this model as it relates to advanced capabilities, but doing so in a safe and protected way.”
OpenAI has also confirmed that it is deprecating GPT-Image-1.5 as the default model across its suite, though it will remain accessible via the API for legacy support. This transition signals OpenAI’s confidence that the 2.0 model is a superior replacement for both casual and high-value creative tasks.
The most significant technical advancement in Images 2.0 is the integration of OpenAI’s “O-series” reasoning capabilities.
Historically, image models have operated as black boxes: you provide a prompt, and a single output is generated. Images 2.0 introduces an “agentic” approach.
When a user selects a “Thinking” model within ChatGPT, the system no longer simply “draws”; it researches, plans, and reasons through the structure of an image before the first pixel is rendered.
During a live press briefing, Li demonstrated this reasoning by uploading a complex PowerPoint file regarding internal product strategies.
Rather than merely creating a related image, the model synthesized the document’s core data, identified the correct logos, and produced a professional poster that preserved the specific stylistic inputs of the original file.
In my brief testing — I was given access last night and tested it on a few generations this morning — ChatGPT Images 2.0 is the first image model from OpenAI and one of only two (Nano Banana 2 being the other) that can seemingly accurately reproduce a map of the extent of the Aztec, Maya, and Inca empires at their respective heights along with a fully legible legend, making it useful for educational or internal training purposes on global knowledge and geography.
This reasoning capability also allows the model to search the web in real-time to ensure visual accuracy for current events or specific technical artifacts.
This is supported by a significantly more recent knowledge cutoff of December 2025, a major leap from previous iterations that struggled with modern context.
The underlying architecture has been “revamped from scratch,” according to Research Lead Boyuan Chen. While Chen declined to confirm if the model uses a traditional diffusion or auto-regressive technique, he described it as a “generalist model” or a “GPT for images” that can handle 3D-style perspective shifts and complex spatial reasoning through simple text prompts.
The product experience for Images 2.0 is defined by three major pillars: typography, linguistic diversity, and sequential consistency.
One of the most persistent “tells” of AI-generated imagery has been the inability to render legible text. OpenAI claims Images 2.0 marks a “step change” in this department. The model is now capable of producing readable typography even in dense compositions, such as scientific diagrams, menus, or infographic posters.
A look at the provided “Magazine Cover” sample (Open Scifi) illustrates this precision: every headline, volume number, and even the “Display until” date on the barcode is rendered with crisp, professional alignment that mirrors human-designed layouts.
This capability extends into the “Thinking” mode, where the model can even generate three-page educational visuals—complete with quizzes—that maintain a consistent instructional flow.
OpenAI has also addressed a long-standing Western bias in AI imagery. Images 2.0 is described as a “polyglot” model with significant gains in non-Latin script rendering. Specifically, the model now supports high-fidelity text generation in Japanese, Korean, Chinese, Hindi, and Bengali.
In the “Global Language” diagram provided, which explains the water cycle, the model successfully renders complex Korean characters (Hangul) within an educational layout.
The text is not just translated; it is “rendered correctly but with language that flows coherently,” ensuring that labels and explanations feel natively integrated into the design.
For creators working on storyboards or brand campaigns, the most impactful new feature is the ability to generate up to eight distinct images from a single prompt. Crucially, these images maintain “character and object continuity” across the series.
Li noted that this solves a “cumbersome” workflow where users previously had to prompt one image at a time and manually stitch them together. This feature enables the creation of entire manga sequences, children’s books, or a family of social media graphics that share the same visual DNA.
OpenAI’s rollout strategy reflects a clear push toward professional and enterprise adoption. While the base model is available to all users—including those on the free tier—the advanced “Thinking” and “Pro” capabilities are reserved for paid tiers.
Free Users: Have access to the base ImageGen 2.0 model for standard tasks.
Plus and Pro Users: Can access “Thinking” capabilities, which include tool use, web search, and multi-image generation.
Pro Users: Receive additional access to “ImageGen Pro” models for more advanced image generation.
API Developers: Can integrate gpt-image-2, which supports resolutions up to 4K (currently in beta) and flexible aspect ratios ranging from a wide 3:1 to a tall 1:3.
Pricing in the API is as follows, echoing GPT-Image-1.5, the predecessor model, but actually shaving off $2 on the output side:
Image
$8.00 for inputs
$2.00 for cached inputs
$30.00 for outputs
Text
$5.00 for inputs
$1.25 for cached inputs
$10.00 for outputs
What is clear so far is that OpenAI is describing three practical layers of access, even if it has not published a precise tier-by-tier matrix.
The baseline is ChatGPT Images 2.0, which OpenAI’s blog post states is available to all ChatGPT and Codex users and includes the core model improvements: better instruction following, stronger text rendering, multilingual gains, broader aspect ratios, and more polished, production-usable outputs.
Above that is “thinking”, which the release defines more concretely: when a thinking model is selected, the system can take more time, use the web, analyze uploaded materials, reason through layout before generating, and produce multiple distinct images at once, including up to eight coherent outputs with continuity.
In the briefing, Li also framed thinking and Pro as “juiced-up” versions of the base model with tool use, and said these advanced modes are slower, not faster, because they do more reasoning and search behind the scenes. What remains unclear is the exact feature boundary between Thinking and Pro.
The materials say Pro users get access to more advanced image generation, but they do not spell out whether that means higher quality, higher limits, higher resolution, more outputs, or some other advantage distinct from thinking itself.
For enterprise users, the safest way to think about the differences is not as three totally separate products, but as a spectrum from fast default generation to slower, more agentic, more structured generation.
If a team needs quick creative drafts, marketing concepts, simple graphics, or everyday image edits, the base Images 2.0 model appears to be the relevant default.
If the task involves factual grounding, transforming internal documents into explainers, creating multi-image sets, or maintaining consistency across a sequence of assets, the more important distinction is whether the organization has access to thinking-enabled outputs.
Until OpenAI provides a clearer Pro-versus-Thinking breakdown, enterprise buyers should treat “thinking” as the meaningful functional upgrade and treat “Pro” as a possibly higher-end access tier whose exact incremental benefits still need clarification before procurement or workflow planning.
OpenAI’s says ChatGPT Images 2.0 offers a”multi-layered stack” of safety protocols, including:
Provenance: Adhering to industry standards for watermarking so that AI-generated images are identifiable.
Model Safeguards: Using advanced perception models to filter out harmful or abusive content for both adults and children.
Active Monitoring: Enforcing user policies through real-time reporting.
Li emphasized that while their philosophy is to “maximize user creativity,” they maintain strict policies against election interference.
The shift from Images 1.5 to 2.0 is more than a resolution bump. By integrating reasoning, OpenAI is attempting to solve the “intent gap” that has plagued AI art since its inception.
When you ask an AI for an “infographic about supply and demand,” you aren’t just looking for a picture; you are looking for a logical layout of information.
The “Interior Design” sample (Japandi Furnishing Concept) highlights this systemic thinking. The model didn’t just generate a room; it created a cohesive floor plan, a color palette, a list of materials, and “inspiration” shots that all adhere to a singular aesthetic.
This is what OpenAI calls moving from a “tool” to a “visual system”. However, this increased capability comes with a trade-off in speed.
For the professional user, this is likely a worthwhile exchange: waiting an extra minute for a “production-ready asset” is still significantly faster than the hours required for manual design.
As ChatGPT Images 2.0 rolls out, it marks the beginning of an era where AI doesn’t just assist in making art, but in conducting “economically valuable creative tasks”.
Whether it can truly replace the intentionality of a human designer remains to be seen, but with 2K resolution, multilingual fluency, and the ability to “think” before it acts, OpenAI has certainly closed the distance.
Beats cables already bring elements that Apple’s own versions don’t, like a bunch of colors. Now, a new longer-than-ever version is here.
Satechi’s new ChargeView GaN Charger sits on a desk and is able to charge up to four devices through its four USB-C ports distributing up to 140W of total power output.
The latest phone from Oppo includes dual 200MP Hasselblad cameras and a 10x optical telephoto lens. But it doesn’t come cheap.
False reports – including a fake CNN screenshot – are circulating on social media, claiming that a product based on honey can cure Alzheimer’s disease.
Sennheiser says professional users like sound engineers and producers want headphone with a tight and accurate bass as well as a supremely comfort fit.
Apple is about to get a new CEO, with John Ternus taking over from Tim Cook on Sept. 1. Cook will become Executive Chairman.
A new leak says that more iPhones will lose compatibility with the next software than almost any update before. Is your iPhone on the list?