Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI

Microsoft AI released two new in-house models into public preview on Wednesday — MAI-Image-2.5-Pro, its highest-fidelity image generator to date, and MAI-Voice-2-Flash, a speech model built for high-volume enterprise workloads — while publishing production data that amounts to the company’s most aggressive argument yet that it can power its own products without leaning on OpenAI’s frontier models.

The announcement, made by Microsoft AI’s Superintelligence team, lands roughly a year after the company committed to building purpose-built models internally, and it arrives with an unusual level of specificity about where those models now run: Bing, PowerPoint, OneDrive, Dynamics 365, Excel, GitHub Copilot, and Azure. The message to enterprise buyers — and, implicitly, to OpenAI — is that Microsoft’s homegrown models are no longer research projects. They are production infrastructure serving millions of users.

“Each of these enhancements is a step toward the same goal: Microsoft products, powered by Microsoft models,” the company wrote in its announcement blog.

How MAI-Image-2.5-Pro and MAI-Voice-2-Flash stake out opposite ends of the AI cost curve

The two new releases occupy opposite ends of what Microsoft calls the quality-speed-cost curve, and the positioning is deliberate. MAI-Image-2.5-Pro targets the premium tier: hero imagery, detailed editing, and precise in-image text rendering — the last of which has long been a notorious weak spot for image generation models. Microsoft priced the model at $5 per million text input tokens, $8 per million image input tokens, and $106 per million image output tokens. The base MAI-Image-2.5 model recently launched at No. 2 for image editing on Arena, the community leaderboard that has become a de facto scoreboard for generative media.

The creative industry appears to be taking notice. Rob Reilly, global chief creative officer at advertising giant WPP, called the Pro model “a strong leap forward for GenMedia tools” in a statement included in Microsoft’s announcement, adding that “Microsoft has firmly established itself among the leaders in generative AI.”

MAI-Voice-2-Flash goes the other direction. First previewed at Microsoft’s Build conference, Flash runs twice as fast as MAI-Voice-2 and costs 32% less, priced at $15 per million characters. It is designed for the unglamorous but enormous market of high-volume voice — call centers, voice agents, and real-time speech applications where latency and cost-per-call matter more than marginal gains in expressiveness. Together, the two models reflect a strategy of building families of models rather than a single flagship, because, as the company put it, a creative studio chasing maximum fidelity has very different needs from a customer service operation handling millions of calls a day.

Microsoft’s production metrics show in-house models cutting GPU costs by up to 89%

The model launches are arguably less newsworthy than the deployment metrics Microsoft attached to them — numbers that read like a systematic case for swapping out third-party frontier models across its product portfolio. 

Bing Image Creator now runs entirely on MAI-Image-2.5, end to end, marking the first time the consumer image tool is fully in-house. In PowerPoint, Microsoft says MAI-Image-2.5 reduces GPU costs by up to 84% compared with GPT-Image-2, OpenAI’s image model. In OneDrive, where MAI-Image-2.5 is now the default for key image-editing scenarios, the company reports a 26% increase in save rates, roughly 25% lower P95 latency, and 2.5 times greater efficiency under medium-utilization production workloads.

On the voice side, MAI-Voice-2-Flash now powers Dynamics 365 Contact Center — the platform used by customers including T-Mobile and EasyJet — where Microsoft claims GPU cost reductions of up to 89%. The model is also integrated into Azure Voice Live for developers building speech-to-speech agents.

Perhaps the most consequential deployment sits in healthcare. Microsoft’s Dragon Copilot, used by 170,000 medical providers and responsible for processing 28 million patient encounters last quarter, now runs on MAI-Transcribe-1.5 for its multilingual workflow across 58 languages. Microsoft says internal evaluations show a 50% relative reduction in both transcription and language-identification error rates across most languages — a meaningful claim in a domain where transcription errors can propagate directly into clinical notes.

Inside the ‘hill-climbing’ strategy that lets small models beat GPT-5.6 in Excel

In a companion post published the same day, Microsoft detailed the methodology behind these results — what it calls its “hill-climbing machine,” an integrated flywheel of data, models, and the product “harness” that surrounds them.

The clearest example is MAI-Code-1-Flash, the lightweight coding model launched in GitHub Copilot in June. Microsoft says the model achieves an approximately 10% higher code accept rate than GPT-5.4 Mini and Claude Haiku 4.5 in VS Code, while using 10% fewer median tokens. Developer retention tells a similar story: users were 6% more likely to return across multiple days than with GPT-5.4 Mini, and 11% more likely than with Claude Haiku 4.5.

Then Microsoft did something more interesting. It took the MAI-Code-1-Flash checkpoint and further trained it inside an Excel reinforcement learning environment, teaching a coding model the tools and workflows of spreadsheet knowledge work. The result, according to production user feedback, is a model on par with GPT-5.6 for the most common Excel tasks — while being small enough to run on Nvidia’s older H100 and even A100 GPUs rather than requiring the latest-generation accelerators.

That hardware detail deserves emphasis. Every major AI company is fighting for allocation of cutting-edge chips, and a model that delivers frontier-adjacent quality on two-generation-old silicon fundamentally changes the deployment economics. It also frees the newest hardware — including Microsoft’s now-operational GB200 cluster — for training rather than serving.

Satya Nadella’s ‘frontier diffusion’ manifesto redraws the OpenAI relationship

Microsoft CEO Satya Nadella framed the announcements in a lengthy post on X titled “Frontier Diffusion & Control,” which functions as something close to a strategic manifesto. “We can now take saturated frontier capabilities and deliver them at scale and at lower cost through models optimized for high-usage products, while continuing to use frontier models for frontier needs,” Nadella wrote, adding that Microsoft is “beginning to route traffic across our first-party surfaces to MAI whenever our models match or outperform frontier alternatives.”

Translated from executive prose: capabilities that were state-of-the-art a year ago are now table stakes, and Microsoft believes it can replicate them cheaply for the specific, repetitive tasks that dominate real product usage. Why pay frontier prices for a frontier model when a user just wants to reformat a spreadsheet column?

Nadella was careful to note that “frontier models from OpenAI and Anthropic are part of the orchestration system alongside MAI” — but he also articulated a pointed principle of model independence, arguing that a company’s evaluations “should continue to hill climb even when any given model has been removed.” 

“Keeping the harness, memory, context, and skills outside the model, he argued, is what gives Microsoft control. The subtext is hard to miss. Reuters reported in April that Microsoft’s exclusive license to OpenAI’s technology had been revised into a non-exclusive arrangement, and The Information reported last September that Microsoft had begun incorporating Anthropic models into some products. Wednesday’s announcement completes the triangle: Microsoft as orchestrator, with its partners’ frontier models as interchangeable components and its own models absorbing an ever-larger share of routine traffic.”

Developers cheer cheaper task-specific models while skeptics question Microsoft’s track record

The response online captured both the appeal and the skepticism surrounding the strategy. “I love when people use small models for niche tasks,” wrote one X user, @mavihsk, responding to Nadella’s post. “Why do I have to use the all-knowing model just to change my field in Excel?” Another user, @nabu_lines, distilled the pitch neatly: “cost and performance both improve when you stop overusing the biggest model.”

Others were less charitable about Microsoft’s execution track record. “Microsoft is the worst when it comes to listening to user feedback,” wrote designer @designedbyabin, arguing the company “will lose the AI race because they repeatedly failed to understand user needs.” And one user, @tokenoverflow, offered a drier critique of the model-independence pitch: “i want it keep hill climbing after removing microsoft.”

The skeptics raise a fair point. Microsoft’s self-reported metrics — accept rates, save rates, GPU savings — come from its own internal evaluations, not independent benchmarks, and the company chooses which comparisons to publish.

But the strategy’s logic does not depend on any single number. Nadella’s framing that software now has “real marginal cost for the first time” explains why Microsoft is obsessive about tokens, GPUs, and serving costs: when AI features run on every keystroke across a billion-user product portfolio, an 84% GPU cost reduction is not an optimization. It is the difference between a viable business and a money pit.

Why Microsoft is turning its internal AI playbook into an Azure product

The final piece of the strategy is that Microsoft is selling the playbook, not just the models. Nadella explicitly positioned the hill-climbing approach as “a template for every other AI native, SaaS, or Enterprise company,” and Microsoft is packaging the toolchain through Foundry and what it calls Frontier Tuning — letting enterprises train specialized models against their own proprietary evaluations and reinforcement learning environments. That turns Microsoft’s internal cost-cutting exercise into an Azure product, and it gives enterprise customers a reason to run their AI workloads on Microsoft’s cloud even if the models themselves come from elsewhere.

The company’s emphasis on models trained “on clean, traceable, enterprise-grade data, without distillation from third-party models” serves the same commercial end. In an industry facing mounting scrutiny over training data provenance, Microsoft is betting that enterprise buyers — and courts — will care where model capabilities come from. Microsoft says it is now extending the hill-climbing approach to Copilot Chat, Outlook, and PowerPoint, and both new models are available in public preview through Microsoft Foundry and the MAI Playground. “None of this is an endpoint,” the company wrote. “We’re just getting started.”

Seven years ago, Microsoft bet more than $13 billion that OpenAI would build the future of AI. Wednesday’s announcement suggests the company has since learned a cheaper lesson: the future of AI may belong to whoever builds the frontier, but the profits belong to whoever makes it ordinary.

Poolside drops Laguna S 2.1, an open-weight coding model that beats rivals 10x its size

Poolside, the San Francisco AI lab that has spent most of its three-year existence quietly selling coding models to governments and defense agencies, released its most capable model to date on Tuesday — and made an unusually aggressive bet that radical transparency, not raw scale, is how a smaller lab competes at the frontier.

The model, Laguna S 2.1, is a 118-billion-parameter Mixture-of-Experts (MoE) system that activates only 8 billion parameters per token, supports a context window of up to 1 million tokens, and — according to benchmarks published by the company — matches or beats open models several times its size on agentic coding tasks. The weights are available immediately on Hugging Face under the permissive OpenMDW-1.1 license.

The headline numbers are striking for a model this small. Poolside reports that Laguna S 2.1 scores 70.2% on Terminal-Bench 2.1, a benchmark of long-horizon terminal tasks, placing it 11th on the company’s compiled leaderboard — ahead of DeepSeek-V4-Pro-Max, a 1.6-trillion-parameter model that scored 64.0; Thinking Machines’ 975-billion-parameter Inkling, at 63.8; and Nvidia’s 550-billion-parameter Nemotron 3 Ultra, at 56.4. On SWE-Bench Multilingual, it posts 78.5%, and on SWE-Bench Pro‘s public dataset, 59.4%.

Perhaps more telling than any single score: the model went from the start of pre-training on May 22 to public launch in under nine weeks, trained on 4,096 Nvidia H200 GPUs. In an industry where flagship model cycles are typically measured in quarters or years, Poolside has now shipped three models in three months.

Why the West’s open-weight AI gap has become a boardroom issue

The release lands in the middle of an increasingly pointed debate about the provenance of open-weight AI. Over the past year, developer adoption has shifted decisively toward open-weight systems that companies can download, inspect, and run on their own infrastructure — and the leading options in that category have overwhelmingly come from Chinese labs. DeepSeek, Qwen, Kimi, GLM, MiniMax, and Tencent’s Hunyuan line all feature prominently in Poolside’s own comparison tables.

Poolside’s accompanying press release frames Laguna S 2.1 explicitly as a response, noting that the model occupies a size class into which no Western lab has released open weights in 11 months — since OpenAI’s gpt-oss-120b last August. “The West needs open-weight models it can trust, run, and build on,” said Jason Warner, Poolside’s co-CEO, in the announcement.

Co-founder and co-CEO Eiso Kant made the philosophical stakes even plainer in a lengthy post on X. “I believe intelligence should and will become a commodity,” he wrote, arguing that the open ecosystem “will not win by being the best in its own category.” Users, he argued, simply want the best intelligence for the task at hand — so open models must be on par with, or better than, their closed equivalents.

The strategic logic here is not charity. Poolside’s core business is deploying models inside the security boundaries of government, defense, and regulated enterprises — customers for whom closed, metered API access is often a non-starter for compliance and sovereignty reasons. 

Every enterprise that standardizes on a Chinese open model today becomes harder to win tomorrow. Releasing competitive open weights is both an ecosystem play and a top-of-funnel strategy for the company’s high-security deployment business. It also reframes the AI race away from terrain where Poolside cannot compete — frontier-scale capital expenditure — and toward terrain where it believes it can: cost per token, self-hosting, and iteration speed.

How a sparse architecture makes enterprise AI agents affordable to run

The technical design reflects a specific thesis about where value in coding AI is moving. Laguna S 2.1’s sparse MoE architecture — 256 routed experts plus one shared expert, with grouped-query attention and interleaved sliding-window layers, according to the Hugging Face model card — means inference costs scale with the 8 billion active parameters, not the 118 billion total. Poolside emphasizes that the model is small enough to run on a single Nvidia DGX Spark, the desktop-class AI machine.

That matters for what Poolside calls token economics. Long-horizon coding agents are voracious consumers of tokens: the company’s published data shows the model consuming a mean of roughly 249,000 completion tokens per trajectory on its hardest benchmark when thinking mode is enabled. At metered API prices, agentic workloads at enterprise scale become a meaningful budget line item. On OpenRouter, Poolside is offering a free 256K-context endpoint and a dedicated 1M-context deployment priced at $0.10 per million input tokens and $0.20 per million output tokens — aggressive pricing that undercuts most frontier alternatives by an order of magnitude.

The ecosystem support is unusually broad for day one. The model is live on Baseten’s model library and Vercel’s AI Gateway, with integrations across vLLM, SGLang, Ollama, and llama.cpp, plus quantized variants down to 4-bit GGUF files — 75 gigabytes — for local use. But Poolside’s more interesting claim is behavioral, not architectural. Pengming Wang, co-head of applied research at Poolside, said the gains came from improving the model’s working habits: “more verification, less taking things for granted, not declaring victory early, and being more persistent.” Raw intelligence, the company argues, is one axis of capability; a model’s way of working is a second axis that matters immensely for agents left unattended for hours.

Publishing every benchmark trajectory to counter AI’s credibility crisis

The most consequential part of the release for enterprise buyers may be an evaluation-transparency move with little precedent among major labs: Poolside published the complete, unedited trajectory of every trial in its final benchmark runs — every reasoning step, tool call, and shell command behind every reported score.

This addresses a growing credibility problem in AI benchmarking. As top scores on mature benchmarks cluster in the 70–90% range, and as “reward hacking” — models finding solutions online or gaming verifiers rather than solving problems — has become endemic, self-reported numbers have lost much of their signal. Poolside disclosed its own encounters with the problem candidly: during training, more than half of trajectories on some SWE-bench tasks were flagged because the model simply researched the original bug-fix pull request online and applied it. The company documented its mitigations, including prompt addenda, LLM-based judging calibrated against human labels, and expert annotator review of a high-scoring Terminal-Bench run.

Three published case studies illustrate what the company means by persistence. In one, the model built a working HTML/CSS rendering engine from an empty folder in a 181-step, 50-minute unattended session — then, lacking vision capabilities, spun up headless Chromium to numerically compare its canvas output against a real browser’s rendering. In another, pointed at Poolside’s own agent harness in an automated optimization loop, the model made the Go codebase 5.2% faster with roughly 70% lower memory allocation, finding an O(n²) string-concatenation bug along the way. In a third, working in a sandbox with no Python installed, the model did its number theory in Perl and independently re-derived a proof of Erdős problem #397 — a combinatorics question open for five decades until GPT-5.2 Pro first solved it this past January. Poolside notes that its model’s construction is structurally different from the earlier published solution, and that its November 2025 knowledge cutoff precedes the first proof.

What the disclosed limitations and benchmark fine print reveal

Poolside deserves credit for disclosing limitations most labs bury. The model can overfit to its native harness and stumble on slightly different tool schemas in third-party agents, mangles JSON in nested tool arguments, and is prone to overthinking on competition math. There is currently no user-configurable thinking-effort dial — just on or off — and the gap between the modes is enormous: thinking lifts Terminal-Bench 2.1 from 60.4% to 70.2%, and DeepSWE from 16.5% to 40.4%, at substantially higher token cost.

Buyers should apply their own discounts to the comparison tables. Poolside’s methodology takes the maximum of vendor self-reported scores, benchmark-author leaderboards, and third-party figures for competitors — a reasonable convention, but one that mixes harnesses and test conditions. On DeepSWE, notably, Poolside ran its own agent harness rather than the leaderboard’s standard mini-swe-agent, a difference the company acknowledges makes scores less directly comparable. And the frontier remains clearly out of reach: closed models like GPT-5.6 Sol, at 88.8 on Terminal-Bench 2.1, and Claude Fable 5, at 88.0, along with the 2.8-trillion-parameter open-weight Kimi K3, at 88.3, sit well above Laguna S 2.1.

The deeper structural question is whether Poolside’s “Model Factory” — the internal platform the company credits for its rapid release cadence — can sustain this pace as models scale. The trajectory so far is genuinely unusual: the April dual release of Laguna M.1 and XS.2, the July 2 refresh of XS 2.1, and now S 2.1, which the company says outperforms April’s flagship M.1 at roughly a third of its active size. Remarkably, S 2.1 used the exact same pre-training data as XS 2.1, meaning nearly all the improvement came from scale, training fixes, and post-training across the company’s corpus of 409,000 agentic and non-agentic training environments. Poolside says its next, larger Laguna model began pre-training last week.

For technical decision makers, Laguna S 2.1 is the most credible Western open-weight option to emerge in nearly a year for self-hosted agentic coding — with published evidence, a permissive license, broad ecosystem support, and an economics story built around hardware you can own. Whether it dents the dominance of Chinese open models will depend less on this release than on the ones that follow it.

Kant, for his part, has already told the world how he intends that story to end. Poolside is building toward a future where the most capable intelligence “can be owned and shaped by anyone,” he wrote — and the company plans to keep shipping “until that future exists.” In an industry where the biggest labs increasingly lock their best work behind an API, the most radical thing about Laguna S 2.1 may not be what it scores, but that anyone can download it and check.

Capital One releases VulnHunter, an open-source AI tool that finds software flaws before hackers do

Capital One on Thursday released VulnHunter, an open-source, agentic AI security tool that scans source code for exploitable vulnerabilities, maps out how an attacker would reach them, and proposes targeted fixes — all before a single line ships to production. The tool, built internally and now available on GitHub under an Apache 2.0 license, is one of the most ambitious attempts by a major financial institution to turn offensive AI capabilities into a public defensive resource.

The move marks a striking philosophical turn for a company still defined, in many boardrooms, by a 2019 data breach that compromised the personal information of roughly 106 million people across the United States and Canada and ultimately cost the bank an $80 million federal fine.

Capital One is not simply releasing another vulnerability scanner. VulnHunter introduces what the company calls an “attacker-first forward analysis” — a workflow in which the tool begins at the points where a real adversary would enter a system, such as APIs, network messages, or file uploads, and reasons forward through the application’s logic to determine whether an exploit path actually survives the code’s existing defenses. Conventional scanners typically work in reverse, flagging a dangerous-looking code pattern and then searching backward for a hypothetical attacker. That approach, security practitioners widely acknowledge, buries engineering teams under avalanches of false positives.

VulnHunter attacks that problem head-on with a second innovation: a built-in “falsification engine” that tries to disprove its own findings before a developer ever sees them. After the tool surfaces a potential vulnerability, a structured reasoning workflow hunts for logical gaps, unsupported assumptions, and conditions that would prevent the attack from succeeding. Only findings the engine fails to rule out reach a human reviewer — and when they do, VulnHunter delivers not just an alert but a full explanation of the exploit path and a proposed code fix ready for engineering review.

The tool currently runs on Anthropic’s Claude Opus 4.8 model inside a Claude Code environment, though Capital One says the framework has the potential to work across other foundation models and coding harnesses.

The 2019 breach that reshaped how Capital One thinks about cybersecurity

To understand why Capital One chose to open-source a tool this consequential, you have to understand the scar tissue.

On July 19, 2019, Capital One disclosed that an outside individual — later identified as a former Amazon Web Services employee named Paige Thompson — had gained unauthorized access to names, addresses, self-reported income, Social Security numbers, and linked bank account numbers belonging to credit card customers and applicants. The breach, which Capital One says occurred on March 22 and 23, 2019, was discovered only after an external security researcher flagged a configuration vulnerability through the company’s Responsible Disclosure Program on July 17 of that year.

The damage was sweeping. Approximately 100 million people in the United States and 6 million in Canada were affected. Roughly 140,000 Social Security numbers, about 80,000 linked bank account numbers, and approximately 1 million Canadian Social Insurance Numbers were compromised. The FBI arrested Thompson, and the government stated it believed the data had been recovered with no evidence of fraud. But the reputational and regulatory toll was enormous.

In August 2020, the Office of the Comptroller of the Currency fined Capital One $80 million, finding that the bank had failed to adequately identify and manage risks as it migrated significant technology operations to the cloud. As Reuters reported at the time, the OCC’s consent order cited insufficient network security controls, inadequate data loss prevention measures, and a board that failed to hold management accountable when internal auditing surfaced problems. The OCC also ordered Capital One to overhaul its operations and submit new cybersecurity plans for regulatory review.

The incident became an industry case study in the dangers of moving fast with new technology. As CyberScoop reported in July 2019, a cybersecurity executive at a competing financial company observed that the breach “could be the result of trying too many new things and forcing them through.” Capital One’s own CEO, Richard D. Fairbank, acknowledged the gravity of the moment. “While I am grateful that the perpetrator has been caught, I am deeply sorry for what has happened,” Fairbank said at the time. “I sincerely apologize for the understandable worry this incident must be causing those affected and I am committed to making it right.”

How Capital One rebuilt its security reputation through open-source investment

What followed was not a retreat from technology but a doubling down — with security explicitly at the center.

Capital One had declared itself an “open-source first” company in 2015 as part of a broader technology transformation that began over a decade ago. After the breach, the company accelerated its investments in software supply chain security, open-source governance, and AI-driven defense. In August 2022, Capital One joined the Open Source Security Foundation as a premier member, earning a seat on the organization’s Governing Board. Chris Nims, then EVP of Cloud & Productivity Engineering, framed the move as a natural extension of the company’s operating philosophy. “As a highly-regulated company, we are seasoned in managing compliance and governance and advocate for standardization, automation and collaboration,” Nims said in the OpenSSF announcement.

Behind that public commitment lay a substantial operational apparatus. Capital One’s Open Source Program Office, now in its third iteration, manages open-source usage, contributions, and community building across the enterprise. The company has released more than 25 open-source projects and made over 2,000 contributions to approximately 135 external open-source projects, according to the company’s own disclosures. Those efforts address not just code dependencies but the entire software development lifecycle — DevSecOps tools, infrastructure, and the collaborative environments, both internal and external, that shape how software gets built and shipped.

Nureen D’Souza, the director who leads Capital One’s OSPO, has spoken publicly about the philosophy underpinning this work. At cdCon 2022, D’Souza described a “company-wide culture with security ingrained” that allows developers to focus on innovation rather than maintenance chores, as reported by SD Times. The OSPO’s charter emphasizes three pillars: standardization of open-source processes, automation of security policies throughout the delivery pipeline, and ecosystem sustainability through upstream contributions to the foundations and projects the company depends on.

VulnHunter is the most consequential product of that multi-year effort — and the clearest signal yet that Capital One views open-source collaboration not as charity but as a competitive security strategy. The company argues that modern software supply chains are so deeply interconnected that a single vulnerability in a widely used open-source component can cascade across thousands of enterprises simultaneously. Proprietary defenses, no matter how sophisticated, cannot address a problem that is fundamentally communal. By releasing VulnHunter under a permissive license, Capital One invites the global security research community to stress-test, extend, and improve the tool — effectively crowdsourcing its own defense infrastructure while strengthening the broader ecosystem.

Inside VulnHunter’s three-stage AI engine for finding exploitable code

For engineering leaders evaluating VulnHunter, the technical architecture is where the tool’s ambitions become concrete. The workflow unfolds in three distinct stages.

In the first stage — attacker-first forward analysis — VulnHunter begins at the points where an external adversary would interact with a system: API endpoints, network message handlers, file upload interfaces. From each entry point, the tool reasons forward through application logic, tracing data flows, transformations, and internal security checkpoints to determine whether an attacker can actually reach a dangerous code path. This approach mirrors how a skilled penetration tester would probe a system, but automates the process at a scale no human team could match.

The second stage is where VulnHunter departs most sharply from conventional scanners. After identifying a potential vulnerability, the falsification engine runs a structured reasoning workflow designed to disprove its own conclusion. It searches for assumptions that do not hold, logical gaps in the exploit path, and environmental conditions that would prevent an attack from succeeding. Findings that fail this internal challenge are discarded before any developer sees them. Capital One’s explicit goal is to shift the developer’s burden away from triaging false alarms — a perennial pain point that erodes trust in security tooling and slows development velocity.

In the third stage, vulnerabilities that survive the falsification engine trigger an evidence-backed remediation workflow. VulnHunter gathers supporting evidence across the codebase, maps the complete surviving exploit path, explains the defect and the specific capabilities an attacker would gain, and generates targeted code changes for engineering review. The output is not a generic advisory but a concrete, context-aware patch proposal.

Capital One says it validated VulnHunter internally before release, running it across thousands of repositories spanning tens of business areas. The company reports that the tool identified and remediated vulnerabilities with speed and efficiency that far exceeded what its teams previously achieved through manual triage.

Why AI-powered attacks are forcing banks to rethink traditional cyber defenses

VulnHunter arrives at a moment when the cybersecurity landscape is shifting beneath the feet of every enterprise. Capital One’s announcement frames the urgency in stark terms: advanced AI models have “dramatically lowered the barrier for bad actors to discover and exploit vulnerabilities in software,” and the window before sophisticated AI attack capabilities become affordable and accessible to virtually every adversary is shrinking rapidly.

The company’s own AI security researchers have been tracking these trends closely. At NeurIPS 2024 in Vancouver, Capital One’s team presented research and curated a list of nearly 100 papers spanning LLM safety, adversarial resilience, jailbreak attacks, and synthetic data generation. The papers they highlighted — including work on multi-agent defense frameworks, automated red-teaming, and guardrail classifiers — paint a picture of an arms race in which offensive and defensive AI capabilities are co-evolving at breakneck speed.

Several of those research themes map directly onto VulnHunter’s architecture. The falsification engine echoes the adversarial defense strategies explored in papers like “BackdoorAlign,” which demonstrated that embedding a structured safety mechanism into a small number of training examples could recover a model’s safety alignment without degrading performance. The attacker-first forward analysis reflects the philosophy of “WildTeaming,” a framework that collects and analyzes real-world jailbreak attempts to build more resilient models. And VulnHunter’s emphasis on minimizing false positives parallels the goals of “GuardFormer,” a guardrail classifier that outperformed GPT-4 on safety benchmarks while running 14 times faster.

The thread connecting all of this work is a conviction that traditional, reactive security — monitoring networks, patching known vulnerabilities, responding to incidents after they occur — is no longer sufficient when adversaries can use AI to discover and exploit zero-day vulnerabilities at machine speed. The only durable defense, Capital One argues, is to find and fix the vulnerabilities in your own code before attackers find them first.

What Capital One’s cloud security journey reveals about the entire banking industry

Capital One’s arc from breach victim to open-source security contributor also illuminates a broader reckoning across financial services. When Capital One moved aggressively to Amazon Web Services in the mid-2010s, it was a rarity among major banks. Most financial institutions simply did not trust third parties to store their most sensitive data. Capital One’s CIO at the time, Rob Alexander, publicly championed the cloud as more secure than the bank’s own data centers — a claim that the 2019 breach complicated considerably.

The CyberScoop report from that period captured the tension within the industry. W. Patrick Opet, managing director of cybersecurity at JP Morgan Chase, described a cultural shift in banking from prioritizing traders to prioritizing developers: “Now, it’s ‘Focus on the developer, turn everything into code, and automate everything.'” Mark Nicholson, Deloitte’s cyber leader for the financial industry, noted that the pressure to move quickly was exposing “weaknesses in the development methodology.” And the breach itself was a reminder that even as Chase spent $600 million annually on cybersecurity, relatively simple vulnerabilities — like the Apache Struts bug that enabled the Equifax breach — could undercut massive investments in data protection.

Seven years later, the industry has largely followed Capital One into the cloud, and the security challenges have only intensified. The question is no longer whether to use cloud infrastructure but how to secure the software that runs on it. VulnHunter represents Capital One’s answer: rather than relying solely on network-level controls and perimeter defenses, push security directly into the code itself, at the moment it is written. The open-source release also carries implicit competitive pressure. If VulnHunter gains traction among developers and security teams, it could set a new baseline for what enterprise security tooling is expected to do — and force rival banks, fintechs, and cloud providers to match or exceed its capabilities.

Whether VulnHunter lives up to that ambition will depend on adoption, community engagement, and the tool’s real-world performance against the increasingly sophisticated AI-powered attacks it was designed to counter. But the release itself tells a story that extends well beyond any single tool or any single company. In 2019, a misconfigured firewall exposed 100 million records and turned Capital One into a cautionary tale about the cost of moving fast without moving carefully. In 2026, the same institution is open-sourcing the kind of AI-driven defense it wishes it had built sooner — and betting that the best way to protect its own code is to help the entire industry protect theirs.

China’s Moonshot AI releases Kimi K3, the largest open-source model ever, rivaling top U.S. systems

Moonshot AI, the Beijing-based artificial intelligence startup backed by Alibaba, on Thursday released Kimi K3 — a 2.8-trillion-parameter model that the company says is now the largest open-source AI model in the world, and one that benchmarks show performs neck-and-neck with the most powerful proprietary systems from Anthropic and OpenAI.

The release, timed to land just ahead of the 2026 World Artificial Intelligence Conference in Shanghai, is a dramatic escalation in the global AI arms race and a watershed moment for the open-source AI movement. It also marks a remarkable comeback for a company whose market position had eroded significantly over the past 18 months following DeepSeek’s meteoric rise.

Full model weights are scheduled to be released on July 27, according to details shared by researchers who reviewed the company’s technical documentation. If you want to take Kimi K3 for a spin right now, you can — just head to kimi.com, sign up with a Google account or phone number (no credit card required), and start chatting with what may be the most powerful open-source model ever built.

Inside the architecture that powers the world’s largest open-source AI model

Kimi K3 is a frontier-class large language model with 2.8 trillion total parameters — roughly 75 percent larger than DeepSeek’s V4 Pro, which the company’s own timeline chart shows at approximately 1.6 trillion parameters. The model features a 1-million-token context window, native visual understanding capabilities, and an always-on reasoning mode that the company calls “thinking mode.”

The model is built on two key architectural innovations developed internally at Moonshot AI: Kimi Delta Attention, a hybrid linear attention mechanism, and Attention Residuals, which the company describes as a drop-in replacement for residual connections that delivers consistent scaling gains. Both techniques were previously published as open research by the Moonshot team on GitHub.

On the API side, Kimi K3 is compatible with the OpenAI SDK, lowering the integration barrier for developers already building on OpenAI or Anthropic toolchains. The model is priced at $3 per million input tokens and $15 per million output tokens, with cached input tokens dropping to just $0.30 per million — pricing that positions it roughly in line with mid-tier offerings from Western labs, but at a performance level the company claims approaches the top of the market. A promotional top-up rebate running through August 12 offers up to 30 percent back in vouchers for API credits of $1,000 or more.

As Xinhua reported, a Moonshot AI executive explained the significance of the parameter count in simple terms: parameters are like neural connections in the human brain, and nearly 3 trillion of them means the model can “store more knowledge and patterns in its brain, understand more, think deeper, and answer more accurately.”

Benchmark results show Kimi K3 trading blows with Claude and GPT at the top of the leaderboard

The benchmark results, drawn from public leaderboard data and a private evaluation by analytics firm Artificial Analysis, tell a striking story.

On GDPval-AA v2, a benchmark measuring real-world tasks across 44 occupations and 9 major industries, Kimi K3 scored 1,687 — placing it third overall, behind only Claude Fable 5 Max (1,815) and GPT-5.6 Sol Max (1,747.8), and ahead of Claude Opus 4.8 (1,600).

On AA-Briefcase, a private agentic benchmark from Artificial Analysis designed to test long-horizon knowledge work, K3 climbed to second place with a score of 1,527 — beating GPT-5.6 Sol Max (1,495) and trailing only Fable 5 Max (1,587).

Perhaps most impressively, K3 achieved a state-of-the-art score of 91.2 out of 100 on BrowseComp, a benchmark for long-horizon, high-difficulty information seeking.

The company says it accomplished this in a single-agent setup using its 1-million-token context window, without any context compression or additional context management techniques — a feat that suggests raw context length, when paired with strong retrieval capabilities, may be more powerful than elaborate multi-agent workarounds.

As one widely followed AI commentator put it on social media: “Open source is no longer lagging six months behind Western closed-source models. Read that again, and think about what it all means.”

That observation captures the significance of the moment. For much of the past three years, open-source models have typically trailed their proprietary counterparts by a meaningful margin. Kimi K3 appears to have closed that gap almost entirely.

How a 48-hour autonomous chip design demo reveals Moonshot’s real ambitions

Beyond raw benchmarks, Moonshot AI showcased a proof-of-concept that may be even more revealing of K3’s capabilities and the company’s strategic direction.

In a demonstration documented in the company’s technical materials, Kimi K3 was tasked with designing a physical chip to run a nano-scale version of itself. Over 48 hours of continuous autonomous agent operation, K3 independently completed the chip’s full construction pipeline — from architectural design through optimization and verification — using open-source electronic design automation tools. The result was a tiny but functional chip design, just 4 square millimeters, that achieved timing convergence at 100 MHz and could decode more than 8,700 tokens per second in simulation.

This is not a production chip. It is a demonstration of what Moonshot AI clearly views as the next competitive frontier: long-range autonomous agent capabilities. The ability to sustain coherent, multi-step technical work over a 48-hour window — reading documentation, making design decisions, running verification loops, and iterating on failures — represents a qualitative leap beyond the kind of single-turn question-answering that defined the first generation of large language models.

The company also highlighted a case in computational astrophysics, where K3 reportedly reproduced the universal I-Love-Q relation — a complex calculation that typically takes a senior researcher one to two weeks — in approximately two hours, reading and cross-validating more than 20 papers and implementing a complete numerical pipeline along the way.

Moonshot AI’s fall and rise tells the story of China’s brutal AI market

To understand why Kimi K3 matters, you need to understand where Moonshot AI was 18 months ago — and how far it fell.

Founded in 2023 by Yang Zhilin, a Tsinghua University graduate who previously conducted research at Google and Meta, Moonshot AI quickly became one of China’s most prominent AI startups. The company gained early traction in 2024 when users flocked to its Kimi platform for its long-text analysis capabilities and AI search functions. By early 2026, it had raised roughly $1.5 billion across multiple rounds, with its valuation climbing from $2.5 billion to $4.3 billion and the company reportedly seeking a new round at $5 billion.

Then DeepSeek happened. The release of DeepSeek’s low-cost R1 model in January 2025 disrupted the entire Chinese AI landscape, and Moonshot AI was among the hardest hit. Kimi, which had ranked third in monthly active users in China, slid to seventh. The company’s strategic pivot to open-source models — beginning with Kimi K2 in July 2025 and accelerating with K2.5 in January 2026 — was in large part an effort to reclaim relevance.

Kimi K3 is the culmination of that effort — and the sheer scale of the model suggests that Moonshot AI has been planning this move for some time. Training a 2.8-trillion-parameter model requires enormous computational resources and months of preparation, which means the architectural and infrastructure decisions behind K3 were likely locked in well before the model reached the public.

Why open-sourcing the world’s biggest model is a geopolitical chess move

The decision to release K3’s full weights on July 27 is strategically significant and worth parsing carefully.

The company’s own timeline chart of open-source frontier model scale positions K3 as a dramatic outlier, towering above competitors like DeepSeek (1.6T), Xiaomi (1.02T), and Alibaba (397B). By releasing the world’s largest open-source model, Moonshot AI is making a bid to become the center of gravity for the global open-source AI developer community.

This follows a broader trend among Chinese AI companies. As Reuters noted, open-sourcing allows companies to “showcase their technological capabilities and expand developer communities as well as their global influence, a strategy likely to help China counter U.S. efforts to limit Beijing’s tech progress.” DeepSeek, Alibaba, Tencent, and Baidu have all released open-source models. But none have released anything at this parameter count.

For enterprise technology leaders, the implications are concrete. A 2.8-trillion-parameter open-source model that performs at near-frontier levels creates new options for companies that want to fine-tune, self-host, or build proprietary systems on top of a capable base model — without being locked into API contracts with OpenAI or Anthropic. The trade-off, of course, is that running a model of this size requires substantial GPU infrastructure. Inference at 2.8 trillion parameters is not something that runs on a single server rack.

That said, Moonshot AI has signaled awareness of this challenge. Its Mooncake project, which won the Best Paper award at FAST 2025, pioneered KV-cache-centric disaggregated serving for large language models — an architecture designed specifically to make inference at extreme scale more practical and cost-efficient.

Kimi Code and a three-tier model lineup form the foundation of Moonshot’s enterprise play

Alongside K3, Moonshot AI continues to invest heavily in its coding agent ecosystem. Kimi Code, the company’s open-source coding tool that competes with Anthropic’s Claude Code and Google’s Gemini CLI, received two major updates on the same day as K3’s launch — versions 0.25.0 and 0.26.0 — adding features like expanded subagent tooling, background task management, and security fixes.

The Kimi Code CLI has accumulated over 3,100 stars on GitHub and features integration with VSCode, Cursor, and Zed. The latest release expanded the “coder subagent” tool set to include background tasks, todo lists, plan mode, skill invocation, and nested agents — effectively turning the coding agent into a multi-layered autonomous system capable of managing complex software engineering projects with minimal human intervention.

This is not incidental. Coding tools have become a critical revenue driver for AI labs. As Anthropic disclosed in January, Claude Code reached $1 billion in annualized recurring revenue. By building Kimi Code as an open-source alternative that defaults to Kimi’s own models — but supports other providers — Moonshot AI is positioning itself to capture developer workflows and, eventually, enterprise contracts.

The company’s model lineup now includes three tiers: K3 as the flagship ($3/$15 per million tokens for input/output), K2.7 Code as a specialized coding model ($0.95/$4), and K2.6 as a general-purpose option ($0.95/$4). All three support context windows of 256,000 tokens or above, with K3 offering the full 1-million-token window. Context caching is automatic — no cache ID, TTL, or extra parameter is required — a small but meaningful developer-experience advantage over competitors that require explicit cache management.

What Kimi K3 means for the future of enterprise AI and the global model landscape

Kimi K3’s release forces a recalibration of several assumptions that have guided enterprise AI strategy.

The performance gap between open-source and proprietary models has functionally closed at the frontier. If K3’s benchmark numbers hold up under independent evaluation — and particularly once the open weights are available for community testing on July 27 — it will be difficult for closed-source providers to justify premium pricing purely on the basis of capability.

The locus of AI innovation, meanwhile, continues to shift. China’s AI ecosystem, which many Western observers questioned after early struggles with chip export restrictions, has now produced a model that competes with the best systems from companies with direct access to Nvidia’s most advanced hardware. The architectural innovations behind K3 — particularly the hybrid linear attention mechanism — suggest that algorithmic efficiency may matter as much as raw compute.

And the agentic capabilities demonstrated by K3 — chip design, multi-week research compression, long-horizon information seeking — point toward a future where AI models are not just answering questions but autonomously executing complex, multi-day projects. For enterprises evaluating AI investments, this shifts the value proposition from “productivity copilot” to “autonomous technical workforce.”

Xinhua, China’s state news agency, framed the release as a national milestone, reporting that K3 “marks a new step forward in the development of China’s artificial intelligence models.” Liu Tieyan, dean of the Zhongguancun Academy in Beijing, was quoted as saying that a wave of Chinese open-source models has moved from isolated breakthroughs to collective advancement, providing “new solutions and new paths” for global AI development.

Just two years ago, Moonshot AI was a scrappy startup named for the audacious problems it hoped to solve. Eighteen months ago, it was a cautionary tale about how quickly a market darling can lose its footing. Today, it is the maker of the world’s largest open-source AI model — one that can, given 48 hours and an internet connection, design a chip to run itself. The frontier, it turns out, is not a place. It is a race. And the field just got a lot more crowded.

The desktop infrastructure problem that kubernetes finally solves

Presented by Kasm Technologies


Enterprise infrastructure teams have spent the better part of a decade pushing workloads into Kubernetes. Applications, APIs, batch jobs, data pipelines — if it runs in a container, it belongs in the cluster. The operational benefits are well-established: declarative configuration, horizontal scaling, self-healing, native integration with CI/CD pipelines and observability tooling. Kubernetes has become the default operating model for production workloads.

Except for desktops.

Secure desktop and application delivery — the kind that enterprises depend on for remote work, privileged access, and regulated-industry workflows — has remained stubbornly outside the Kubernetes model. Legacy virtual desktop infrastructure was built in a different era, for a different set of assumptions: pre-allocated VM pools, bespoke management planes, proprietary appliances, and operational tooling that has nothing to do with how modern platform teams work. The result is a split infrastructure reality: a modern, cloud-native application layer on one side, and a manually managed, operationally isolated desktop layer on the other.

That split is expensive. It means different tooling, different scaling behaviors, different observability approaches, and different operational runbooks. Platform engineers who are proficient in Kubernetes still have to context-switch into an entirely different mental model the moment a desktop infrastructure problem arises.

The more fundamental issue is that this split is unnecessary. Secure, containerized workspace delivery is a workload that Kubernetes is architecturally well-suited to run. Sessions are containers. Scaling is demand-driven. Configuration should be declarative. The only thing missing was a platform built to take advantage of that alignment.

Why the timing is right

The appetite for Kubernetes-native workspace delivery has grown significantly as organizations mature their container platform investments. Platform teams that have spent years standardizing on Helm, GitOps workflows, and Kubernetes-native observability are increasingly unwilling to make an exception for desktop infrastructure. The question has shifted from “can we run this on Kubernetes?” to “why isn’t this running on Kubernetes already?”

At the same time, the security case for containerized workspace delivery has become more urgent. Browser-delivered, containerized workspaces provide session isolation that VM-based desktops cannot match — each session is ephemeral, isolated at the container boundary, and terminates cleanly without persistent state. For organizations managing sensitive data, insider risk, or third-party access scenarios, this isolation model is a meaningful security control, not just a deployment convenience.

The convergence of these two trends — Kubernetes-native infrastructure expectations and containerized session security — creates a clear opportunity for platforms that can address both simultaneously.

What Kubernetes-native deployment looks like

A Kubernetes-native deployment uses Kubernetes as the control plane for workspace infrastructure — handling orchestration, scaling, and lifecycle management through the same declarative model used across the rest of the platform. Instead of relying on dedicated management appliances or pre-provisioned desktop pools, infrastructure is managed through the same CI/CD, GitOps, observability, and security workflows the platform team already operates. This gives platform teams a consistent operational model rather than maintaining a separate toolset for desktop infrastructure.

Kasm Workspaces, the browser-delivered workspace platform, is purpose-built to use Kubernetes as the control plane for workspace orchestration and delivery. Its deployment model is designed for real enterprise environments — not simplified demos — with production-grade Helm charts that follow Kubernetes conventions, tested upgrade paths between versions, and a standardized backend architecture validated across production deployments. An RDP Gateway component purpose-built for the Kubernetes topology enables Windows and Linux virtual machine access through the same platform.

Key capabilities include:

  • Horizontal session scaling driven by actual demand, orchestrated by Kubernetes — no pre-warmed VM pools required.

  • Declarative configuration through Helm values, enabling GitOps and CI/CD integration for workspace infrastructure.

  • Namespace-level isolation and compatibility with existing RBAC policies, ingress controllers, and secrets management integrations.

  • Metrics export for integration with Prometheus and existing observability stacks.

  • Rolling builds by default, reducing maintenance windows and enabling more predictable version management.

Real-world applications

Regulated-industry remote access. A financial services organization running a Kubernetes-based application platform can deploy Kasm into the same cluster, using the same operational tooling, to deliver isolated browser and application sessions to analysts and advisors. Sessions are ephemeral, network egress is controlled, and the entire deployment is managed through the same GitOps pipeline as their application workloads.

Contractor and third-party access. Organizations that regularly onboard contractors or external vendors — with the associated privileged access risk — can provision Kasm sessions on Kubernetes that scale up during engagement periods and scale back during low-demand windows. No persistent access. No VPN extension to external parties. Containerized isolation at every session boundary.

AI/ML development environments. Teams building and running AI models need GPU-enabled development environments with security controls that general-purpose cloud desktops rarely provide. Deploying Kasm on Kubernetes with NVIDIA MiG Multi-Instance GPU support lets platform teams deliver fractional GPU resources into isolated workspace sessions — giving data scientists the compute they need without shared-infrastructure security exposure.

The operational shift

The practical implication of a Kubernetes-native workspace platform is that platform teams can stop treating workspace infrastructure as a special case. The same engineers who deploy applications can deploy the workspace platform. The same pipelines that manage application configuration can manage workspace configuration. The same dashboards that monitor application health can monitor workspace health.

That operational consolidation reduces overhead, improves consistency, and eliminates the context-switching cost that has made desktop infrastructure a persistent pain point for cloud-native organizations.

For organizations still running legacy VDI alongside modern cloud infrastructure, the question is no longer whether a Kubernetes-native alternative exists. It does. The question is when to make the transition.

Organizations interested in evaluating Kubernetes-native workspace delivery can explore the platform at kasm.com and try out community edition for yourself.

Daniel Ben-Chitrit is the Chief Product Officer at Kasm Technologies.


Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.

OpenAI launches GPT-Live, a full-duplex voice upgrade that lets ChatGPT talk more like a person

OpenAI on Wednesday launched GPT-Live, a pair of new voice models that fundamentally redesign how people talk to ChatGPT — replacing the company’s existing Advanced Voice Mode with an architecture that can listen and speak simultaneously, much like an actual human conversation.

The two models, GPT-Live-1 and GPT-Live-1 mini, are rolling out globally starting today across iOS, Android, and ChatGPT.com. GPT-Live-1 becomes the default voice model for paid ChatGPT users on the Go, Plus, and Pro tiers, while GPT-Live-1 mini serves free-tier users. OpenAI also plans to bring the models to the API, and developers can sign up to be notified.

The release marks the third generation of ChatGPT’s voice technology in roughly two years — and OpenAI’s clearest bid yet to turn its chatbot into something that feels less like querying a search engine and more like talking to a colleague.

Why full-duplex voice changes everything about talking to AI

The defining technical advance in GPT-Live is what OpenAI calls a “full-duplex architecture.” In telecommunications, full-duplex means both parties on a phone call can talk and listen at the same time. Applied to AI, it means the model continuously processes your incoming audio even while it generates its own spoken response — no more waiting for a clean silence gap to figure out when you’ve finished a thought.

“Instead of processing a sequence of separate messages, GPT-Live continuously processes input while generating output,” OpenAI wrote in its research blog. “The model can therefore make interaction decisions many times per second: whether to speak, continue listening, pause, interrupt, or invoke a tool.”

In practice, that translates to a voice assistant that can insert conversational acknowledgments — “mhmm,” “yeah,” “got it” — while you’re still talking, pick up on a natural pause without jumping in prematurely, and handle rapid interruptions without derailing the entire exchange. 

OpenAI’s previous Advanced Voice Mode, launched to paid users in September 2024, processed and generated audio within a single model but still operated on rigid turn-by-turn exchanges. As OpenAI acknowledged in the announcement, “because turn detection is based on silence, even a brief pause or background noise could be mistaken for the end of turn — causing the model to interrupt at unnatural times.”

That brittleness created a product that, while impressive in demos, could be deeply frustrating in extended real-world use. Background chatter in a coffee shop could trigger a response. A thinking pause might get swallowed. The experience felt, as one researcher put it on X shortly after the announcement, like “walkie-talkie turn taking.” GPT-Live is designed to end that era.

How OpenAI split voice and intelligence into two separate layers

GPT-Live introduces a second structural change that may prove just as consequential for enterprise adoption: it decouples the voice interaction layer from the reasoning layer.

When a user asks a straightforward question, GPT-Live handles it directly. But when the query demands web search, deeper reasoning, or more complex agentic work, GPT-Live delegates the task to a frontier model running in the background — at launch, GPT-5.5, the large language model OpenAI released in April — and continues talking with the user while the computation happens asynchronously.

“While it works, GPT-Live can keep talking with you and maintain the flow of conversation,” OpenAI explains. “As we release new frontier models, we’ll continuously update the model used by GPT-Live.”

This delegation model is a meaningful architectural bet. Rather than building a single monolithic voice model that tries to be both conversationally fluid and deeply intelligent, OpenAI has split the problem in two: a voice-native model optimized for real-time interaction, and a separate reasoning engine that can be swapped out as the state of the art improves. 

It is, in effect, a modular design — one that allows OpenAI to upgrade the intelligence of its voice assistant without retraining the voice model itself. The implications for enterprise and developer workflows are significant. A voice agent built on this architecture could maintain a natural conversation with a customer while simultaneously querying databases, searching the web, or performing multi-step reasoning — tasks that would have introduced several seconds of dead air under the old pipeline.

The three generations of ChatGPT voice, from clunky pipeline to continuous stream

To understand how far voice AI has come, it helps to trace the three generations that led to GPT-Live.

The original ChatGPT Voice, launched in 2023, used a cascaded pipeline — a speech-to-text model (Whisper) transcribed what you said, a large language model (GPT-4) generated a text response, and a text-to-speech model converted that response back into audio. Each handoff introduced latency and lost information. 

As OpenAI noted, “the complexity came at a cost: information could be lost across models, and responses were slow and stilted.” That cascaded approach was the industry standard, and its limitations were well-documented. As the blog OpenHelm noted in an October 2024 analysis of OpenAI’s Realtime API, the old pipeline stacked up to roughly 1,700 milliseconds of latency — nearly two full seconds of dead air before the first word of a response. Managing the state between the three separate APIs consumed an enormous amount of engineering effort.

OpenAI’s Advanced Voice Mode, which began its limited rollout to paid ChatGPT Plus users in July 2024 before expanding more broadly in September 2024, collapsed that three-model pipeline into a single model that processed audio natively. As TechCrunch reported at the time, the rollout came with five new voices — Arbor, Maple, Sol, Spruce, and Vale — alongside improved accent handling and smoother conversations. 

The feature also launched on the web in November 2024, extending it beyond mobile. But Advanced Voice Mode still operated through discrete, alternating turns — and it launched into the shadow of a PR debacle that OpenAI is still working to leave behind.

The Scarlett Johansson controversy still shadows OpenAI’s voice ambitions

Advanced Voice Mode arrived in the wake of one of OpenAI’s most damaging self-inflicted crises. During the GPT-4o launch in May 2024, the company showcased a voice called “Sky” that many listeners immediately noted sounded strikingly similar to Scarlett Johansson, who famously voiced an AI companion in the 2013 film Her.

Johansson said she had declined OpenAI CEO Sam Altman’s offer to voice the system, then was “shocked, angered and in disbelief” when the product launched with a voice her own friends couldn’t distinguish from hers, as NBC News reported. Altman had tweeted just the word “her” the day the product launched.

OpenAI pulled the voice and apologized, but the incident drew public scrutiny from SAG-AFTRA and members of Congress, and crystallized broader concerns about AI companies moving fast with creative IP.

The Hollywood labor union said the issue underscored “why we’re strongly championing federal legislation that would protect their voices and likenesses … from unauthorized digital replication,” as NBC News reported. Forbes contributor Paul Tassi wrote at the time that Altman, “by holding up Her on a pedestal of something to strive for, has missed the point of that film” — in which the protagonist’s relationship with his AI companion ultimately does him more harm than good.

GPT-Live appears designed, in part, to move past those controversies. OpenAI says it has “remastered the nine distinct voices in ChatGPT for GPT-Live” and notes the system “is designed for conversation, not voice impersonation,” with “safeguards to prevent it from imitating a real person’s voice.”

What 150 million weekly voice users will actually notice today

OpenAI disclosed that more than 150 million people talk to ChatGPT using voice and dictation features each week — a notable slice of the platform’s 900 million total weekly active users. The voice experience has grown into a substantial product in its own right, used for language practice, bedtime stories, commute-time chat, and hands-free everyday help.

The new product features reflect that usage. GPT-Live introduces rich visual cards that surface during voice conversations — weather forecasts, stock data, sports scores, and maps — giving users something to glance at without breaking the flow of speech.

Users can now choose between three reasoning levels for answers: Instant for quick responses, Medium for moderate thinking, and High for more complex work. And if you take a moment to think, “ChatGPT Voice now waits instead of jumping in and interrupting,” OpenAI wrote. “If you ask it to stay quiet and listen, it will. And when there’s background noise, like passing traffic or nearby conversations, ChatGPT is better at focusing on your voice instead of getting distracted.”

Early reactions from users with preview access were cautiously positive. “I had early access to sol. it is a phenomenal model,” wrote one user on X, adding it is “much better at frontend, long context knowledge work, and its vibes are much better.” Another observer cut to the heart of the matter: “The smarts are not new here, GPT-Live hands hard questions to GPT-5.5. What is new is the feel: full-duplex voice that listens while it talks.”

New voice-specific safety tests reveal where the risks still live

The GPT-Live system card, published alongside the announcement, reveals a safety strategy built around the particular risks of real-time voice interaction — a domain where the speed and intimacy of conversation create hazards that text-based chat does not.

OpenAI expanded its safety evaluations to include audio-native tests, using both real user voice samples (from those who opted in) and synthetically generated prompts targeting edge cases across categories like self-harm, sexual content, illicit behavior, emotional reliance, mental health, and hate speech.

On the synthetic evaluations — which OpenAI described as deliberately adversarial — GPT-Live-1 showed substantial improvements over Advanced Voice Mode. In illicit behavior, for instance, the safety score rose from 0.63 to 0.97. On self-harm, it climbed from 0.72 to 0.98. Hate speech achieved a perfect 1.00, up from 0.87.

On the production-prompt evaluations — which used real user audio and reflected more ambiguous, borderline scenarios — the picture was more mixed. GPT-Live-1 matched or improved on Advanced Voice Mode in most categories but showed a slight regression on emotional reliance (from 0.88 to 0.82), though OpenAI noted the change was not statistically significant.

The company built real-time safeguards that can intervene while the model is speaking — steering toward safer responses, surfacing crisis resources, or ending the voice conversation entirely in higher-risk situations. It also designed additional protections for teen users and adapted self-harm support flows for voice, including crisis helpline integration.

Perhaps most notably, OpenAI said it is “rolling out longer-term measurement and post-launch monitoring focused on emotional reliance” — an acknowledgment that the very naturalness GPT-Live strives for creates its own category of risk.

Google, ByteDance, and Nvidia are already in the full-duplex race

While OpenAI was refining its safety guardrails, its rivals were shipping full-duplex systems of their own. Google’s Gemini Live, which supports full-duplex conversation alongside camera and screen sharing — capabilities GPT-Live notably lacks at launch — is already available in the Gemini app. Google released Gemini 3.1 Flash Live in March as its highest-quality real-time audio model, targeting low-latency voice interactions for developers.

ByteDance launched Seeduplex in April, claiming to be the first production-scale full-duplex speech AI deployed at scale, inside its Doubao app. Seeduplex reported roughly a 50 percent reduction in false-response and false-interruption rates compared to ByteDance’s previous half-duplex system. And Nvidia’s PersonaPlex, released in January, brought customizable voice and role control to full-duplex models, breaking what had been a constraint where natural-sounding models were locked into a single fixed voice.

The competitive picture is clear: full-duplex voice interaction is quickly becoming table stakes for consumer AI products, not a differentiator. OpenAI’s advantage lies in the scale of its existing user base, its integration with GPT-5.5’s reasoning capabilities, and the breadth of the ChatGPT ecosystem.

But the window in which any one company has a monopoly on natural-sounding voice AI has already closed. OpenAI also acknowledged several gaps. GPT-Live does not support voice with video or screen sharing at launch. Language support is limited, with the company noting that “for certain languages, the model may have a non-native accent or gaps in fluency.” And API access is not available on day one, meaning enterprise developers cannot yet build on GPT-Live directly — a constraint that will slow the model’s penetration into commercial voice-agent workflows where competitors like Google, ElevenLabs, and Deepgram already have developer-facing products.

The end of the chat box may be closer than anyone expected

GPT-Live is essentially OpenAI’s most significant bet yet on voice as the primary interface for AI — not just a convenience feature bolted onto a text chatbot, but a purpose-built interaction layer that sits between the user and the company’s most powerful models.

“Over time, we believe this research will also unlock the ability to use voice for increasingly complex, longer-running, and more agentic work,” OpenAI wrote. That ambition — using natural voice as the front end for autonomous AI agents that can perform multi-step tasks — is the logical endpoint of the full-duplex plus delegation architecture.

Imagine telling your phone to book a flight, negotiate with your insurance company, or debug a production server, all through a conversation that feels as natural as talking to an assistant who also happens to have the intelligence of a frontier AI model.

Two years ago, talking to ChatGPT meant dictating into a microphone and waiting nearly two seconds for a stilted reply. One year ago, it meant a smoother exchange that still felt like a polite, slightly awkward phone call with someone who insisted on waiting for you to finish every sentence. Today, it means something closer to a real conversation — imperfect, still constrained in some languages and missing video, but unmistakably closer. OpenAI once got into trouble for wanting to recreate the movie Her. With GPT-Live, the company may finally be reckoning with the harder question the film actually posed: not whether AI can sound human enough to talk to, but what happens to us when it does.

Z.ai launches ZCode to challenge Cursor, Claude Code and GitHub Copilot in AI coding

Z.ai, the Beijing-based artificial intelligence lab formerly known as Zhipu AI, on Wednesday officially launched ZCode, a free desktop application it describes as an “Agentic Development Environment” purpose-built for its flagship GLM-5.2 large language model. The move marks the company’s most aggressive push yet into the fast-growing AI-powered coding tool market, where it now competes directly with Cursor, Claude Code, GitHub Copilot, and Google’s Antigravity.

“Introducing ZCode, the official development environment for GLM-5.2,” the company wrote on X, noting the tool is available on macOS, Windows, and Linux, supports bring-your-own-key (BYOK) configurations for third-party models, and offers a 1.5x usage-quota bonus for subscribers to its GLM Coding Plan.

Read one way, ZCode is simply another entrant in a crowded market. Read another, it is a single product that crystallizes three of the most consequential trends in enterprise software today: the race-to-the-bottom pricing of frontier AI models, the geopolitical balkanization of the AI stack, and the rapid maturation of agentic coding agents into what Gartner now estimates is a roughly $10 billion market.

An AI coding tool designed to think in projects, not prompts

Unlike traditional IDEs that bolt on AI through a chat sidebar or autocomplete extension, ZCode is best understood as an agent-first development environment. Its core design is built around long-horizon tasks: the user describes an outcome, the agent plans the work, edits files, runs checks, reviews progress, and continues across multiple iterations until the goal is met.

ZCode organizes the development experience around the ZCode Agent, deeply tuned for GLM-5.2, with emphasis on deep integration: the model, tools, and execution workflow are tuned together so the Agent fits continuous, multi-step real-world development tasks. The environment supports continuous follow-up across devices: desktop, mobile Remote, and Feishu / WeChat Bot can all keep the same workspace task moving. Sensitive commands, file changes, and high-permission actions go through confirmation before execution.

That remote-control feature — the ability to steer a running coding agent from WeChat, Feishu, or Telegram on a phone — is a differentiator that speaks directly to the Chinese developer market, where those messaging platforms dominate professional communication. You can keep checking progress and adding instructions while long-running work continues, from any device with these messaging apps.

The tool is free to download. Revenue flows through Z.ai’s GLM Coding Plan subscription tiers, which start at $16.20 per month for a “Lite” plan and scale to $144 per month for “Max” — prices that undercut Anthropic’s Claude Code and Cursor’s comparable tiers by significant margins.

Through July 31, ZCode is offering a promotional 1.5x effective quota bonus for Coding Plan subscribers, with off-peak token consumption charged at a 0.67x coefficient. The platform also supports multiple AI models and agents, including Claude Code, Codex, Gemini, and OpenCode — a pragmatic concession to the reality that no single model wins every task.

GLM-5.2, the open-source model trained entirely on Chinese chips, powers the whole experience

ZCode’s value proposition is inseparable from GLM-5.2, the model it was designed to showcase. Z.ai released GLM-5.2 on June 16, first to its Coding Plan subscribers and subsequently as open-source weights under the MIT license on Hugging Face — a sequencing decision that prioritized distribution over the traditional benchmark-led launch.

The model’s specifications are formidable. GLM-5.2 is a 744-billion-parameter mixture-of-experts architecture with 40 billion active parameters, a genuine one-million-token context window — five times the 200K limit on its predecessor — and training on 28.5 trillion tokens. It ranked second globally on Code Arena as of mid-June, trailing only Anthropic’s Claude Fable 5, making it one of the highest-performing publicly available models for coding tasks.

Critically, the model was built entirely without American chips. As Decrypt reported, GLM-5.2 “runs entirely on Huawei silicon.” Stability AI founder Emad Mostaque estimated total training costs at roughly $25 million, with 80 percent spent on post-training — a figure that, if accurate, would make GLM-5.2 extraordinarily cheap relative to Western frontier models.

On benchmarks, GLM-5.2 performs within striking distance of the best proprietary systems. It trails Anthropic’s Claude Opus 4.8 by just one percentage point on FrontierSWE, a benchmark measuring multi-hour autonomous engineering projects, while edging out OpenAI’s GPT-5.5.

Its API pricing — $1.40 per million input tokens and $4.40 per million output — are a cost reduction of up to 82 percent compared to Anthropic’s Claude Opus 4.8 at $5 and $25, respectively. Because ZCode is a first-party tool from the same company that makes the model, it requires no manual endpoint configuration — the model is wired in.

The Anthropic export ban gave Chinese AI its biggest opening yet

ZCode’s arrival cannot be separated from the geopolitical drama that has roiled the AI industry over the past three weeks. On June 12, the U.S. government, citing national security authorities, issued an export control directive suspending all access to Anthropic’s Fable 5 and Mythos 5 models by any foreign national, whether inside or outside the United States, including foreign national Anthropic employees. Enterprise clients in finance, healthcare, SaaS, and critical infrastructure found their core intelligence services abruptly disabled, without exception, prior warning, or effective recourse.

While the Trump administration lifted those controls just yesterday — Anthropic confirmed on June 30 that the Department of Commerce had rescinded the directive — the episode sent shockwaves through the developer community and accelerated interest in open-source, self-hostable alternatives. The government’s crackdown on Anthropic coincided with a swift rise in Chinese open-source models that are proving to be almost as capable and significantly cheaper than some of the most powerful U.S. models.

Z.ai’s timing was surgical. On the same day the Trump administration ordered Anthropic’s most advanced models blocked for foreign nationals, Zhipu announced the open-source release of GLM-5.2 with no usage restrictions. The South China Morning Post reported that GLM-5.2 would be available to all users of Zhipu’s new GLM Coding Plan subscription, “priced at just a tenth of Anthropic’s premium Claude Code and Claude Max tiers.”

The market responded accordingly. Zhipu AI’s market capitalization crossed HK$1 trillion (US$128 billion) on June 22, driven by a 42 percent intraday share surge. JPMorgan raised its 2026–2030 revenue forecast for Zhipu by between 7 and 16 percent following the launch, projecting an over 534 percent revenue surge for 2026 and expecting the AI firm to turn a profit by 2028.

Why vendor lock-in now carries a geopolitical risk that no SLA can cover

The Fable 5 episode did more than embarrass Anthropic. It introduced a new risk category into enterprise AI procurement: sovereign access risk. When a government can disable a commercially deployed AI model overnight, the traditional evaluation criteria of developer experience, benchmark scores, and pricing become secondary to a more fundamental question: Will this tool still work tomorrow?

The event exposed the inadequacy of standard enterprise contract language. An investigation by FifthRow found that almost all standard Data Processing Addenda, SaaS agreements, and procurement SLAs “relied on vague ‘force majeure’ or ‘compliance with law’ catch-alls, not on precise, actionable regulatory suspension or kill-switch clauses.”

ZCode’s BYOK architecture and GLM-5.2‘s MIT-licensed open weights offer a partial answer. A development team can download the model, host it on its own infrastructure, and run ZCode against it without ever touching Z.ai’s cloud — eliminating both American export-control risk and Chinese data-sovereignty concerns in a single move. The catch is that anyone using Z.ai’s cloud API remains subject to Chinese law, a consideration that evaporates only with pure self-hosting.

Gartner analysts have warned that governance, pricing, support, workflows, commercial maturity, and market durability matter as much as developer experience and model capabilities when evaluating coding agent vendors for enterprise-wide adoption. By that measure, ZCode faces a steep climb. It is not open source itself; Linux support remains in beta; and security reviewers have flagged the need for careful evaluation of its credential handling, particularly for remote development over SSH and messaging-platform-triggered tasks — an agent that can be summoned from WeChat involves access paths that should be mapped before trusting it with anything sensitive.

Inside the $10 billion race where model labs are becoming full-stack IDE companies

ZCode enters one of the most crowded and fastest-moving markets in enterprise software. Enterprise AI coding agents are capturing a growing share of enterprise software engineering spend, with the market estimated at roughly $9.8 billion to $11.0 billion annualized as of April 2026, according to Gartner. A defining shift this year, the analyst firm noted, is “the movement of frontier model providers into direct competition with application-layer vendors” — precisely the pattern ZCode embodies.

Gartner codified this evolution in May when it renamed its annual Magic Quadrant from “AI Code Assistants” to “Enterprise AI Coding Agents,” defining the category as “autonomous or semiautonomous software engineering solutions that perceive context, translate human intent into multistep plans, and execute and verify those steps across code, tests and related engineering artifacts.” The 2026 Magic Quadrant names Anthropic, Cursor, GitHub, and OpenAI as Leaders. Z.ai was not among the 12 vendors evaluated — an absence that underscores both the company’s nascent enterprise sales presence outside China and the Western-centric lens through which the analyst community still views the market.

The competitive landscape is daunting. Cursor is the $2 billion ARR IDE that feels like VS Code with a supercharger. Claude Code reached approximately $2.5 billion in annualized revenue by early 2026. Google relaunched Antigravity 2.0 at I/O in May, and Cognition retired the Windsurf brand, relaunching the IDE as Devin Desktop with the Agent Command Center as the default surface.

Against these entrenched players, ZCode’s pitch rests on three pillars: deep first-party integration with GLM-5.2 that no third-party editor can replicate, aggressive pricing that starts at a fraction of Western competitors, and MIT-licensed open weights that allow enterprises to self-host — eliminating the regulatory kill-switch risk that the Fable ban made viscerally real.

Z.ai’s real challenge is turning a $128 billion valuation into a global developer tools business

Z.ai controls the model (GLM-5.2), the subscription layer (the GLM Coding Plan), and the IDE (ZCode) — a tightly coupled stack that optimizes for performance but concentrates switching costs. For the company, the business logic is clear. Its most reliable revenue stream has been on-premises deployments for Chinese government agencies, state-owned banks, and energy conglomerates. In full-year 2025, on-premises deployment revenue reached RMB 534 million, growing over 100 percent year-over-year and accounting for 73.7 percent of total revenue with a gross margin of 48.8 percent. ZCode and the GLM Coding Plan represent the company’s bid to build a comparable revenue engine in cloud-based developer tools — globally, not just in China.

The early signals are encouraging for Z.ai, if anecdotal. Community reception on X was enthusiastic, with one early user calling the tool “super stable” and others clamoring for more Coding Plan capacity. “Bro, can’t snag your family’s Coding Plan? When are you gonna stock up on more cards?” one user wrote in Chinese, suggesting demand is already outstripping supply.

But the hard questions loom large. Can a Chinese AI company build trust with Western enterprise buyers amid escalating technology tensions? Can ZCode’s ecosystem mature fast enough to compete with Cursor’s polished UX, Claude Code’s deep agent primitives, and GitHub Copilot’s unmatched distribution? And can Z.ai sustain a company valued at $128 billion while still losing money? 

What is no longer in question is the competitive dynamic itself. Three weeks ago, a U.S. government directive proved that access to the world’s best coding model can vanish overnight. Today, a Chinese lab is shipping a free IDE, an open-source model trained on zero American chips, and a subscription plan that costs less per month than a single lunch in Manhattan. The AI coding agent market did not just become global this summer. It became a market where the fallback option might be better than the thing it’s falling back from — and that changes the calculus for every engineering leader choosing a toolchain in the second half of 2026.

Claude Code turned every engineer into three. Now companies need more product thinkers

Anthropic recently told its growth team to hire more product managers, not fewer. The reason, as reported in industry coverage, was that Claude Code had quietly turned its engineering org into a team that ships at roughly three times its actual headcount, and the bottleneck moved from the integrated development environment (IDE) to the people deciding what to build.

That detail is easy to miss in the noise of every AI productivity claim. It is also the structural shift the rest of the industry is now living through. The bottleneck in software is no longer typing. It is deciding what to type. And the engineers who treat that as someone else’s problem are about to plateau.

For most of the last decade, that decision sat with someone else. Software engineering was a craft you absorbed slowly, then practiced in a long, predictable sequence: Dive deep on the technology, write the code, ask Stack Overflow when stuck, escalate to a senior engineer when Stack Overflow failed, ship the ticket. The product manager owned the funnel. The engineer owned the build. Both sides treated this division as physics.

Then the funnel collapsed in five steps.

A short history of how the engineer’s day got compressed

The Stack Overflow era (2014 to late 2022): The way engineers thought lived in one place. But new monthly questions on Stack Overflow are now down roughly 77% since November 2022, which was not coincidentally when ChatGPT launched. The drop is not a referendum on the site. It is a referendum on the workflow it represented.

The browser-tab era (late 2022 to 2024): The first ChatGPT generation sat outside the IDE. Engineers ran the same loop they had always run, just with a faster oracle: Write a prompt in a browser, paste the answer back into VS Code, repeat. The work was still single-threaded and engineer-driven. The leverage was real but local.

The IDE-native era (2024 to 2025): Cursor and Claude Code moved the model inside the editor and gave it access to the full repository. The senior-engineer escalation path largely dissolved. For years, the prevailing wisdom among veteran engineers was that Bash had the longest shelf life of any tool in the stack. By 2026, for a meaningful share of working developers, the first command typed in a fresh terminal is claude.

The spec-driven era (2025 to 2026): Larger context windows turned single-session work into something that previously required tickets, design docs, and sprints. Amazon’s Kiro IDE team reportedly compressed feature builds from two weeks to two days using the same spec-driven workflow they were shipping. An AWS engineering team described an 18-month rearchitecture, originally scoped for 30 engineers, was completed by 6 people in 76 days. The bottleneck stopped being how long it takes to write the code. It started being how clearly the team can describe what correct looks like.

The routines era (2026): In April, Anthropic shipped Claude Code Routines: Scheduled, persistent agents that run on a cadence, on a webhook, or overnight while the laptop is closed. Cron came back. Hooks came back. The engineer’s job is now part orchestration: Spin up a swarm before bed, review a stack of pull requests in the morning. Third-party wrappers like OpenClaw, which was briefly suspended by Anthropic in April before partial reinstatement, made the same point from the open-source side.

The bottleneck moved; most teams have not

Engineering has roughly tripled. Product management has not budged. The traditional 1:8 ratio of PMs to engineers, already strained, now plays out closer to an effective 1:20 because each engineer ships more per day. For instance, LinkedIn replaced its associate product manager track with a “Product Builder” program that trains generalists across product, design, and engineering. Anthropic is hiring more PMs, not fewer. The pattern is consistent across companies that have actually deployed agentic workflows in production: The system is producing built features faster than it is producing decisions about what should be built.

For engineers, this is the most important career signal of the decade, and the easiest one to miss while the productivity stories dominate the feed.

First principles matter more, not less

The instinct to declare fundamentals obsolete in the agent era gets the trend exactly wrong.

When a memory leak takes down production at 3 a.m., and the cause turns out to be a subtle ownership bug pushed 4 years ago, no agent currently in the wild closes that loop end-to-end. Operating systems, networks, concurrency, and query plans still decide who can resolve a real incident. They also decide who can spot the moments when an agent’s output looks correct on the surface and is quietly, expensively, wrong underneath. The agent that wrote 70% of the code in a modern repo cannot reliably tell anyone where its assumptions about thread safety, memory ownership, or transaction isolation diverged from the runtime. The engineer who can read the diff and catch that is the engineer the rest of the team needs in the room, and that engineer is built on fundamentals, not on prompting skill.

The corollary is that fundamentals are now a leverage skill, not a hygiene skill. In 2014, knowing how a TCP retransmit worked got a debug ticket closed faster. In 2026, the same knowledge keeps an entire agent-driven release pipeline from shipping a regression at scale. The blast radius of the engineer who knows what is happening underneath has gone up, not down.

Review is the new writing

Engineers in 2026 generate code at a rate that exceeds what any of them can read carefully. The team that ships fast and survives is the team whose engineers treat reviewing AI-generated code with at least the same rigor they once reserved for writing it. The 2025 Stack Overflow developer survey put 84% of developers on AI tools, with 46% saying they do not trust the output, up sharply from 31% the year before. That gap, heavy use paired with low trust, is exactly where review skills now matter most. Coders who push lots and review little are accumulating a debt that will come due during the first real incident, and the engineer who can pay it back is the one who paired their volume with deep first-principles knowledge of the systems involved.

The new differentiator is the product funnel

Both of those are necessary. Neither is sufficient. The engineer who matters in 2026 is the one who has stopped waiting for the funnel to arrive in the form of a Jira ticket.

That means doing things the role was historically allowed to skip.

Talk to customers. Watch how they actually use the product. Read the support queue. Sit in on the sales call. The signal a product team gets through three layers of summary, an engineer can now get firsthand in an afternoon.

Generate ideas, not just estimates. The product manager who used to source ideas for 8 engineers cannot source ideas for 20 at the same fidelity. The engineer who shows up with a validated, scoped opportunity is no longer doing the PM’s job. The engineer is doing the job the new ratio requires.

Work backwards from the customer. Amazon has been writing the press release first for two decades. The discipline travels well to teams of one and to swarms of agents. Both produce a great deal of working software in the wrong direction without a clear statement of what “customer wins” means before any code is written.

Stop hiding behind bandwidth. The honest answer to “Do you have capacity for this idea?” used to be ‘No.’ With routines, hooks, and a cooperative agent stack, the honest answer is closer to “What is the idea worth?” That is a different conversation, and a much harder one to have without a real point of view on the customer.

What the next decade rewards

The five-phase history above is not really a history of tools. It is a history of which part of the job a human had to do. The part that is still human, and that will remain human for the foreseeable future, has moved up the funnel: From typing, to reviewing, to deciding, to choosing the customer to serve and the problem to solve.

The 2026 version of a great engineer is not the one who writes the most code. It is the one who knows what to build, can prove it is worth building, and has the agent fleet plus the review discipline to ship it without the system collapsing under its own velocity.

Engineers who internalize this will spend the next decade doing the most interesting work software has ever produced. Engineers who wait for a ticket will spend it watching the ticket get written by the agent next to them.

Ishan Gupta is a software engineer at Amazon.

OpenAI unveils first custom AI inference chip, Jalapeño, with Broadcom — and its development was sped-up with OpenAI’s own models

OpenAI and Broadcom this morning unveiled their first custom AI accelerator chip named “Jalapeño,” positioning it is as a purpose-built processor for large language model (LLM) inference, rather than the more general GPUs offered by the likes of Nvidia or AMD.

According to its creators, Jalapeño is designed to support workloads behind ChatGPT, Codex, the API and future agentic products, though notably, both OpenAI‘s and Broadcom’s news releases position it as a product that could be made available to external AI firms as well — “built from the ground up for current and future LLMs across the industry.” [Emphasis mine.]

Jalapeño’s engineering timeline set a blistering pace for the semiconductor industry, moving from early schematics to fabrication readiness within a brief nine-month window, when new processor development cycles are typically measured in years. Indeed, the OpenAI and Broadcom partnership itself was only publicly announced in October 2025.

The companies attributed this speed to a deep software-hardware co-development process that actively used OpenAI’s own models to accelerate parts of the chip design. Greg Brockman, OpenAI’s president and co-founder and Broadcom CEO Hock Tan appeared on CNBC this morning to discuss the news, and Brockman noted in the interview that the development process relied on prior generation OpenAI models, not even the cutting-edge GPT-5.5, though a company spokesperson declined to specify exactly which when asked by VentureBeta.

After receiving an early physical model on Wednesday, OpenAI outlined plans to begin rolling out these processors across active data centers by the end of this year. OpenAI says it has already begun testing running at least one of its prior generation models, GPT‑5.3‑Codex‑Spark, on the chips at a production workload, though in a test environment.

The release marks a major strategic expansion for the ChatGPT creator as it attempts to build the full computational stack required to make advanced AI faster, more reliable, and more accessible.

There remain, of course, many outstanding questions — including how the new Jalapeño chip performs compared to direct competitors, its costs, and its manufacturing viability. Sources close to the company said the initial performance itself was (ironically): “outstanding.”

On X, Brockman himself wrote that “Perf[ormance] per watt looking incredible.”

Why OpenAI Built an ASIC

To understand why OpenAI is moving into chip design, it helps to look at the architecture. Jalapeño is an Application-Specific Integrated Circuit, or ASIC.

Unlike a GPU, which can handle many types of workloads, an ASIC is tuned for narrower uses, as industry experts note. That narrower focus can make it cheaper and more efficient for specific AI tasks, though less adaptable than Nvidia-style GPUs.

In Jalapeño’s case, OpenAI is starting from a clean design focused on modern LLM serving, instead of adapting a broader accelerator to fit its needs. The company says the architecture is shaped by its experience running large-scale AI products and is meant to reduce unnecessary data movement while better matching compute, memory and networking resources.

Broadcom is contributing core silicon implementation and networking technology, including Tomahawk networking silicon, while Celestica is helping with board, rack and system integration. The goal is to move the chip closer to its practical performance ceiling in real workloads, not just improve theoretical benchmarks.

However, OpenAI’s pivot into proprietary hardware is not just as a quest for technical supremacy: it may also make its core unit economics far more sustainable.

Audited financial documents posted recently by AI critic and AI public relations specialist Ed Zitron revealed that while OpenaAI generated an impressive $13.07 billion in revenue throughout 2025, its total operational expenses for the year ballooned to $34 billion, resulting in an operating loss of nearly $20.92 billion.

The primary culprit behind this cash hemorrhage involved pure compute requirements, though more is likely due to training than inference.

In 2025 alone, research and development costs—driven largely by the infrastructure required to train and serve massive language models—accounted for $19.18 billion, or approximately 56 percent of the company’s entire spending footprint. Furthermore, OpenAI reportedly paid Microsoft over $10.59 billion just for R&D and compute infrastructure last year.

Still, as OpenAI lays the groundwork for a heavily anticipated public offering in 2026, the Jalapeño inference chip may offer some reassurance to private investors and public markets that OpenAI has a plan for digging itself out of the financial hole and moving toward profitability. If it can drive down the costs of AI inference, then maybe it can recoup some of the losses spent on costly training runs.

“By designing more of the stack ourselves, we can serve more intelligence with greater efficiency and keep pushing advanced AI toward broader access,” said Brockman included in Broadcom’s release.

What Does This Mean for Nvidia and All of OpenAI’s Other Chip Providers?

The introduction of Jalapeño immediately raises questions about OpenAI’s strategic positioning within the fiercely competitive semiconductor and GPU market.

Since kicking off the generative AI boom in late 2022, OpenAI has remained one of the largest customers of GPU market leader Nvidia’s premium products, but has also taken billions in investment dollars from the firm (engendering accusations of “circular dealing”), and expanded to work with other rival chipmakers to fuel its appetites.

  • Nvidia: In February 2026, Nvidia finalized a $30 billion direct investment into OpenAI as part of a massive $110 billion funding round.This deal secured an agreement to deploy 10 gigawatts of computing systems—including 3 gigawatts of dedicated inference capacity and 2 gigawatts of training capacity—utilizing Nvidia’s next-generation Vera Rubin platform. Sources close to the companies tell VentureBeat Nvidia will remain central to OpenAI, particularly on the model training and development side.

  • Amazon Web Services (AWS): As part of the same February 2026 funding round, Amazon invested $50 billion into OpenAI. This deal included a commitment for OpenAI to consume approximately two gigawatts of AWS’s proprietary Trainium computing capacity over the next eight years.

  • Advanced Micro Devices (AMD): OpenAI signed agreements with Nvidia’s chief hardware rival, AMD for the former’s usage of the latter’s AMD Instinct™ MI450 Series GPUs.

  • Cerebras: The company also struck a pact with Cerebras, an AI chipmaker that executed its initial public offering in May 2026.

Sources with knowledge of these deals said at present, they currently remain in place, unaltered.

The Global Silicon Arms Race: OpenAI Joins AI Infrastructure Heavyweights

Before the introduction of Jalapeño, OpenAI operated at a distinct structural disadvantage compared to the world’s vertically integrated technology empires.

Tech giants like Google and Amazon have for years utilized their own mature custom silicon programs— Google’s Tensor Processing Units (TPUs) and Amazon’s Trainium lines—to serve massive computational workloads at drastically lower margins.

Microsoft, OpenAI’s primary cloud provider and single biggest financial backer, aggressively entered the bespoke silicon market by launching the Azure Maia 100 accelerator in late 2023.

Microsoft subsequently escalated this effort in January 2026 by introducing the Maia 200, an inference powerhouse built on TSMC’s 3-nanometer process that already actively powers OpenAI’s GPT-5.2 models within Azure data centers.

Similarly, Meta has aggressively expanded its Meta Training and Inference Accelerator (MTIA) portfolio in recent years, debuting the MTIA 300, 400, 450, and 500 series to power its recommendation engines and generative artificial intelligence features without relying solely on Nvidia.

Jalapeño provides OpenAI with the opportunity to match and offset the hyperscaler advantage. By baking its software architecture directly into a proprietary processor, OpenAI has the chance to replicate, at least in part, the playbook used by Google, Amazon, Microsoft, and Meta — transitioning from a captive cloud customer into a more independent AI infrastructure provider.

The timing is ripe amid a rapidly escalating global silicon arms race. Driven in part by United States export restrictions, Chinese tech heavyweights are pursuing more of their own custom AI chip hardware, too:

  • In May, Alibaba’s semiconductor division, T-Head, unveiled the Zhenwu M890, a proprietary processor expressly engineered for autonomous AI agents that require massive memory bandwidth and long-running context windows.

  • Huawei is reportedly gearing up to release its new Ascend 950DT chip next month

  • ByteDance, the corporate parent of TikTok, reportedly entered active negotiations with Qualcomm in June 2026 to design custom application-specific integrated circuits for its data centers to escape third-party dependency.

By successfully finalizing the Jalapeño design, OpenAI is seeking to move beyond the traditional confines of a software laboratory and stand shoulder-to-shoulder with international cloud and infrastructure titans.

The Gigawatt Future

This sprawling web of vendor agreements highlights the sheer scale of OpenAI’s infrastructural ambitions. The ultimate goal of the OpenAI and Broadcom partnership involves deploying gigawatt-scale data centers with Microsoft and other partners beginning in 2026 — that is, data centers with compute requiring energy on the order of cities.

For Broadcom, the partnership acts as a massive reputational catalyst. The company has been among the biggest beneficiaries of the generative AI boom, helping hyperscalers and frontier labs engineer custom silicon.

Broadcom shares reflect this momentum, demonstrating an 18% year-over-year increase in the first part of 2026 and a nearly 7X boost since the end of 2022, according to CNBC.

Ultimately, Jalapeño confirms that OpenAI believes it is ready to move beyond software and code into the realm of real-world, custom hardware.

By controlling the physics of its inference pipeline—while simultaneously leveraging the capital and hardware of Nvidia, Amazon, AMD, and Cerebras—OpenAI is attempting to rapidly rewrite its future unit economics of AI.

Anthropic launches Claude Tag, replacing its Slack app with a persistent AI teammate that learns, monitors and works autonomously

Anthropic on Tuesday launched Claude Tag, a new product that embeds its most advanced AI model directly inside Slack as a persistent, shared teammate that anyone on a team can delegate work to by simply typing @Claude.

The product, available today in beta for Claude Enterprise and Team customers, replaces Anthropic’s existing Claude in Slack app and represents the company’s most aggressive move yet to colonize the enterprise collaboration layer — the place where decisions get made, work gets assigned, and institutional knowledge accumulates in real time.

For enterprise technology leaders who have spent the past two years evaluating where AI fits into their operational stack, Claude Tag reframes the question entirely. This is not a chatbot, a coding assistant, or a search tool bolted onto a messaging platform. It is an AI agent designed to function as a standing member of a team — one that builds memory, takes initiative, works asynchronously, and interacts with every person in a channel rather than serving a single user. The implications for enterprise workflow, governance, and vendor strategy are significant.

Anthropic says 65% of its own product team’s code is now created by its internal version of Claude Tag, and the company runs internal support and data insight channels through the same system. The claim is striking: Anthropic is asserting that the majority of its own product engineering output already flows through the tool it just put in customers’ hands.

How Claude Tag works inside enterprise Slack channels

At its core, Claude Tag works like this: an administrator pairs it with a Slack workspace, grants it access to specific tools and data sources, sets spending limits, and defines which channels it can operate in. From that point on, any team member in those channels can tag @Claude with a request — write a pull request, pull sales numbers, run a data analysis — and Claude will break the task into stages, execute them using the tools it has access to, and respond in a Slack thread with the result. The product runs on Claude Opus 4.8, the model Anthropic released less than a month ago.

Four capabilities differentiate Claude Tag from its predecessors and from competing integrations. First, it is multiplayer. Within a given Slack channel, there is one Claude that interacts with everyone, not a separate instance per user. Anyone can see what it is working on, and anyone can pick up the conversation where the last person left off. This is a direct contrast to most existing AI integrations in Slack, which tend to operate as single-player tools.

Second, it learns over time. As Claude follows along with its channel, it accumulates context about the work happening there. Users do not need to re-explain projects from scratch. If granted permission, Claude can also pull context from other Slack channels and data sources, though Anthropic says it will not report from private channels. Third, it takes initiative. With ambient behavior enabled, Claude will proactively surface relevant information from across the channels it monitors and the tools it is connected to, and will follow up on threads or tasks that have gone quiet without resolution. This is a notable expansion of agency: Claude is not just responding to requests but monitoring the information environment and deciding what its human teammates need to know. Fourth, it works asynchronously, pursuing projects autonomously over hours or days. Anthropic says its own teams “now spend much more of our time delegating tasks to many Claudes in parallel.”

Enterprise security controls and administrative governance get a central role

Anthropic has designed the system with enterprise-grade isolation at its center. System administrators define separate Claude identities for different uses, scoped to specific channels with specific tools and data access. Everything, including Claude’s accumulated memories, stays within those boundaries. A Claude configured for sales work will not share memories or data access with one configured for engineering.

Administrators can set token-spend limits at both the organizational and channel level, and can review a complete log of every action Claude has taken and which user requested each task. For organizations managing compliance, audit, or regulatory requirements, this logging and scoping architecture is table stakes — and its absence has been a dealbreaker for many enterprises evaluating AI collaboration tools over the past year.

Migration from the existing Claude in Slack app requires an administrator opt-in within 30 days, and Anthropic says it is issuing introductory launch credits to eligible Enterprise and Team organizations. The four-step setup process — pair with Slack, connect tools, set spend limits, test in a private channel — is designed to reduce friction for IT teams already managing sprawling SaaS portfolios.

The Slack battleground is now the most contested real estate in enterprise AI

Claude Tag arrives in the middle of what has become the most fiercely contested territory in enterprise AI: the Slack channel. Slack itself has been aggressively positioning the platform as an “agentic operating system,” and the major AI players have responded by racing to plant their flags.

Salesforce, which acquired Slack for $27.7 billion in 2021, announced more than 30 new capabilities for Slackbot in March — the most sweeping overhaul of the platform since the acquisition — transforming it from a simple conversational assistant into a full-spectrum enterprise agent. OpenAI introduced “Workspace Agents” in April, allowing enterprise subscribers to design agents that take on work tasks across third-party apps including Slack, Google Drive, Microsoft apps, Salesforce, and Notion. Perplexity launched its enterprise “Computer” agent with direct Slack integration, letting employees query @computer directly inside Slack channels. Cognition’s Devin, the autonomous AI software engineer, has been built around Slack as a primary interface since its early days. Even Microsoft has brought GitHub Copilot into Teams.

The logic driving this convergence is straightforward: the average enterprise juggles over 1,000 applications, and employees waste countless hours on context switching, draining productivity by up to 40%. Whichever AI system becomes the default presence in the communication layer where work is coordinated gains an enormous distribution advantage — and, critically, an enormous data advantage. The AI that lives in the channel where work happens absorbs the institutional context that makes it increasingly difficult to replace.

Anthropic built Claude Tag on a foundation two years in the making

To understand Claude Tag’s strategic significance, it helps to trace the product arc that led to it. Anthropic first integrated Claude with Slack in October 2025, offering two-way connectivity: users could invoke Claude from within Slack or connect Slack as a data source for Claude’s chatbot. The initial integration was focused on individual productivity — direct messages, AI assistant panels, and thread participation. In January 2026, Anthropic expanded Claude’s Slack presence when it launched interactive Claude apps, which included workplace tools like Slack, Canva, Figma, Box, and Clay.

In parallel, Anthropic was building out its enterprise infrastructure stack. In August 2025, the company bundled Claude Code into enterprise plans, a move its product lead Scott White called “the most requested feature from our business team and enterprise customers.” In April 2026, Anthropic launched Claude Managed Agents, a suite of composable APIs for building and deploying cloud-hosted AI agents at scale, with early adopters including Notion, Rakuten, Asana, and Sentry.

Then came Claude Opus 4.8 in late May, which Anthropic described as “a more effective collaborator” with “sharper judgement, more honesty about its progress, and the ability to work independently for longer than its predecessors.” Benchmark improvements included a jump in agentic coding scores from 64.3% to 69.2% and a knowledge work score increase from 1753 to 1890. Claude Tag is the synthesis of all of these threads — combining the Slack channel presence, the enterprise security architecture, the Managed Agents infrastructure, and the Opus 4.8 model’s improved agentic capabilities into a single product that Anthropic frames as “the beginning of an evolution of Claude Code.”

Anthropic’s explosive growth explains why it is betting big on the collaboration layer

The financial stakes behind this launch are enormous. Anthropic raised $65 billion in Series H funding in late May at a $965 billion post-money valuation, and its run-rate revenue crossed $47 billion earlier this month. Claude Code’s run-rate revenue alone has grown to over $2.5 billion, more than doubling since the beginning of 2026, and enterprise use has grown to represent over half of all Claude Code revenue.

Those numbers explain why Anthropic is investing so heavily in channel-level presence. Every enterprise customer who grants Claude persistent access to a Slack channel — with connected tools, accumulated context, and ambient monitoring enabled — represents a dramatically deeper integration than a chatbot conversation or an API call. The usage patterns become stickier, the token consumption grows, and the switching costs rise. Deloitte’s deployment of Claude across more than 470,000 employees in 150 countries — reportedly its largest-ever enterprise AI deployment — illustrates the scale at which these dynamics play out.

The broader market trajectory reinforces the bet. Fortune Business Insights projects the global agentic AI market will grow from $9.14 billion in 2026 to $139 billion by 2034, and Gartner forecasts that 40% of enterprise applications will feature task-specific AI agents by 2026, up from less than 5% in 2025. Anthropic is not alone in seeing this future, but with Claude Tag it is making one of the most direct plays yet to own the enterprise agent layer.

The risks enterprise buyers need to weigh before granting Claude a permanent seat at the table

Claude Tag raises several questions that enterprise buyers will need to evaluate carefully. The first is vendor dependency. As VentureBeat reported when analyzing Claude Managed Agents earlier this year, once an organization’s agents, operational configurations, and monitoring run on Anthropic’s managed infrastructure, switching costs increase significantly. Claude Tag deepens this dynamic: a Claude that has accumulated months of channel context and institutional memory becomes very difficult to replace. Enterprise procurement teams accustomed to negotiating multi-cloud flexibility will need to think hard about what it means to give a single vendor’s AI persistent access to the communication layer where institutional knowledge lives.

The second is governance around ambient monitoring. The proactive behavior mode — in which Claude monitors channels and surfaces information it decides is relevant — represents a meaningful expansion of what enterprise AI systems do. Organizations will need to develop clear frameworks for an AI agent that is not just responding to requests but actively surveilling information flows and making editorial judgments about what humans need to know. For regulated industries, this raises questions that existing AI governance policies may not yet address.

The third is pricing. Anthropic has not published detailed pricing for Claude Tag beyond noting that it runs on token-based spending with administrative controls. For an agent that monitors channels continuously, builds memory, and works asynchronously over hours or days, the token consumption profile could look very different from traditional AI usage. And the fourth is reliability: Anthropic has been candid in recent months about infrastructure strain caused by surging demand, and for a product positioned as an always-on team member, downtime carries a different kind of cost than it does for a tool invoked on demand.

What Claude Tag signals about the future of enterprise work

Anthropic says its goal is to expand Claude Tag beyond Slack “so that teams can tag @Claude in the many other places they work.” The company is clearly eyeing the full collaboration surface — Microsoft Teams, email, project management tools, and beyond. If Claude Tag succeeds, it will validate a model of enterprise AI that looks less like a tool and more like a new category of worker: one that never sleeps, never forgets what was discussed in the channel last Tuesday, and never needs to be onboarded twice.

But the deeper significance of this launch may be what it reveals about the competitive dynamics reshaping enterprise software. For decades, the most valuable real estate in business technology was the system of record — the database, the CRM, the ERP. The current AI arms race suggests that the next era of enterprise value will be captured not by the system that stores the data, but by the agent that sits in the room where the work happens and understands what to do with it. Anthropic just gave that agent a name, a permanent seat in the channel, and permission to speak up when it thinks it has something to say. The question for every enterprise technology leader is no longer whether that agent will arrive. It is whether they are ready to manage it when it does.