Microsoft AI on Thursday released MAI-Transcribe-2, a speech-recognition model the company says is faster, more accurate, and cheaper than anything OpenAI, Google, or ElevenLabs currently sells. Then it priced the thing at 10 cents per hour of audio.
That figure deserves a pause. When Microsoft AI shipped the first model in this line just five months ago, it charged $0.36 an hour. Thursday’s early-bird price cuts that by roughly 72%. For an enterprise processing 100,000 hours of call-center audio a year — a modest volume for a large bank or telecom — the bill drops from $36,000 to $10,000. At that level, transcription stops being a line item anyone argues about.
The release arrives as Microsoft executes a strategy that would have seemed implausible two years ago: building its own frontier-class models one modality at a time, then steadily swapping them into products that once ran on OpenAI’s technology. Transcription is the modality where that plan has moved fastest, and MAI-Transcribe-2 is its clearest proof point yet. It also offers a preview of how the world’s most valuable software company intends to compete in AI without depending on the partner it spent $13 billion to cultivate.
The model transcribes audio in 60 languages, up from 43 in June’s MAI-Transcribe-1.5 and 25 in April’s original release. It runs on Microsoft Foundry, the company’s model marketplace for developers, and in MAI Playground, its testing environment. Microsoft says it built the model for the messy audio that real businesses generate — background noise, low-quality recordings, overlapping speech — rather than clean studio conditions.
More important than the language count is what Microsoft has bundled into the base product. Speaker diarization sorts out who said what in a multi-person recording, which is the difference between a wall of text and a usable meeting transcript. Word-level timestamps attach a precise time marker to every word, enabling search, editing, and alignment with video. Keyword biasing lets developers feed the model a list of drug names, product codes, or employee names so it stops mangling domain jargon. Automatic language identification means users no longer have to declare the language in advance.
Two features stand out for their specificity. A configurable output style offers a “verbatim” mode that preserves every “um,” false start, and stutter for compliance and legal teams, alongside a “clean” mode that strips fillers for readable captions and notes. And code switching handles conversations that drift between languages mid-sentence; Microsoft explicitly names Hinglish and Spanglish, a nod to the Indian and U.S. Hispanic markets where a single customer-service call might toggle languages a dozen times. Specialty vendors have historically charged premiums for each of these capabilities. Microsoft is including all of them for a dime.
Microsoft makes three performance claims, each resting on a different measuring stick, and technical buyers should understand what each one captures and what it misses.
The first is that MAI-Transcribe-2 ranks number one on FLEURS across 60 languages with an average word error rate of 5.2%. FLEURS is a benchmark Google researchers published in 2022, built from native speakers reading roughly 2,000 sentences in each of 102 languages — about 12 hours of speech per language. It is the standard yardstick for multilingual speech recognition because it lets you compare a model’s Swahili against its Swedish on identical content. Word error rate, its metric, simply counts substitutions, insertions, and deletions against a human reference; 5.2% means roughly one word in 20 is wrong. But FLEURS is read speech, not conversation, and Microsoft’s average has actually risen from the 3.7% it reported for MAI-Transcribe-1.5 in June. That almost certainly reflects broader coverage rather than regression — averaging across 60 languages instead of 43 means folding in low-resource languages where every model struggles — but buyers should request the per-language breakdown.
The second claim is that the model ranks second on the Artificial Analysis word-error-rate leaderboard and defines that firm’s accuracy-latency Pareto frontier. Artificial Analysis is an independent benchmarker that tests models through their public APIs, measuring what a customer actually gets. Its index blends simulated agent conversations, European Parliament speeches, and corporate earnings calls, weighting heavily toward English business speech. In June, the firm ranked MAI-Transcribe-1.5 third at 2.4% WER, behind Alibaba’s Fun-Realtime-ASR-preview and ElevenLabs’ Scribe v2, while calling it the fastest model in the top 10. Climbing to second suggests Microsoft has cleared ElevenLabs. “Pareto frontier” is the phrase practitioners should note: it means no rival beats the model on accuracy without being slower, and none beats it on speed without being less accurate.
The third claim is raw speed — 10 times faster than OpenAI’s GPT-Transcribe, seven times faster than ElevenLabs’ Scribe v2, five times faster than Google’s Gemini 3.5 Transcribe, per Artificial Analysis evaluations. In batch transcription, speed matters less because anyone is waiting and more because throughput is cost. A model running at 300 times real-time needs a fraction of the GPU-hours of one running at 30 times. That efficiency is what lets Microsoft charge a dime and, presumably, still make money.
The pace is the story within the story. On April 2, MAI-Transcribe-1 launched with 25 languages at $0.36 per hour. On June 2, MAI-Transcribe-1.5 arrived with 43 languages, keyword biasing, and a third-place ranking on Artificial Analysis. Today, MAI-Transcribe-2 shipped with 60 languages, diarization, timestamps, code switching, a second-place ranking, and a price of $0.10.
Three releases in five months, each expanding language coverage by roughly 40% while adding features competitors gate behind premium tiers. That cadence is characteristic of a team that has settled on a stable architecture and is now turning the crank on data and scale — the phase where speech models tend to improve quickly and predictably. It is also the cadence of a company that intends to make transcription a commodity before anyone else can.
The organizational bet behind that speed is one Mustafa Suleyman, Microsoft AI’s chief executive, described to The Verge in April. He credited the first model to “a small, focused 10-person team” that had been “liberated from any of the bureaucracy,” with a larger surrounding group handling vendor management and data acquisition.
He also told The Verge the model ran at “half the GPU cost of the other state-of-the-art models,” calling it “a huge cost-saving” for Microsoft. Meta, Amazon, Google, and Anthropic have all experimented with similar flattened structures, The Verge noted. Microsoft’s transcription line is the most visible test yet of whether the approach produces commercial results rather than research papers.
Microsoft has invested more than $13 billion in OpenAI, and hosts OpenAI’s models across Azure, Office, and Copilot. For most of the past four years, the obvious question about any Microsoft-built model has been: why bother? The answer has sharpened over the past year, and it begins with independence.
When Microsoft hired Suleyman from Inflection AI in March 2024, along with most of Inflection’s staff, Salesforce CEO Marc Benioff read it as a declaration of intent. “Microsoft is building their own AI and I don’t think Microsoft will use OpenAI in the future. They’ll have their own frontier models,” Benioff told CNBC in January 2025. “That’s why they hired Mustafa Suleyman.” Benioff had his own motives — Salesforce competes with Microsoft and invests in Anthropic — but events have largely borne him out.
In October 2025, Microsoft and OpenAI restructured their partnership in a deal that, per Microsoft’s own announcement, allowed Microsoft to “independently pursue AGI alone or in partnership with third parties” for the first time. Suleyman told The Verge that renegotiation “unlocked [Microsoft’s] ability to pursue superintelligence,” and Microsoft announced its MAI Superintelligence team weeks later. In April 2026, the companies amended the deal again, ending Microsoft’s exclusive access to OpenAI’s models and eliminating Microsoft’s revenue-share payments, according to reports at the time. Each amendment loosened the tie. Each one was followed by more MAI models.
The second half of the answer is margin. Every prompt Microsoft routes to an OpenAI model carries a cost. Every prompt it routes to its own model on its own GPUs carries a smaller one. In July, Bloomberg reported that Microsoft had begun using MAI models to answer a portion of user prompts in Word and Excel — products it had previously advertised as powered by OpenAI and Anthropic. TechCrunch framed the shift as part of a broader industry pullback on AI spending, with Amazon, Uber, Meta, and Accenture all reportedly trimming.
Transcription is the natural first target for this substitution because the problem is bounded and the metric is objective. Microsoft owns Teams, which generates an enormous volume of meeting audio. It owns Nuance, whose clinical documentation business runs on speech recognition. It owns the Azure speech services that thousands of enterprises already call. Every one of those workloads is a candidate to move onto MAI-Transcribe-2, and every hour that moves is an hour Microsoft no longer pays anyone else for.
Suleyman has been unusually candid that this is the point. Superintelligence, he told The Verge in April, “is really about, ‘Are these models capable of delivering product value for the millions of enterprises that depend on us to deliver world-class language models?'” Whatever one thinks of applying the word “superintelligence” to a transcription API, the commercial logic is plain: build the capability once, deploy it across a dozen products, and stop writing checks to a partner that is increasingly a competitor.
Microsoft’s release names four rivals: OpenAI’s GPT-Transcribe, Google’s Gemini 3.5 Transcribe, OpenAI’s older Whisper V3-Large, and ElevenLabs’ Scribe v2. It does not mention Deepgram, AssemblyAI, Speechmatics, or Rev — the specialists that have sold transcription to enterprises for a decade. Microsoft is positioning against the frontier labs, not the incumbents.
That framing is partly marketing and partly true. The frontier labs have treated speech as a checkbox feature of broader platforms, priced accordingly, and a dedicated model that beats them on speed by five to 10 times while matching their accuracy is a genuine differentiator. But the specialists will feel the price pressure most acutely. At $0.10 an hour, Microsoft is pricing at or below where many of them sell high-volume enterprise contracts, and it is bundling diarization, timestamps, and 60 languages into the base rate. The specialists’ remaining moat is domain depth — medical vocabularies, legal formatting, industry-specific integrations — and Microsoft’s keyword biasing feature is aimed squarely at it.
The one competitor Microsoft conspicuously does not claim to beat on accuracy is Alibaba, whose models have posted leading numbers on independent leaderboards for much of 2026. TechCrunch reported in July that some U.S. companies had begun evaluating Chinese models as cheaper alternatives despite security concerns. Microsoft’s pitch to those buyers is implicit but unmistakable: comparable accuracy, faster inference, lower price, and a vendor your compliance team already trusts.
For all its specificity on benchmarks, the release leaves several practical questions open. The first is duration: Microsoft calls $0.10 per hour a launch offer without naming an end date or a standard rate, and anyone building a cost model should get both in writing. The second is streaming. The release emphasizes batch throughput and long-form audio but says nothing about real-time transcription, which voice agents and live captioning require. Artificial Analysis maintains a separate streaming leaderboard, and Microsoft’s silence on it is notable.
The third is per-language accuracy. A 5.2% average across 60 languages could mean 3% on major languages and 12% on low-resource ones, so buyers with specific needs should test those languages directly. The fourth is diarization quality. Word error rate does not measure speaker attribution; a transcript can have near-perfect WER and still assign every other sentence to the wrong person. The release offers no diarization error rate or comparable metric.
The fifth is data handling. Enterprise transcription touches medical records, legal privilege, and financial disclosures, and the release says nothing about data residency, retention, or whether audio submitted to Foundry feeds future training. Microsoft’s April announcements described training data as a mix of human-curated recordings, contractor-recorded noisy audio, and “vast amounts of data from the open web,” per The Verge — a description that should prompt pointed questions from regulated industries. None of these gaps is unusual for a launch announcement, but they are exactly the questions that separate a leaderboard win from a production deployment.
Step back from the speech-recognition details and a pattern emerges that extends well beyond transcription. Microsoft’s AI unit now ships models for images, voice, transcription, code, reasoning, and cybersecurity. At Build in June, it announced seven new MAI models in a singlekeynote. Each follows the same playbook: target a well-defined modality, optimize aggressively for inference cost, price below the frontier labs, distribute through Foundry, and quietly swap the model into Microsoft’s own products.
This is not an attempt to build one model that beats GPT or Gemini at everything. It is an attempt to build a portfolio of specialized models that, in aggregate, let Microsoft serve most of its enterprise workloads without paying anyone else — and to sell the surplus capacity to everyone else at prices the specialists cannot match. Transcription happened to be the first modality where the approach fully matured, but the release notes for MAI-Transcribe-2 read less like a product announcement than a template.
Suleyman has spent two years talking about “humanist superintelligence” and AI assistants that are “accountable to them, on their side.” The vocabulary is lofty. The execution is a spreadsheet. Five months ago, Microsoft charged 36 cents to turn an hour of speech into text. On Thursday it charged a dime, threw in six features its rivals sell separately, and claimed the top spot on the industry’s standard multilingual benchmark. The company that spent $13 billion learning what frontier AI costs has decided it would rather own the factory than rent the output — and now it is selling the output for less than the rent.
MAI-Transcribe-2 is available now through Microsoft Foundry and MAI Playground.
Perplexity is launching Portable Computer today, a version of its agentic “Computer” platform that runs entirely on hardware users already own — starting with Nvidia’s DGX Spark desktop supercomputer and Linux machines equipped with Nvidia RTX GPUs.
The launch, developed in close partnership with Nvidia, is one of the most aggressive attempts yet to move serious AI agent workloads off the cloud and onto local devices. The model, the user’s files, and the work itself can all stay on the machine. Work completed locally consumes no billing credits, and the company says every task starts on the device by default — with the system asking permission before sending any individual step to a more powerful frontier model in the cloud.
“We’ve basically brought the exact same UI to a fully local app,” said Nate, Perplexity’s vice president of engineering for infrastructure and enterprise, during a press briefing Monday. “This incorporates the entirety of the agent harness and inference and everything needed to do work locally.”
For Nvidia, which has spent the past two years selling the world on trillion-dollar AI data centers, the announcement signals something subtler but strategically important: the chipmaker believes local AI has crossed a threshold from hobbyist curiosity to practical tool — and it wants to sell the hardware that runs it.
“Local AI reached an inflection point,” said Nader, Nvidia’s director of developer technology, who focuses on developer tooling and open source. “For the longest time, it was hobbyists and enthusiasts, and they were running these quantized models that were quantized down to be super tiny… And while that’s cool, it’s not super practical. But all that changed with a lot of these new open source models that have come out that are super useful.”
Perplexity Computer, the company’s agentic platform for knowledge work, orchestrates AI models, files, tools, and web access to complete multi-step tasks — reviewing folders of documents, analyzing data, producing reports, and pushing results into business systems. Portable Computer replicates that experience locally: the local models, agent harness, inference engine, tools, app connectors, and a security sandbox come packaged together in a single system. That bundling is the point. With most local AI stacks today, users must assemble and operate those pieces separately — downloading model weights, standing up an inference server, wiring together tools, and tuning performance.
“Historically it’s just been really painful to bring up the local AI stack,” Nate said. “With Portable Computer, we really focused on just making this a really straightforward experience where you can get up and running very quickly.”
In one demo Monday, the system played the role of a retail investor reviewing a folder of 1099s and investment documents — the kind of sensitive financial material many users would hesitate to upload to a cloud service. Running a 27-billion-parameter Qwen model at full GPU utilization on a DGX Spark, the agent reviewed each document and flagged cases where the hypothetical investor was paying unnecessary fees. The interface element that normally displays a running tally of cloud credits “is just parked at zero,” Nate noted, “because all of this is happening on the device.”
A second demo showed the hybrid side of the product. Playing a startup founder, Nate asked the agent to analyze a CSV of user funnel data locally, then push the finished analysis to a Slack channel using Perplexity’s connector ecosystem — proof that local-first does not mean disconnected.
The system also connects to Google Drive, Gmail, and GitHub, and can escalate to a frontier cloud model when the local model hits its limits. At launch, users can set up Qwen 3.8 27B or PPLX 27B, a version Perplexity has post-trained on its own harness, with Nvidia’s Nemotron 3.5 Lightning coming soon.
Portable Computer arrives today for Pro, Max, Enterprise Pro, and Enterprise Max subscribers on Linux, with Windows support following in September. Any RTX GPU with at least 24GB of VRAM — roughly a GeForce RTX 3090 or newer — clears the bar, a threshold Nate called “sort of the floor where we really want to make sure that we can deliver a great experience, but balance that with making it broadly available.”
Alongside the launch, Perplexity published a research paper arguing that effective local agents require the model and the agent harness — the scaffolding of prompts, tools, and orchestration logic around the model — to be designed together. The core insight: general-purpose harnesses assume a frontier model that can absorb enormous contexts, navigate sprawling tool surfaces, and plan over long horizons. Small local models buckle under those demands.
Perplexity found empirically that although models like Qwen 3.8 27B advertise 260,000-token context windows, they begin to struggle beyond 100,000 tokens. So the company built a deliberately minimal harness: a succinct system prompt, a small set of core tools, and capabilities that load and unload as on-demand “skills” rather than sitting permanently in context. It converted popular connectors like Gmail and GitHub from token-hungry MCP servers into compact command-line tools, added self-verification hooks that monitor the health of a task, and enforced always-on OS-level sandboxing. If the sandbox is unavailable, the harness disables itself rather than running tools unprotected — a contrast with open-source harnesses that run commands with the user’s full permissions by default.
The benchmark results Perplexity reports are striking, though they come from the company’s own evaluations. On its internal Local Knowledge Work Bench — 53 tasks spanning deep research, financial analysis, and document creation, which Perplexity says it plans to open-source — Computer running Qwen 3.8 27B on a DGX Spark scored 82.6%, versus 77.6% for the open-source Pi harness and 74.0% for Hermes running the identical model.
Perplexity’s post-trained PPLX 27B pushed the score to 85.4%. The gaps widen dramatically on harder tasks: on BrowseComp, a web research benchmark, Computer hit 66.7% accuracy versus 50.2% for Pi and 43.9% for Hermes, while using 51% less wall time and 70% fewer tokens than Pi. On multimodal document understanding, Computer scored 65.1% against Hermes’ 34.6% and Pi’s 13.9%.
The strategic logic behind the launch becomes clear when you consider how AI workloads have changed. Chat was bursty — a question, an answer, done. Agents are different.
“With agents, you want these agents always on if you can. You want the agents to really consume as many tokens as they can,” Nader said. “What we’re seeing is an insatiable demand for tokens, and that’s something that makes local AI so great. As you saw through all these demos, you were not metered by the token. You were not paying for the token. So it’s really killer for agents.”
This reframes the value proposition of local hardware. An agent that runs for hours reviewing documents, verifying its own work, and iterating on analyses would rack up substantial API bills in the cloud. On a device the user already owns, the marginal cost of those tokens approaches zero. Perplexity’s paper makes the enterprise version of this argument explicitly: as agents scale across individual workflows and entire organizations, token expenditure and data movement “become increasingly difficult to govern.” Local-first execution addresses both at once — spend, because inference is free, and privacy, because sensitive tokens never leave the device boundary.
Perhaps the most commercially interesting result concerns the hybrid middle ground. On Terminal Bench 2.1, a challenging coding benchmark, the fully local Qwen model scored 59.6% at essentially zero marginal cost. Letting it escalate to a Claude Opus 5 “advisor” in the cloud raised the score to 73.0% at an estimated $0.415 per task. Running the frontier model alone scored 82.4% at $0.65 per task. Escalation, in other words, recovered roughly three-fifths of the gap to frontier performance at about two-thirds of the cost — and the user decides when that trade is worth making. Before any advisor call, the harness runs a PII classifier over the outgoing context and shows the user exactly what would leave the device. The remote model returns text guidance only; it never touches local files or tools.
Jason Hiner of The Deep View pressed the companies on how Portable Computer relates to existing local inference tools like Ollama. Nate’s answer drew a clear line: the tools solve different layers of the problem.
“The majority of the effort here has been at the agent harness level,” he said, noting that the system uses vLLM to host model inference underneath, with an advanced mode for users who want to plug in their own inference endpoint. “We’ve heavily post-trained both the Qwen and Nemotron models that we’re working with in order to really get the best possible results… Our focus has been on really honing the whole stack, top to bottom, of the model inference and the harness together.”
Nader put it more colorfully. “Just getting inference running really quickly on a Spark — there’s a smooth path. You can use Ollama. You can get that set up. But then, as you start to do more complicated, more agentic things, then suddenly you need more perf. You start looking at different models. You start looking at different harnesses, and it’s kind of like the ocean. The deeper you go, the deeper it gets.”
The appliance-like pitch appeared to land with at least one attendee. Ben, who described struggling to set up his own DGX Spark despite being an engineer — “this experience sucks, we have to fix it” — said the product feels like the unlock “needed for people to really feel and understand what agentic means, and you need the right UX to make it happen.” Nvidia also emphasized that the hardware scales: connecting two Sparks over shared memory runs frontier-class open models like DeepSeek’s latest, and four can run GLM 5.2 or Nemotron Ultra. “I’ve even seen eight Sparks get connected,” Nader said.
The launch extends a partnership that has been building for more than a year. In June 2025, Nvidia and Perplexity announced a collaboration to bring sovereign AI models to European publishers and telecoms, part of CEO Jensen Huang’s continent-hopping campaign to convince governments that, as the Associated Press reported from VivaTech in Paris, “every country needs a national intelligence infrastructure.” The sovereign AI pitch — that data “belongs to your people, your country, your culture,” in Huang’s words — is philosophically the same argument Portable Computer makes at the scale of a single desk: intelligence you control, running on hardware you own.
There is a self-interested logic for both companies. Perplexity, which has raised capital at steadily escalating valuations while facing legal pressure from publishers over its content practices — including a lawsuit filed by The New York Times in December 2025 and an earlier public dispute with Forbes — gets a product whose economics don’t depend on metering every token, and a differentiated wedge into privacy-sensitive enterprises in law, healthcare, and finance. Nvidia gets a killer app for DGX Spark, a device that, by the admission of attendees at Monday’s briefing, has been easier to buy than to use. When one reporter asked whether a Spark might ship with Portable Computer and a Nemotron model preinstalled, Nader demurred without ruling it out: “That would be cool… the goal is just making sure that it’s a super smooth experience for every user.”
Questions remain. Perplexity’s most impressive numbers come from its own internal benchmark, and the company acknowledges that compact models still trail the frontier meaningfully on hard reasoning tasks — advisor escalation “narrows but does not fully close the gap.” The launch is Linux-only for now, the 24GB VRAM floor excludes the vast majority of consumer PCs, and Apple silicon — home to some of the most enthusiastic local AI tinkerers — is conspicuously absent from the roadmap. “We’re very focused right now on Nvidia hardware,” Nate said when asked.
But the direction of travel is unmistakable. Perplexity’s researchers describe the launch as part of “a broader shift in which increasingly capable agents move from remote infrastructure to individual and local devices,” and both companies are betting that advances in chips and open models will keep expanding what a box on a desk can do. During Monday’s demos, the most telling detail wasn’t a benchmark score — it was that credit counter in the corner of the screen, sitting motionless at zero while the agent churned through a folder of tax documents. For two years, the AI industry has measured its ambitions in gigawatts and tokens per dollar. Portable Computer proposes a different meter, one that never runs.
IBM is announcing today at the annual Hot Chips conference what may be the most consequential change to mainframe architecture in decades: a processor whose cores can natively execute both IBM’s own instruction set and Arm’s — switching between the two in nanoseconds.
The chip, which will power the next generation of IBM Z and LinuxONE systems, is the first dual-architecture mainframe processor ever built. It is designed to let enterprises run the vast and fast-growing ecosystem of Arm-native Linux software — including the AI frameworks that increasingly define modern infrastructure — directly alongside the z/OS transaction-processing workloads that anchor the world’s banks, insurers, and governments.
“As technology enthusiasts on both sides, we’re really excited about being what I would consider one of the most powerful commercially available processors that’ll be dual architecture,” Tina Tarquinio, chief product officer for IBM Z and LinuxONE, told VentureBeat in an exclusive interview ahead of the announcement.
The announcement marks the first hardware milestone from the strategic collaboration IBM and Arm unveiled in April, and it offers an unusually direct answer to a question that has shadowed the mainframe for years: can the machine that processes most of the world’s regulated financial transactions remain a first-class citizen in an AI era built largely on other people’s silicon?
The most striking engineering decision is what IBM chose not to do. The company could have bolted a handful of standalone Arm cores onto the side of its processor — a simpler design that other chipmakers have used for heterogeneous computing. Instead, IBM built every core on the chip to be bilingual.
“On this chip are 11 cores, and each core can dynamically switch back and forth between Arm software mode and traditional Z software mode,” Jacobi explained in an exclusive interview with VentureBeat. “That enables us to run the mission-critical enterprise software right next, on the same chip, to the much broader software ecosystem of Arm applications.”
The mechanism relies on the open-source KVM hypervisor. Enterprises can run Arm64 Linux virtual machines and Linux on Z virtual machines side by side, and as the hypervisor dispatches each virtual machine onto a physical core, the core flips into the corresponding mode. The performance penalty, Jacobi said, is effectively zero. “That switch takes about the nanosecond scale,” he said. “Because you’re running for many milliseconds in the virtual image, this switching overhead sort of amortizes to zero — pretty much no impact at all.”
Traditional z/OS workloads run in a separate partition on the same chip, outside KVM — meaning a bank’s core ledger, its fraud models, and a modern Arm-native monitoring stack can all share the same silicon, the same memory fabric, and the same reliability guarantees. Jacobi was candid that IBM debated the easier path and rejected it. “We’re really not addressing their need if we just have a few, I’d say, loosely Arm cores in the corner of the chip,” he said. “It really needed to be deeply integrated into the entire system design for it to have the same qualities of service that clients are used to.”
The specifications underscore that this is no compromise design. Built on a leading-edge 2-nanometer process node, the chip runs its 11 high-performance cores at a base frequency above 5.7 GHz — extraordinarily fast by industry standards — with on-chip AI inference accelerators for in-transaction fraud detection, a dedicated data processing unit for I/O acceleration, and a large cache architecture. Full systems will scale to hundreds of cores and tens of terabytes of memory. “That’s really, really fast compared to what you otherwise get in the industry,” Jacobi said. “It’s just another example of how mainframe technology is not old technology. It’s very modern, leading-edge technology.”
The strategic logic behind the chip is about software, not hardware. IBM’s s390x architecture runs an enormous share of the world’s mission-critical transactions, but the broader universe of enterprise software — monitoring tools, security agents, cloud-native middleware, and above all the AI stack of PyTorch, ONNX Runtime, and container workloads — was built for x86 and, increasingly, for Arm. By Arm’s own estimates, close to half of the compute shipped to major hyperscalers in 2025 was Arm-based, driven by AWS Graviton, Google Axion, and Microsoft’s Arm silicon. Arm counts more than 22 million developers worldwide.
Porting each application to s390x has been a grinding, one-ISV-at-a-time effort, and Tina Tarquinio, chief product officer for IBM Z and LinuxONE, described the calculus bluntly. “No matter how great our ecosystem team is, we would never be able to work with all of them and port them all,” she told VentureBeat. “There’s a lot of ISVs out there, and so we wanted to make a fundamental, big step-function forward. We took a swing from a technology point of view.”
Notably, she said customers weren’t asking for a dual-architecture chip per se — they were asking for outcomes. “I wouldn’t say our clients were saying, ‘Can you please make me a dual-architecture environment?’ But they were saying, ‘Help me get these surround workloads, or different types of workloads, to run in a quicker-to-market fashion.'”
The compatibility promise is ambitious: Arm Linux binaries should run unmodified. “The new Arm capabilities are designed to be 100% binary compatible,” Jacobi said. “Once you have, for example, Red Hat Linux for Arm, and you have applications that run on Red Hat Linux for Arm, they will run on the system without modifications.” Arm defines the instruction set architecture and supplies validation tooling to guarantee that IBM’s implementation behaves identically to every other Arm chip — while IBM designs and builds the silicon entirely in-house. “Very good partnership. Very solid engineering partnership as well,” Jacobi said of the collaboration.
IBM is also previewing the next generation of its Spyre AI accelerator at Hot Chips, and the pairing is not coincidental. The current architecture already offers two tiers of AI: an on-processor accelerator, introduced with the Telum chip in 2022, that handles ultra-low-latency inference such as fraud scoring inside a payment transaction, and the Spyre accelerator card sitting in the I/O subsystem for heavier models.
The new Spyre raises the ceiling considerably. “We’re also bringing a much higher performance chip that is capable of running large language models for agentic workflows,” Jacobi said — both AI-ops workflows that administer the system itself and business workflows “for things like document understanding and insurance adjudication.” The new accelerator will ship with high-bandwidth memory to feed those models.
Here the dual-architecture bet and the AI bet converge. Enterprises want to run inference next to their data; the data lives on the mainframe; and the AI tooling is overwhelmingly Arm-native. Mohamed Awad, Arm’s executive vice president for cloud AI, framed the announcement in exactly those terms: “As AI scales, more of the computing landscape is converging on Arm. Bringing Arm compute and its software ecosystem to these platforms will extend that momentum into mission-critical enterprise infrastructure to give organizations greater choice in how they deploy AI.”
The timing tracks with where enterprise AI actually stands. McKinsey’s most recent State of AI survey found that while 88% of organizations now use AI in at least one business function, nearly two-thirds have not yet scaled it across the enterprise — and the companies capturing the most value are those redesigning core workflows rather than running detached pilots. For regulated industries whose systems of record sit on IBM Z, running AI where the transactions happen is arguably the most direct route to that kind of integration.
Buyers will need patience. The chip will debut in the successor to the z17, which shipped in the second quarter of 2025, and IBM holds to a roughly three-year product cadence — pointing to a launch around 2028. But Tarquinio insisted the program is well past the concept stage. “It’s more than being on the drawing board. We’re full steam ahead on the whole system,” she said, adding that IBM will release more details in the run-up to launch.
For IBM’s installed base, the reflexive question is whether embracing Arm signals a slow sunset for the traditional architecture. Both executives pushed back hard. “This is a big and. It is not an or,” Tarquinio said. “I have a roadmap that goes out 10 or 15 years of hardware systems. Many of our teams are working on this next system; many are also working on the one after that, and the one after that.”
Jacobi cast the move as continuity rather than rupture. “The traditional mainframe that we have today as a z17 system is not just a faster version of what we built 25 years ago,” he said. “We didn’t have pervasive encryption capabilities. We didn’t have on-processor AI capabilities. Adding the Arm capability is the next big iteration in this continuous evolution.”
The competitive subtext is the cloud. Asked why an enterprise would run Arm workloads on a mainframe instead of a hyperscaler, Tarquinio pointed to the platform’s availability numbers: “We’re talking eight nines of availability — that’s 0.3 seconds of downtime a year. If you’re running your ledger, if you’re running your fraud detection, any of these mission-critical apps, you want that.” The pitch, she said, is fit for purpose: match the infrastructure to the SLA, not the fashion.
There are real caveats. IBM’s own press release notes that statements of future direction “represent goals and objectives only.” The Arm support is Linux-only for now, and the hardest engineering — running a foreign instruction set at production performance, with mainframe-grade fault detection and recovery, under real customer workloads — remains to be proven over the next two years.
But the ambition is unmistakable. For sixty years, the mainframe has survived every wave of technology that was supposed to kill it — minicomputers, client-server, the cloud — by absorbing what it needed from each. Now IBM is attempting its boldest act of absorption yet: teaching the machine that runs the world’s money to speak the language of the AI era, fluently and natively, on the same silicon. “Bringing something that’ll really be first of its kind in production,” Tarquinio said, “showcases again what IBM is capable of from a technology point of view.” The mainframe, it turns out, isn’t being left behind by the future. It’s learning to run it.
Serval is making Catalyst, its AI agent for building enterprise automations, generally available Thursday and enabling it by default for customers — allowing teams of AI agents to decide what should be automated and then build the automation itself.
Catalyst sits above Serval’s AI-native service management platform as an admin-facing “super agent.” It can inspect ticket history, standard operating procedures or natural-language instructions, identify recurring work, and draft the workflows, skills, forms, access policies, journeys and dashboards needed to automate it.
Serval is also using Catalyst to create background agents that continuously inspect connected systems for emerging problems and propose fixes before an employee files a ticket.
That distinction matters because enterprise service management vendors are rapidly converging on AI-assisted workflow creation.
ServiceNow’s Build Agent can already translate natural-language instructions into full-stack applications, flows, scripts and other platform metadata, while its AI Agent Advisor can analyze instance records to identify automation opportunities. Atlassian’s Rovo can generate Jira automation flows from plain-English requirements, and Freshworks offers Freddy AI Agent Studio for creating service agents that act across Freshservice workflows.
So Serval’s claim to differentiation is narrower — and potentially more consequential — than simply “we use AI to build workflows.” Catalyst is designed as a single administrative layer that can move from discovering an opportunity, to assembling multiple kinds of governed automation, to creating proactive agents that keep looking for new work to automate.
“You just started with a single prompt, and now you’ve got enterprise-grade workflows ready to deploy that are going to solve all password resets for the entire company,” Serval co-founder and CEO Jake Stauch told VentureBeat in an exclusive interview.
Serval says Catalyst analyzes existing help desk data before an organization has decided what to automate. If it finds a repetitive category of requests, it can draft the automation required to resolve those requests and stage the result for administrator review. Users can also upload an SOP or spreadsheet and ask Catalyst to turn the documented process into an executable system.
Serval’s documentation says Catalyst can build workflows, author help desk skills, create onboarding and offboarding journeys, configure access-management policies, construct dashboards, investigate operational issues and debug failed workflow runs. Unlike Serval’s earlier workflow builder, Catalyst is intended to become the primary interface for configuring the platform; the company says its long-term goal is that anything an administrator can do through the UI should also be possible through Catalyst.
The actual workflows are code-backed. In a demonstration, Stauch showed Catalyst taking a request to build password-reset workflows, detecting connected systems including Okta, Google Workspace and Microsoft Entra, and generating the underlying TypeScript needed to perform those actions. Administrators could then add approvals or restrict who was allowed to run the workflow.
Serval is not building its own foundation model. Stauch said in the interview that the company uses models from “frontier labs,” runs evaluations to determine which models work best for particular jobs, and is deliberately model-agnostic. “You can swap different models in,” he said, adding that Serval also works with enterprises that build their own models.
Stauch provided more detail in a May 2026 interview with Sequoia Capital, saying Serval was using both OpenAI and Anthropic models. He said OpenAI’s GPT models had performed best for end-user interactions and tool calling, while Anthropic’s Sonnet and Opus models were producing the strongest results for the code-generation side of Serval’s automation system — the workload most directly relevant to Catalyst. Serval continuously runs evals rather than automatically moving every workload to the newest model release, Stauch said.
That architecture makes the underlying LLM less central to Serval’s differentiation. The company’s own documentation now lets organization administrators supply their own OpenAI or Anthropic API keys, including a compatible custom endpoint, while Stauch said the broader architecture can accommodate different models.
The materials do not, however, establish that every Catalyst user gets a self-service menu for arbitrarily choosing an individual model. Serval’s pitch is instead that its proprietary value sits in the harness around those models: enterprise context and memory, integrations, generated code, permissions, approvals and the controls governing what an agent can actually do.
That code-generation model is central to Serval’s pitch against ServiceNow. Stauch argues that legacy ITSM deployments often accumulate custom tables, business rules, workflows and platform-specific expertise that make seemingly simple automation changes expensive to implement. Serval, by contrast, wants administrators and business teams to describe the outcome they need and let the model generate the implementation.
But ServiceNow is no longer standing still on that front. Its current Build Agent similarly creates applications and code from natural-language prompts, supports flow design and testing, and operates inside ServiceNow’s governance framework. ServiceNow’s AI Agent Studio lets customers create agents and agentic workflows, while AI Agent Advisor is explicitly designed to analyze operational records for automation candidates.
The competitive question is therefore shifting from “who has generative AI?” to how many separate tools, configuration concepts and specialists are required to get from an observed operational problem to a production automation.
Serval is effectively arguing that Catalyst compresses those steps into one conversational surface and a smaller platform model. ServiceNow, by comparison, now has a powerful but broader set of AI and development surfaces spanning Build Agent, AI Agent Studio, AI Agent Advisor, Workflow Studio and AI Control Tower. That breadth is an advantage for customers already deeply invested in ServiceNow, but it also illustrates the complexity Serval is attacking. ServiceNow itself notes that Build Agent is aimed at admins and developers who understand and can support what it generates.
Atlassian is moving in the same direction from a different starting point. Rovo can generate “if this happens, then that happens” automation flows from natural-language descriptions, while Jira Service Management increasingly supports agents that triage, investigate and execute service work.
Freshworks’ Freddy AI Agent Studio likewise emphasizes agents that resolve requests end-to-end, with prebuilt IT and HR agents and more than 30 workflow templates.
Catalyst’s differentiator, then, is not that rivals cannot generate an automation from a sentence. It is Serval’s attempt to make the entire automation lifecycle itself agentic.
That approach becomes clearest with Serval’s background agents.
Rather than waiting for a help desk request, a background agent can run on a schedule across connected systems, correlate signals and draft a remediation. In one customer example provided by Serval, an agent correlated network incidents across two offices using switch telemetry, DHCP data and historical tickets, ruled out hardware and wireless interference, traced the issue to configuration drift, and generated a remediation workflow for an administrator to approve.
“Most AI agents today wait for an employee to ask a question or submit a ticket,” Stauch said. “We believe the future is AI that acts before an employee ever submits a request.”
That framing also highlights a philosophical difference in Serval’s pitch. The startup does not want service management to revolve around creating, routing and tracking better tickets. It wants the system to eliminate as many requests as possible by turning repeated support work into executable automation.
“A lot of the code written in enterprises has nothing to do with software engineering,” Stauch explained. “It’s actually internal automations and other scripts for the company, and so we use that technology to build a better service management platform.”
Serval’s pitch to enterprises is that it can largely automate those scripts. And the governance model is critical because Catalyst can generate code and potentially initiate changes across production systems. Serval says Catalyst inherits the permissions of the user operating it and remains scoped to that user’s team workspace.
Everything it builds starts as a draft, and organizations can restrict publishing privileges or require formal review and approval before an automation becomes active.
Those controls also extend to the enterprise data Catalyst examines. Stauch said Serval is intended to operate as the customer’s system of record and told VentureBeat that “they own all the data.”
Serval’s current Master Services Agreement is more precise: customers retain rights, title and interest in both their “Customer Materials” — a category that includes records, documents, workflows, prompts, inputs and configurations — and the output Serval generates from them. Serval receives the rights necessary to process that information to provide, maintain, support and secure the service.
Serval also says it does not retain or use customer materials, inputs or outputs to train, fine-tune or improve its own or third-party AI models.
Its Data Processing Addendum identifies Serval as the processor of customer personal data and allows processing for operating the service, responding to support requests, diagnosing issues and protecting the platform, while authorized subprocessors can also be involved. Serval’s acceptable-use terms say it maintains a current list of AI subprocessors and model providers for customers.
Where that data resides can vary by deployment. Stauch said customers can use Serval as a cloud SaaS service, run it on-premises or place it in their own VPC. Serval’s self-hosting documentation now describes two fuller options: a Serval-managed single-tenant deployment inside an AWS account owned by the customer, or a self-managed deployment on the customer’s Kubernetes cluster in any cloud or on-premises environment.
In the AWS option, Serval says it operates the installation without persistent IAM access to the customer’s AWS account.
There are therefore two distinct access boundaries for enterprise buyers to consider.
At the Catalyst level, the agent can only reach data, integrations and automations available to the user and team workspace under which it is operating.
At the platform level, Serval and authorized subprocessors necessarily process customer information to deliver and support the service, subject to the company’s contractual confidentiality and data-processing terms.
That makes Stauch’s informal statement that Serval “doesn’t touch” customer data better understood as an ownership and deployment claim, rather than a literal assertion that the service never processes it.
Customer deployments provide some evidence that the faster-build thesis can translate into operational changes, although the metrics come from Serval’s own case studies.
Corporate expense and financial technology firm Ramp says in a Serval case study that Catalyst has made workflow building 50% faster and helped extend Serval across roughly 10 teams, including IT, finance, facilities, people and talent, legal and business operations. In one hardware replacement program, Serval says Ramp automated 600 laptop replacements and saved 150 hours, leaving approval as the principal human step.
The more telling Catalyst example may be what happened afterward. Ramp had already automated laptop replacement when Catalyst suggested splitting its shipping logic into separate office and home workflows to reduce errors. The company also says employees outside IT now use Catalyst for analytics, bulk ticket operations, workflow troubleshooting and HR process automation.
Other Serval deployments show the broader operating environment Catalyst is meant to configure. Mercor says it has onboarded more than 4,000 external experts through Serval automations and expanded the platform across seven teams. Together AI says Serval automates 95% of its just-in-time infrastructure access requests, with approval and auditing controls around sensitive access. Perplexity says Serval automatically handles more than half of its incoming IT requests and all employee onboarding.
Those deployments extend beyond Catalyst itself, but they demonstrate the type of cross-system automation substrate Catalyst is now being asked to build and maintain.
Serval says more than 90% of customers adopted Catalyst as their starting point for automation during beta. Catalyst is generally available Aug. 20 and will be enabled by default for all Serval organizations.
Pricing is customized depending on the size of the deployment and is not publicly listed on Serval’s website or documentation.
Serval describes a single platform fee and typically runs a pilot to determine expected deployment and usage.
Stauch said the software license can be similar to ServiceNow’s, but argues total cost of ownership can be substantially lower because customers require fewer implementation and maintenance services.
“The total cost of ownership is going to be dramatically less — usually half as much, sometimes 10 to 20% of the total cost of ownership of ServiceNow,” Stauch said. “But the actual software license fee is not necessarily going to be all that different.”
Serval was founded in 2024 by Stauch and CTO Alex McLeod, former Verkada product and engineering leaders, after they repeatedly heard IT customers complain about overburdened help desks and the limitations of established IT service-management software.
Serval has positioned itself as an AI-native alternative to platforms such as ServiceNow and Jira Service Management, combining help-desk ticketing, access management, asset management and workflow automation within a single system.
Serval and Sequoia Capital describe the company’s goal as moving IT software beyond merely recording and routing requests toward resolving them automatically.
The company can operate as an organization’s primary IT service-management system or add automation to an existing one. Its publicly identified customers include Perplexity, Mercor, Clay, Verkada and Together AI.
Serval says customers can automatically resolve more than half of their incoming IT requests; its Together AI case study reports automation of 95% of that customer’s just-in-time access requests.
Investor interest accelerated rapidly in late 2025. Serval announced a $47 million Series A led by Redpoint Ventures in October, bringing its funding at that point to $52 million.
In December, it raised another $75 million in a Sequoia-led Series B at a $1 billion valuation, lifting total capital raised to approximately $127 million; Redpoint, Meritech Capital and General Catalyst also participated.
Serval told Reuters that revenue had grown 500% since August 2025 and that it was expanding beyond IT into operational work performed by human resources, finance and legal departments.
For enterprise buyers, Catalyst’s biggest test will be whether its compression of the automation lifecycle survives contact with large, messy, highly customized environments.
ServiceNow can now generate applications and discover automation opportunities with AI. Atlassian and Freshworks are adding increasingly capable agentic automation to their own service platforms. Serval therefore cannot rely on natural-language creation alone as its moat.
Its stronger wager is that an AI-native platform can make the administrative layer itself agentic: continuously finding repetitive work, building the necessary resources across the service stack, exposing generated code for review, and proposing the next automation before an administrator has opened a workflow designer.
If Catalyst works at that scope, the competitive unit is no longer the ticket — or even the workflow. It is the system that keeps turning an enterprise’s operational history into new automation.
Cursor began rolling out Origin, its own code hosting platform, to paid users on Monday morning. Roughly three and a half hours later, GitHub’s status page lit up with what became a six-hour-and-forty-two-minute global degradation — error rates near 20% across pull requests, issues and the API, and near 50% on archive and raw file downloads, according to GitHub’s incident log. Enterprise single sign-on went down with it: SAML, OIDC, SCIM provisioning and Team Sync all failed. So did Copilot.
The developer internet did what the developer internet does.
“You can now host your repos in Cursor Origin and deploy to Vercel via Cursor Origin which is itself hosted on Vercel,” Vercel chief executive Guillermo Rauch posted on X. “And unlike GitHub, it’s online 😁” Asked why he was smiling, Rauch replied: “trying to make light of the situation. We ourselves are stuck because of github rn!”
Matt Palmer, who works at Cursor, quote-tweeted his own company’s launch with the day’s best line: “We were going to ship this earlier, but GitHub was down.” A GitHub outage, in other words, delayed the launch of a GitHub competitor.
Product launches get locked weeks in advance, and no evidence suggests Cursor timed this one. But the coincidence did the company an enormous favor, because it dramatized the argument Origin exists to make. For eighteen years, choosing where to host your team’s source code has been the least interesting decision an engineering organization makes. Cursor is betting that AI agents have made it interesting again — and for technical decision makers, that is the real news here. Not a new product, but a new procurement question with a governance problem attached.
Origin lives in a new Codebase tab inside Cursor. Teams name a codebase, which becomes part of its URL, then push to it over the command line. From there they get the machinery you would expect from a forge — the service layer that wraps Git and handles storage, permissions, checks and merges. Every repository comes with pull requests: timelines, commits, checks and files changed. Reviewers read the diff, leave comments and merge, without ever opening a browser tab.
What Cursor built around that machinery is the part worth studying. Agents now operate in the same surface as the code and the pull requests they are modifying. “Your code, PRs, and agents are now in the same place,” the changelog reads. A developer can ask questions about the file on screen, hand an agent a review comment and have it revise the pull request in place, or tell it to push a branch — all inside the editor where the code was written.
Three integrations shipped on day one, and the choice of partners is telling. Vercel spins up a preview deployment for every pull request and ships to production on merge, available in public beta for Pro and Enterprise customers, its developer account said. Depot and Buildkite run continuous integration, and critically, both execute existing GitHub Actions workflows unchanged. Buildkite adds native pipelines on top.
That compatibility layer is the whole strategy in miniature. Cursor is not asking teams to rewrite their build system, retrain their engineers or rip out their deployment pipeline. It is asking them to try a second window onto code they already have — which is a far easier request to approve.
More partners are coming, the company said, and the ones it landed first are the ones that matter to a platform team evaluating whether Origin can carry real work. A forge without deployments and CI is a code viewer. A forge that runs your existing Actions workflows and ships previews to the CDN you already pay for is a candidate.
Here is the decision enterprise buyers should study most closely, because it determines whether Origin survives a security review at all.
Cursor does not ask you to leave GitHub. Connect a GitHub organization, pick repositories, and they appear alongside Origin-native ones. “Pushes keep going to GitHub, which stays the source of truth for anything started there,” the changelog says. Access permissions mirror GitHub’s existing read and write settings rather than establishing a parallel system. Pull request conversations sync in both directions — comment in Cursor and it posts to GitHub; reply or react on GitHub and it surfaces in Cursor “within seconds.”
This is a classic wedge, and a well-executed one. Rip-and-replace migration of source control ranks among the highest-risk projects an engineering organization can undertake. It touches continuous integration, compliance evidence, audit trails, branch protection rules, every integration in the toolchain and the muscle memory of every engineer on staff. Almost no chief technology officer approves that for a product in early beta.
A read-mostly mirror that leaves GitHub authoritative approves itself. It costs nothing to try, breaks nothing if abandoned, and quietly relocates the place developers spend their working hours. If Cursor’s review experience proves better — and Cursor spent real money to make sure it would — the source of truth eventually follows the attention.
That money went to Graphite, the code review startup Cursor bought in December 2025 for what Axios reported was well above its $290 million Series B valuation. Graphite built stacked pull requests, the workflow that lets developers keep shipping dependent changes without waiting on approvals. Announcing the deal, Cursor wrote that “the boundary between where you write code and where you collaborate on it feels increasingly arbitrary,” and promised “some more radical ideas we can’t share just yet.” Origin is the radical idea. Graphite co-founder Tomas Reimers unveiled it on stage at Cursor’s inaugural Compile conference in June and leads its development.
The case for an agent-native forge rests on a claim that is easy to state and, unusually for this market, well supported by evidence: writing code stopped being the constraint. Reviewing and integrating it became one.
Google’s 2025 DORA report, drawn from nearly 5,000 technology professionals, found that 90% of developers now use AI at work, spending a median of two hours a day with it, and more than 80% say it made them more productive. But AI adoption showed a positive relationship with software delivery throughput and a negative one with delivery stability. More output, more breakage. The report’s authors describe AI as “an amplifier” that “magnifies the strengths of high-performing organizations and the dysfunctions of struggling ones.”
Trust has not kept pace with volume. Stack Overflow’s 2025 developer survey of 49,009 respondents across 177 countries found 84% using or planning to use AI tools, while trust in their accuracy fell to 33% from 43% a year earlier and distrust climbed to 46% from 31%. Two-thirds named “AI solutions that are almost right, but not quite” as their leading frustration. GitLab’s ninth annual DevSecOps survey, of 3,266 practitioners polled by Harris, put numbers on the operational drag: 73% had hit problems with vibe-coded output, 70% said AI made compliance management harder, and only 37% would let AI handle daily tasks without human review.
The volume climbs regardless. GitHub’s Octoverse 2025 counted 180 million developers, 630 million repositories and 43.2 million pull requests merged per month, up 23% year over year. And RuntimeWire reported the internal figure that best explains Origin’s existence: 35% of pull requests merged inside Cursor were opened by agents running autonomously in cloud virtual machines.
A forge built for humans assumes a pull request represents human intent, opened by someone you can ask what they meant. Once a third of merged changes come from software, the queue stops being a conversation and becomes a scheduling problem. That is a real architectural argument, and it is the strongest thing Cursor has going for it.
The supply-side case for an alternative is simpler: GitHub has been unreliable, and its own executives have said so.
An analysis by LeadDev counted 257 incidents between May 2025 and April 2026, 48 of them major — roughly one significant disruption per week. February was the worst month on record with 37. GitHub Actions alone accounted for 57 outages in twelve months. Chief technology officer Vlad Fedorov has said the platform “wasn’t built for the scale it’s now being asked to handle” and must design for 30 times today’s load. In an April engineering post covered by InfoQ, the company acknowledged it “failed to meet its own reliability standards,” citing rapid growth, tight architectural coupling and inadequate load shedding. Monday’s outage was the seventh incident on GitHub’s status page in fifteen days.
The fatigue is audible. “GitHub really doesn’t feel built for the agent era,” one developer wrote on X as Origin went live. “It goes down way too often, but until now there haven’t been many real alternatives.”
The defections started before Origin existed. The Zig programming language moved to Codeberg in November 2025, citing Actions failures among its reasons. In April, Mitchell Hashimoto announced that Ghostty — a terminal emulator with more than 52,000 stars — would leave too, pointing to near-daily outages that blocked reviews and CI for hours. And The Information reported in March that OpenAI, a company Microsoft holds a large stake in, began building its own GitHub alternative partly because outages left its engineers unable to commit for hours at a time, as Tom’s Hardware relayed.
Microsoft’s structure has not helped. Thomas Dohmke resigned as GitHub chief executive in August 2025 and was never replaced; the unit’s leadership was absorbed into Microsoft’s CoreAI organization under executive vice president Jay Parikh. In a May report, The Information wrote that Parikh had warned deputies that coding tools from Cursor and Anthropic could eventually make GitHub obsolete. GitHub’s own answer to the agent era, Agent HQ, lets customers orchestrate third-party agents from Anthropic, OpenAI, Google, Cognition and xAI inside GitHub — a coherent strategy that concedes the agent layer and keeps the substrate underneath. Origin attacks precisely that substrate.
Cursor’s rise has been extraordinary even by the standards of this cycle. Founded in 2022 by four MIT students, Anysphere raised $8 million from the OpenAI Startup Fund in October 2023, per TechCrunch, then $100 million at $2.5 billion, $900 million at $9.9 billion, and $2.3 billion at $29.3 billion last November. In May, Bloomberg reported annualized revenue of $3 billion and more than 3,000 customers paying at least $100,000 a year.
Then, three days before Origin shipped, Bloomberg reported that SpaceX completed its $60 billion all-stock acquisition of Cursor — an agreement TechCrunch covered in June, days after SpaceX’s record IPO and six months after it absorbed xAI. Cursor now operates inside a division called SpaceXAI. The vendor asking to hold your proprietary source code became, last Friday, a unit of a rocket company with its own frontier-model division and a founder not known for institutional caution.
Jason Andersen of Moor Insights & Strategy raised the model-routing question to Tech Times in June, before the deal closed: “xAI’s models and treatment of guardrails are very different than what Cursor has stood for.” That piece framed the question a chief information security officer now has to answer. When one company controls the editor where agents write code, the host where that code lives and the model those agents run on, what governs what it does with the code?
Cursor has not published an answer. RuntimeWire noted before launch that Origin’s pricing, security architecture, data-handling terms and migration tooling were all unpublished, and Monday’s changelog adds none of them. It says only that Origin reaches “all paid plan users starting today, except enterprise orgs whose admins opt out.” Opt-out, not opt-in — a sentence administrators should read twice.
There is also a track record to weigh. In July, researchers at Mindgard disclosed that Cursor would execute a malicious git.exe planted in a Windows project’s root the moment a user opened it, with no prompt — a repository-poisoning flaw they first reported in December 2025. The Hacker News reported that Cursor declined to patch it, calling the issue out of scope under a shared-responsibility model while conceding it had not “closed the loop with the researcher in a timely manner.” No CVE was issued. The same flaw class turned up unpatched in GitHub Copilot CLI, Google’s Gemini CLI and OpenAI’s Codex — but a vulnerability the vendor declined to fix makes an awkward footnote for a product whose pitch is basically “let us hold your repositories.”
Origin is a beta, not a migration, and treated as one it is worth evaluating. The sync mode gives platform teams a low-risk way to measure whether an agent-native review surface shortens cycle time, without touching a single branch protection rule. But three things deserve resolution before anything authoritative moves.
The first is the default. Origin switches on for paid users unless an enterprise administrator opts out, which means an organization that has not made an affirmative decision about whether proprietary code may be mirrored to a new host has effectively had that decision made for it. Confirming your posture is a Monday-morning task, not a next-quarter one.
The second is the paperwork. Retention, residency, training use, subprocessors and what changes now that Cursor reports into SpaceX are all unpublished, and a product page is not a contract. Until those terms exist in writing, the defensible position is to treat Origin as a convenience layer over GitHub rather than a system of record — which is, conveniently, exactly what its architecture already is.
The third is the exit. Origin’s Actions compatibility and its GitHub-as-source-of-truth design are the properties that make it safe to adopt. They are also the ones most likely to erode as Cursor’s incentives shift toward owning the substrate rather than borrowing it. Ask what egress looks like now, while the mirror is still a mirror.
None of which makes Cursor’s argument wrong. GitHub earned its incumbency by being boring, dependable infrastructure, and it has spent eighteen months being neither while a third of the code arriving at its front door stopped being written by people. Origin is a serious answer to a real problem, built by a team that bought the right company to build it.
But GitHub’s failure and Cursor’s are different in kind, and enterprises should not confuse them. Monday’s outage resolved at 20:22 UTC. Availability is an engineering problem, and engineering problems close. The question of who holds your source code, what they may do with it and who they ultimately answer to carries no such timestamp — and on that one, the company that spent Monday selling trust has yet to publish its terms.
Current space exploration aims to establish permanent structures on the Moon, Mars, and eventually other planetary bodies. Successful lunar missions depend on understanding lunar regolith, the granular material covering the Moon’s surface, whose behav…
Mistral AI wants to turn European AI sovereignty from a talking point into a product — one with a service-level agreement attached.
The French artificial intelligence company announced Tuesday a three-part expansion of its infrastructure business: regional inference endpoints that let customers choose whether their AI workloads run in Europe or the United States, a new “Priority Tier” backed by an uptime guarantee for mission-critical deployments, and a coalition of European enterprises making multi-year compute commitments that Mistral says will underwrite 200 megawatts of infrastructure across Europe by the end of 2027 — and a full gigawatt by the end of 2030.
In a move that may raise eyebrows among sovereignty purists, the company also said it will begin hosting third-party open models on its platform, starting with GLM-5.2 from Z.ai, the Chinese AI lab formerly known as Zhipu.
Taken together, the announcements mark a decisive shift in how Mistral positions itself. The company that built its reputation training open-weight language models is now selling something closer to critical infrastructure: assured capacity, regional control, and contractual reliability for enterprises and governments that want frontier AI without surrendering control over where it runs.
“When we spoke in June, the story was around how Mistral was building a full-stack AI offering,” Timothée Lacroix, Mistral’s co-founder and chief technology officer, told VentureBeat in an exclusive interview ahead of the announcement. “Today, the announcement is about strengthening one part of this infrastructure, which is the inference part.”
That one part, it turns out, comes with a price tag measured in the tens of billions of dollars.
The headline numbers deserve scrutiny, because they imply staggering capital requirements. Mistral currently operates less than 200 megawatts of capacity, according to the company. Details shared with VentureBeat show the near-term buildout resting on three sites: a 44-megawatt facility near Paris that became operational in the second quarter of this year, a 23-megawatt facility in Sweden built in partnership with EcoDataCenter using renewable energy and advanced cooling, and a 10-megawatt site in Les Ulis, France, that came online in the third quarter.
Getting from there to one gigawatt by 2030 is a different order of magnitude. Independent estimates suggest just how different: research firm Epoch AI calculates that a typical one-gigawatt AI data center requires roughly $38 billion in upfront capital expenditure, with servers and GPUs — not buildings or land — consuming the majority of the cost. Goldman Sachs Research pegs next-generation AI facilities at $15 million to $20 million per megawatt before accounting for the chips inside them.
Lacroix did not dispute the scale of the challenge. The investment required for a gigawatt of capacity “is a large investment that requires also a lot of scaling and revenue behind it,” he said.
The urgency, in his telling, comes from a supply crunch that is about to get worse. “More and more, and especially around 2027 and 2028, we see that the demand for AI compute is exceeding what the market has to offer, especially in Europe,” Lacroix said. McKinsey has estimated that meeting global AI demand could require $5.2 trillion in data-center capital expenditure by 2030 — and Europe, by most analyses, is starting from behind.
A company valued at a fraction of its American rivals cannot close that gap with venture capital alone. Which explains the most consequential — and most unusual — piece of Tuesday’s announcement.
Mistral is assembling what it calls an anchor group of enterprises whose long-term commitments will collectively finance infrastructure none of them could justify alone. Those commitments convert into “European Compute Units,” or ECUs — a claim on Mistral-built capacity over multiple years that participants can spend on inference, training, model adaptation, or other AI workloads as their needs evolve.
If that structure sounds more like a power-purchase agreement than a cloud contract, that appears to be the point. Data-center financing increasingly resembles large infrastructure projects — gigawatts, substations, energy agreements — rather than traditional technology spending, and lenders want demand locked in before capital gets deployed. Mistral raised €830 million ($962 million) in debt earlier this year to fund its data center near Paris, TechCrunch reported in March, and pre-committed enterprise demand is exactly what makes that kind of financing repeatable at ten times the scale.
Lacroix was unusually direct about the mechanics. “The entire point of compute units is to have commitment,” he said. “The goal is to have customers commit for around five years, or at least a long time.” Asked what happens if a customer wants out early, he didn’t soften the answer: “There is no getting out.”
What makes a five-year, no-exit commitment palatable, he argued, is flexibility in how the capacity gets consumed. “Typically this can be spent on raw inference that you then feed through any other AI stack. It can be spent on raw compute as managed Kubernetes, and it can be spent at the very top with our full AI offering,” he said. “My hope is that they will use it with our full-stack services and will love it.”
The anchor group already includes some of Europe’s industrial heavyweights. Amadeus CEO Luis Maroto said in a statement that “capacity, deployment control, and operating continuity become increasingly important for all enterprises.” ASML chief Christophe Fouquet — whose company led Mistral’s $13.4 billion (€11.7 billion) Series C last year — called building European AI capacity one of the few industrial endeavors that “will matter more to Europe’s next generation,” while Capgemini’s Aiman Ezzat framed it as “a question of who shapes the future of European industry.” CMA CGM chairman Rodolphe Saadé said the shipping group’s Mistral deployment is “already under way among thousands of employees.”
Commitments of that duration only make sense, of course, if the sovereignty being purchased is real. On that question, Mistral’s announcement contains an asterisk worth reading closely.
The centerpiece product is Mistral Regional Endpoints, now generally available, which let customers pin inference and its associated processing to Europe or the U.S. Alongside it, the new Priority Tier — in public preview — offers committed service levels, custom rate limits, and an uptime SLA for mission-critical workloads.
Mistral claims it is the only European AI lab offering both a choice of processing region and an SLA-backed service tier, and Lacroix said a third option is coming: an endpoint “that stays on Mistral-controlled infrastructure, so on Mistral compute” — for customers who want their inference not just in Europe, but off hyperscaler hardware entirely.
Then comes the fine print. Mistral’s own materials note that in-region inference remains subject to “limited, safeguarded transfers” to sub-processors that may sit outside the chosen region. Pressed on what actually leaves Europe, Lacroix pointed to the connective tissue of modern AI applications: tool calls.
“There are some tool services, like some tool calls, that might be hosted in places where we don’t fully control this,” he said, citing web search as an example. “A few of our web-search providers might not all be in Europe, and in that case, we need to potentially gate that capability.”
His answer to the compliance question — would this satisfy a European bank or a defense ministry? — was that gating is the feature, not the bug. Capabilities that cannot be sourced in-region can be switched off entirely, restricted to certain users or workspaces, or, given sufficient demand, rebuilt with European providers. “Any capabilities that we don’t find a provider for in Europe — if it needs to be done in Europe, we’ll find some way to implement it or find ways to address it,” Lacroix said.
For enterprise buyers, that is a more honest framing than most sovereignty marketing offers: full regional control is available, but the moment an AI agent reaches out to the open web, sovereignty becomes a configuration decision rather than a default. The same pragmatism runs through the announcement’s most surprising line item.
A French national champion — one that has partnered with the French army and positioned itself as Europe’s answer to American AI dependence — hosting a Chinese lab’s model invites an obvious question. Lacroix’s answer was disarmingly matter-of-fact.
“It’s a great model. Everyone loves it. It’s open weight, so there was no good reason for us not to do it, really,” he said, noting that Mistral’s own stack is already built on open-source software like Kubernetes.
On security vetting, he argued that open weights fundamentally change the risk calculus. “The risks in taking a new model, at the layer of the weights, are — at least in my opinion — rather limited,” Lacroix said. “We checked basically all of the safety and compliance evals that we have. We’ll control that model, its outputs, and what it does the same way we do any of our models. We have the same inputs and outputs and monitoring capabilities over all of it.”
The strategic logic is worth unpacking. By hosting third-party open models under European regional controls and the same SLAs as its own, Mistral is repositioning itself from model vendor to sovereign distribution layer — the trusted intermediary through which any open model, regardless of origin, can be consumed by a regulated European enterprise that could never call a Chinese API directly. It is the “model garden” playbook the hyperscalers run with Bedrock and Vertex, executed on European soil with European guarantees.
Customers appear to be reading it that way. “Mistral allows us to run open models under strict regional controls and service commitments, making it easy for us to maintain data residency and compliance requirements,” Matan Griberg, CEO of AI software-engineering company Factory, said in a statement.
Lacroix stressed the move is not a retreat from frontier training: the model Mistral had in training as of June “is still training, and we’re still very excited about it,” he said. But openness to rivals’ models signals where the company now believes its moat lies — not in any single model, but in the infrastructure underneath all of them. Which makes its relationship with the world’s most powerful infrastructure company all the more interesting.
Hovering over every sovereignty claim is Mistral’s deepening relationship with Microsoft. In July, the two companies announced a multibillion-dollar expansion of their partnership under which Microsoft will rent capacity from Mistral’s European data centers to serve its own cloud and AI demand, while adding Mistral Medium 3.5 and OCR 4 to Microsoft Foundry, bringing Medium 3.5 to Copilot Studio, and enabling Mistral models on Azure Local for disconnected, customer-controlled environments. Mistral CEO Arthur Mensch told The Wall Street Journal at the time that two-thirds of Mistral’s customers already work with Microsoft.
How does a company selling independence from U.S. hyperscalers square taking one on as its largest tenant? Lacroix described Microsoft not as a patron but as an anchor customer that de-risks the buildout.
“It allows us to scale different parts of the business differently by building infrastructure with Microsoft as a customer,” he said. “We can scale that team, we can scale our infrastructure, and make sure that we can then, on the side of it, also build for ourselves and for our customers.” He compared the arrangement to the neocloud playbook — companies that built businesses supplying capacity to the hyperscalers themselves. “As that part of our business resembles that of neoclouds, we’re following the same thing.”
It is a genuinely clever inversion: rather than renting American infrastructure, Mistral is renting infrastructure to one of America’s largest companies, using Microsoft’s demand to finance capacity that also serves European sovereignty customers. But the independence has limits no contract can engineer away — the GPUs filling Mistral’s European data centers come overwhelmingly from Nvidia and other American chipmakers, as SiliconANGLE noted in its coverage of the July deal.
Asked directly why a customer should choose Mistral over an EU region on AWS or Azure, Lacroix gave two answers. “The simplest possible answer is capacity. There is more demand than supply right now, and so it adds another option,” he said. The second cuts closer to the pitch: “We are a European provider, and on the region that would be Mistral compute, we are fully independent. That’s a truly differentiated offering than all of the hyperscalers or pure inference companies can provide.”
There has always been a tension at the heart of Mistral’s business: its best-known models are free to download, and open models have historically been difficult to monetize through APIs. Asked how free weights fund a gigawatt buildout, Lacroix offered the clearest articulation yet of the company’s thesis — that the economics of self-hosting are collapsing under the weight of the models themselves.
“When the models were smaller, and we were before the explosion of agentic AI, it was doable for enterprises to host their own — up to, let’s say, 100-billion-parameter dense models — on their premises,” he said. “More and more, with models going into the trillion or more parameters, with the current hardware, and with the increasing amount of tokens that need to be processed, it becomes harder.”
His conclusion was blunt: “I don’t see how, with the current trend of model size and growth of agentic tokens, we keep the full inference on-prem. To me, that is why we think we’re going to monetize our cloud inference.” Inference, he noted, is particularly well suited to the cloud because it “does not need to hold any data” and can be encrypted in transit.
In other words: open weights get Mistral into the enterprise, and the physics of trillion-parameter agentic workloads brings the inference — and the revenue — back to Mistral’s data centers. The thesis will get an expensive test. Mistral has raised roughly $4 billion to date, according to PitchBook data — a fraction of the war chests assembled by OpenAI and Anthropic — and Bloomberg reported in June that the company is in talks to raise about €3 billion at a roughly €20 billion valuation, nearly double its Series C mark. The revenue behind the buildout will have to come from exactly the enterprises Tuesday’s announcement is courting.
And Europe, in Mistral’s telling, is only the first market for what it is selling. Asked whether the framework could be replicated in the Middle East, Asia, or anywhere else anxious about AI dependence, Lacroix didn’t hedge: “It’s completely right. We’re starting this in Europe because it’s also an easier part of the world for us to scale into, especially in the infrastructure. But we definitely want to extend this, depending on customer demand.” Every layer of the stack, he said, “can be controlled, changed, replaced depending on where we operate and what the requirements are — that’s pretty much where we excel.”
That is the wager underneath the SLAs, the compute units, and the Chinese model flying a European flag: in a world where the U.S. and China dominate frontier AI, the durable business is selling everyone else control. To fund it, Mistral is asking Europe’s largest enterprises to sign five-year contracts with no exit — while making a bigger, longer commitment of its own. A gigawatt, after all, is a promise measured in decades. For Mistral, too, there is no getting out.
Presented by Tata Communications
Continuous inference, agent-to-agent communication, and real-time data pipelines are generating unpredictable, always-on traffic that legacy architectures were never built to support. As AI moves from pilot project to operational backbone, the network is emerging as a critical control layer that determines performance, reliability, and cost.
The shift is forcing organizations to question assumptions that have held for decades. Legacy systems were static and rigid, and lacked the ability to manage network demand efficiently or dynamically, while AI-ready networks need to adapt in real time. A study by Cisco notes that 80% of executives believe their company’s competitive survival will depend on agentic AI, and consumer usage of AI is already prevalent and accelerating. This is driving a fundamental shift in how traffic is generated, distributed, and experienced, with implications for service providers and enterprises that manage large-scale networks.
This infrastructure gap is a global concern. A recent Bloomberg study, “The Future-Ready Enterprise,” commissioned by Tata Communications, found that while 3 in 4 leaders consider AI a board-level priority, nearly two-thirds (65%) of enterprises continue to operate on transitional or legacy infrastructure. This disconnect between ambition and reality is a primary obstacle to realizing value from AI investments.
The performance bar has also moved by an order of magnitude. Traditional business applications could tolerate 100 to 500 milliseconds of latency, while mission-critical AI workloads now require latency below 10 milliseconds.
“This isn’t just an incremental improvement,” says Kapil, Vice President, Global Network Services at Tata Communications. “It’s a completely different performance paradigm that breaks traditional network design assumptions, where such extreme low latency was never a primary consideration.”
That gap between what legacy infrastructure can deliver and what AI demands turns network performance into a direct driver of AI reliability and cost. Treating the network as a best-effort transport layer introduces risk that many organizations only discover once a deployment underperforms in production. A model built for real-time fraud detection or supply chain optimization becomes worthless the moment network congestion delays the data it depends on, and Kapil notes that every millisecond of that delay can carry a direct financial or operational cost.
“Relying on a ‘best-effort’ network turns multi-million-dollar AI stack investments into a high-stakes gamble, where performance is left to chance,” Kapil says.
He adds that businesses often underestimate the complexity of using the public internet as a global enterprise network. Performance may look acceptable within a single country, but once data starts crossing borders or connecting to international cloud platforms, the lack of end-to-end control becomes an operational barrier.
Complexity compounds as AI components spread across cloud, edge, and enterprise environments. Organizations often focus on compute power and data infrastructure while overlooking the network fabric that connects them. That blind spot often surfaces as a performance bottleneck created by high-frequency east-west traffic moving between GPUs.
Distribution also widens the surface enterprises have to defend. Applications, users, and partner ecosystems are now spread across cloud, SaaS, edge, and device environments, and Kapil notes that AI-driven malicious bots account for roughly 37 percent of online traffic, making it increasingly difficult to distinguish legitimate users from automated threats. Many enterprises have responded by layering on siloed tools, which has produced fragmentation, inconsistent security, and a lack of unified visibility rather than a coherent defense.
“SASE helps mitigate these risks by converging networking and security into a unified, cloud-delivered architecture,” Kapil says. “This convergence is enabling consistent policy enforcement across cloud, on-premises, and edge environments, while supplying the scalability and proximity needed to secure real-time AI-driven interactions.”
Closing that gap requires organizations to gain far greater visibility into how AI traffic moves across distributed environments and the ability to direct workloads accordingly. Kapil says that demands a different approach to network management.
“Leaders must realize that the network is no longer passive ‘plumbing.’ It must be managed as an active, intelligent platform foundational to the entire AI stack,” he says. “That platform requires real-time observability into how and where AI traffic flows, paired with the control to orchestrate workloads across the most efficient and secure path available.”
It’s the difference between merely connecting systems and unlocking new capability, for instance a seamless shopping experience during a peak sales period or a global sports broadcast streamed without buffering.
This intelligence also changes how infrastructure teams spend their day. The network itself is now software-defined and API-driven rather than fixed by hardware configuration, which Kapil says shifts infrastructure teams away from reacting to outages and toward designing the systems that prevent them.
“Instead of manually re-routing traffic during an outage, the team must define the rules, policies, and business outcomes for an intelligent fabric,” Kapil says. “The network itself then executes those policies automatically and autonomously.”
Tata Communications is putting this principle into practice with its recently launched IZO Data Centre Dynamic Connectivity. The software-defined platform creates a “self-healing, intelligent network” using deterministic multi-path routing to reroute traffic automatically in seconds during a disruption.
The company says the platform transforms resilience from a reactive process into an autonomous capability, providing the predictable, low-latency performance mission-critical AI applications require while reducing operational costs by up to 30%.
Delivering on that intelligence in practice means giving mission-critical workloads dedicated capacity rather than having them compete for it. Reaching that level of consistency also requires enterprises to define performance far more precisely than they have in the past. It’s the shift from vague goals like “high performance” toward deterministic performance criteria where an organization commits to a guaranteed service level, such as latency for a specific workload not exceeding 10 milliseconds 99.999% of the time, for instance.
That same demand for predictability extends into capacity planning. As AI workloads become larger and more dynamic, networking infrastructure must be able to absorb rapid shifts in demand without sacrificing performance or efficiency.
“Without dynamic scalability, enterprises are forced into a false choice: either risk performance-killing congestion or engage in massive, inefficient overprovisioning of their network ‘just in case.’ This is incredibly expensive and unsustainable,” Kapil says.
Building this foundation for the world’s most demanding AI workloads is already underway. For example, Tata Communications is collaborating with Amazon Web Services (AWS) to build one of India’s largestAI-ready networks. This high-capacity, resilient network will connect major AWS infrastructure locations in Mumbai, Hyderabad, and Chennai, providing the ultra-low latency backbone needed to accelerate generative AI adoption and cloud innovation across the country.
He points to a consumption-based model, where software allows bandwidth and network functions to scale instantly with demand, as the operational alternative, since it lets organizations pay only for what they use while still protecting performance during spikes.
CIOs and infrastructure leaders need to reframe the network, not thinking of it as a cost center but as something closer to an insurance policy for an organization’s broader AI investment portfolio. An intelligent network de-risks those investments in three ways:
enabling dynamic scalability that removes the need for overprovisioning
strengthening security and governance through the visibility needed to protect data and models
and providing a flexible, programmable foundation that can absorb future compute demands without a full architectural overhaul.
Getting there does not require enterprises to start from scratch.
Choosing a partner with a proven track record is critical. Tata Communications was recently named a Leader in the Gartner Magic Quadrant for Global WAN Services for the 13th consecutive year, reflecting its completeness of vision and ability to execute. That recognition reflects continued investment in areas such as SASE capabilities for AI-driven security and high-capacity 800G services designed for AI-scale infrastructure.
“We recommend a phased approach that begins with assessing the current state of the network and identifying inefficiencies, then prioritizing upgrades in areas such as AI-ready technologies, seamless data exchange, and advanced security solutions,” Kapil says. “Treating the network as a business enabler rather than overhead gives organizations the scalable, secure, and resilient infrastructure the AI economy will continue to demand.”
Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.
Bright Machines wants to solve one of the least glamorous but most consequential problems in the AI buildout: what happens to quality data when a human being has to touch the production line.
The San Francisco-based manufacturer announced today the Hybrid BRC (Bright Robotic Cell), an expansion of its Bright Factory platform that lets human operators step inside a sensor-monitored robotic cell to perform prescribed assembly steps — without breaking the digital record that tracks every server from its first screw to its shipping label.
It sounds like an incremental hardware update. It isn’t. The Hybrid BRC is a direct answer to a structural weakness in high-stakes electronics manufacturing — one that CEO Sviat Dulianinov quantified in stark terms in an exclusive interview with VentureBeat.
“If you assemble modern AI servers starting with manual operations, your initial yield — first-pass yield — can be as low as 20%,” Dulianinov said. “Then you gradually ramp up and scale, and it can reach the 60s, 65% or so.”
When a single AI server can cost hundreds of thousands of dollars, and hyperscalers are burning billions waiting for infrastructure they can’t deploy fast enough, that number is the whole story. The Hybrid BRC is Bright Machines’ attempt to keep human hands in the loop without letting human error back in the door.
Modern automated assembly lines generate a continuous stream of production data — torque values, placement coordinates, component serial numbers, inspection images. That “data thread” is what lets a manufacturer prove a server was built correctly and, when something fails in the field months later, trace the failure back to a specific station, step, or part.
But automated lines inevitably need manual intervention, and until now manufacturers had two bad options when that happened: stop the line entirely, or pull in-process units off to a separate manual workstation that sits outside the monitored data flow. The first choice kills throughput. The second punches a hole in the production record at precisely the moment when human error is most likely to occur.
The Hybrid BRC eliminates that tradeoff, the company says. The cell incorporates guarded access doors and safety panels directly into the production line. When an operator opens the doors, the robotic arm deactivates, and on-screen instructions guide the operator through each assembly step while the cell’s sensor array — cameras, force feedback, and tooling sensors — continues monitoring for incorrect installs, missed steps, and wrong components, applying the same quality checks used during full automation. The traceability record persists at the serial-number level from start to finish.
The economics driving the design become clear when Dulianinov’s manual-assembly figures are set against what automation delivers. “At robotic operations, yield-per-station level is usually more than 98% with our technology, and even at the line level, we usually get to 97.5%, 97.7% or so,” he said.
First-pass yield measures the percentage of units that come off the line correct the first time, without rework. The gap between a 20% manual ramp and a 98% automated station isn’t a rounding error — it’s the difference between profitability and disaster on hardware this expensive.
That math explains the company’s design philosophy for the Hybrid BRC, which treats the human operator as an escape valve for exceptions rather than a substitute for automation. “The more human stations you introduce, the more you increase the risk of lower yields driving the overall yield down,” Dulianinov said. “That’s why we prefer to start at least with 50% automation, and then move to at least 80%.” Speed follows a similar pattern: “On the line level, robots can be faster than humans from like 50 to 100%” in throughput terms, he said.
The AI infrastructure conversation usually revolves around chip supply, power availability, and data center construction. Dulianinov argues that assembly — the unglamorous work of turning chips and motherboards into racked, tested, deployable compute — is a quietly enormous drag on deployment timelines.
“When you have the chips and you have the motherboards, you want to be as fast as possible to deploy that in the data center,” he said, describing greenfield deployments where power and buildings already exist. Getting hardware built, tested, and often rebuilt when quality falls short “could be months,” he said. “With more technology used for this, as our tech, we believe that we can cut it by at least a third.”
A company executive on the call added an anecdotal but telling data point: the servers Bright Machines produces are “flying out into production” rather than sitting stacked in warehouses awaiting deployment — evidence that assembly capacity, not just chips or power, gates hyperscaler timelines. The stakes are asymmetric, the executive noted, because the largest hyperscalers lose millions of dollars per day when servers fail or arrive late. That is why customers are less interested in buying boxes than in buying assurance — and why an unbroken data thread has become a product in its own right.
The Hybrid BRC is not vaporware. Dulianinov said the company already operates a number of the hybrid lines in the U.S. and has “built more than 10,000 compute nodes” through the new stations. This year, he said, Bright Machines plans to manufacture “more than half a gigawatt of compute capacity.”
Who’s buying? Don’t ask. “We cannot unfortunately name customers. That’s the toughest part of our job,” Dulianinov said. “They’re pretty secretive because, as you can imagine, everything data center related is IP related.”
He did offer growth figures: customers grew “more than 3x this year” versus the prior year, driven by what he called the intersection of “physical AI, AI infrastructure buildout, and onshoring.” The demand is spilling into real estate — the company is moving from its 16th Street San Francisco offices to a Burlingame space this fall that executives described as three to four times larger. Overall, the company says it has deployed more than 130 microfactories across 10-plus countries, served more than 60 customers, and produced more than 300,000 servers.
Asked how the Hybrid BRC’s traceability claims stack up against operator-guidance and inspection software vendors like Tulip and Instrumental, Dulianinov drew a sharp line around business models.
“Tulip is just a company that does interface for operators. Instrumental, they focus on inspection. It’s just pieces of the puzzle,” he said. “We, as a technology-enabled manufacturer, we actually run this whole operation… We put our lines, put our software, put our data on the floor, our people, and run it from the beginning to the end.”
The right comparison set, he argued, is contract manufacturing giants like Flex, Jabil, and Foxconn — companies that own the full production process but historically built it on manual labor that generates little data. Bright Machines’ differentiation, he said, is that robot data, sensor data, and now human-station data all flow through one orchestration layer into a single environment the company calls Bright Insights.
That positioning is notable given the company’s origins. Bright Machines was carved out of contract manufacturer Flex eight years ago, and its history has had turbulence: the company planned to go public in 2021 via a SPAC merger at a reported $1.6 billion valuation, according to contemporaneous reporting by The Wall Street Journal and CFO Dive, before the deal fell through. It rebounded in June 2024 with a $126 million Series C — $106 million in equity led by funds managed by BlackRock with participation from Nvidia, Microsoft, Eclipse, Jabil, and Shinhan Securities, plus $20 million in venture debt from J.P. Morgan — bringing its total raised past $400 million, per the company’s announcement at the time.
For technical decision makers, two governance questions loom over any system that instruments human work this closely, and Dulianinov addressed both directly.
On data ownership, he drew a clean boundary: “Everything related to the customer and inspection of their devices and parts obviously would be protected and owned by the customer.” Process and robotics data, he said, stays with Bright Machines to fuel continuous improvement across its platform.
On worker surveillance, he pushed back on the framing. High-IP electronics floors — especially those touching aerospace, defense, or government workloads — already prohibit workers from carrying personal electronics, he noted. “People who know those floors, they know that this is part of the game,” he said, adding that employees “actually appreciate” the traceability because it underpins the security mission: “If you build a data center for the government, and then you build servers somewhere in China, you cannot guarantee how exactly it was built and what component was put there.” In his telling, the monitoring isn’t about watching workers — it’s about being able to prove, component by component, that American-built AI infrastructure is what it claims to be.
The Hybrid BRC‘s modular design carries strategic weight beyond quality assurance. Because the cells are software-defined and snap together like building blocks, Bright Machines says it can retool lines for new hardware generations in days or weeks rather than months — “we can introduce it within a day” for minor design changes within a product family, Dulianinov said, though a jump from air cooling to liquid cooling remains “a big jump.” In an industry where new chip architectures now arrive on a roughly annual cadence, changeover speed is arguably as valuable as yield; a production line that takes six months to retool is obsolete before it amortizes.
But Dulianinov’s closing argument was about labor arithmetic, not machinery. “We need to build in the U.S., and you don’t have 3 million people to bring up manufacturing in the U.S.,” he said, referencing the massive workforces of Shenzhen-scale electronics plants. “So you need to solve it with AI software and robots, and that’s our thesis… It’s not just robots on the floor — it’s also creating jobs. All the robots, and some people on the floor.”
Lior Susan, founder and CEO of Eclipse and chairman and co-founder of Bright Machines, framed the announcement in the same terms: “The future of manufacturing isn’t choosing between automation and flexibility — it’s combining both in the same digital production environment.”
For all the talk of gigawatts and yield curves, the Hybrid BRC amounts to an admission wrapped in an innovation: even in the most automated factories on Earth, humans still have to open the door and reach inside. Bright Machines’ wager is that the winners of the AI infrastructure race won’t be the manufacturers who eliminate the human hand — but the ones who never lose sight of it.
The Model Context Protocol, the open standard that has quietly become the connective tissue between AI agents and the world’s software, is getting its largest update since Anthropic released it twenty months ago — a sweeping architectural revision that its maintainers and backers say finally makes agentic AI ready for massive enterprise production deployments.
The update, released today under the stewardship of the Agentic AI Foundation (AAIF), a directed fund under the Linux Foundation, finalizes MCP’s transition to a fully stateless architecture, hardens its authentication model against a known class of attacks, establishes a formal 12-month deprecation policy, and graduates two headline capabilities — interactive server-rendered interfaces and long-running asynchronous tasks — into official protocol extensions.
The changes may sound arcane. Their consequences are anything but. According to the announcement, running MCP at scale has historically required “sticky routing” or shared state to maintain continuity across sessions — an operational burden that made large production deployments complex even when the underlying capabilities were simple. The new release removes that bottleneck entirely, letting organizations run MCP servers behind standard load balancers using the Kubernetes and cloud-native DevOps tooling they already operate.
“Some people jokingly call it a v2, and I think in spirit that’s accurate,” David Soria Parra, MCP’s co-creator and a lead maintainer at Anthropic, told VentureBeat in an exclusive interview. “It’s probably the biggest change we’ve ever made to the protocol, and with that, it’s a big step up in maturing it for use by really big players.”
To understand why the industry’s largest companies pushed for this release, it helps to understand what was broken. Under the old design, an MCP client — the AI application making requests — had to maintain a persistent session with a specific server instance. In modern cloud environments, where fleets of interchangeable compute nodes spin up and down behind load balancers, that requirement was poison. If the specific server holding your session state disappeared, your agent’s work disappeared with it.
“Before, you needed to have a session store and manage session IDs — and if one of your compute pods went down, all of a sudden the requests would start failing,” said Den Delimarsky, a lead maintainer of the protocol, in an interview with VentureBeat. “That’s not going to be a problem with the new version of the protocol. That’s a huge unlock, and it’s one we collaborated with folks across many companies to put together.”
Mazin Gilbert, executive director of the AAIF and a veteran of Google and AT&T, framed the change in historical terms — comparing it to the architectural decision that made the web itself possible. “That stateless capability enables your MCP client to speak to a load balancer that connects with any server. You don’t need the stickiness,” Gilbert told VentureBeat. “You could not have the internet we have today if my browser couldn’t speak to any website — with any server supporting that connection. You can switch between servers behind a load balancer.”
Gilbert said the constraint had become the primary blocker for companies trying to move AI agents from pilots into production. “I’ve come across companies who are deploying tens of thousands of agents, and you cannot do that without having to go in this direction,” he said. Crucially, he argued, the obstacle was never the AI itself: “It wasn’t the technology, it wasn’t the business case, it was really these fundamental changes that were required.”
The tension is nearly as old as the protocol. A public design discussion opened by MCP co-creator Justin Spahr-Summers on GitHub in December 2024 — just weeks after launch — flagged that MCP’s long-lived, stateful connections were limiting for serverless deployments, and sketched three possible paths forward, including the fully stateless option the protocol has now largely embraced.
Engineers from Vercel, Cloudflare, Shopify, and Amazon weighed in over the following months, a preview of the multi-vendor collaboration that would eventually define the project. The core maintainers formally committed to the direction at a December 2025 meeting on the future of MCP transports, according to the announcement.
Protocol design is a game of trade-offs, and the maintainers were unusually candid about what this one cost. First, payloads get bigger. “A lot of the state doesn’t disappear, but it’s moved back and forth with the server on the wire, at the actual transport layer,” Soria Parra explained. “You get bigger payloads in return for statelessness — but luckily they’re very compressible and very well understood, and still fairly small in comparison to an HTTP request on the web.”
Second, a handful of rarely used capabilities are gone or narrowed. Out-of-band server logging — where a server could push informational log messages to a client at any moment — no longer works in the new model. The team did its homework before cutting it: “As part of the whole exercise, we scraped all of GitHub and looked at who is using it — and it’s basically nobody,” Soria Parra said. Those affected amount to “probably a handful of people — quite literally a handful of people.”
He even allowed himself a moment of engineering self-deprecation. “I’m sad that things I thought were useful turned out not to be useful,” he said. “I think one of the bigger trade-offs was more about my ego than any actual limitation of the protocol.”
Delimarsky argued the shift is less a removal of state than a deliberate transfer of responsibility. “With statelessness, we did shift the responsibility of creating and managing state to the developers — but very intentionally so,” he said. Under the old protocol, “a lot of folks had a hard time understanding: Do I need to use this? Where do I use this? How do I use this? Removing that burden basically says: look, now you can manage state in the way that makes sense for your environment.”
For most developers, migration should be nearly painless, because the vast majority of the ecosystem builds on official SDKs in TypeScript, Python, C#, Rust, Java, and other languages, which will absorb the changes. “One of the key things we constantly do is double-check that the upgrade path is minimal — to the point where any model in the world will probably one-shot it for you,” Soria Parra said — a telling remark in itself, reflecting an era in which protocol maintainers now design migrations to be trivially executable by AI coding assistants.
Perhaps the most enterprise-flavored feature of the release isn’t code at all. It’s a policy. The new formal deprecation framework guarantees developers a minimum of twelve months between a feature’s formal deprecation and its earliest possible removal — the kind of stability contract that lets a Fortune 500 engineering organization commit to a specification without fearing silent breakage.
The number wasn’t picked arbitrarily. “We consulted with folks like Google, Microsoft, and Amazon to find out: in your deployment environment, what’s the right path for making these kinds of changes?” Delimarsky said. “Twelve months seemed like the reasonable middle ground.” He stressed that features are not being torn out on a whim: “It’s not about ripping stuff out of the protocol just because we don’t like it. There’s a very, very strong industry pull behind these changes.”
Soria Parra added that the maintainers’ own telemetry supports the figure — most of the ecosystem upgrades within six to eight months — and stressed that the window functions more as a listening period than a countdown clock. “It just says that in 12 months we are open to remove it, but both Den and I can change our minds based on feedback,” he said. “I think it’s more of a feedback period than a definite period.”
Gilbert sees the policy as one leg of a three-legged stool of enterprise trust, alongside open standards and stateless scale. “There are companies deploying things at a smaller scale, but they’re slowed down because of MCP’s authorization gap, because of identity, because of — do they trust the deprecation policy? Things could change basically any day,” he said. Those companies, he argued, “are going to benefit not because of the statelessness. They’re going to benefit because of the security.”
The release also ships significant authorization hardening, aligning MCP’s auth specification with how OAuth 2.0 and OpenID Connect are actually deployed in practice. Most notably, the protocol now enforces mandatory validation of the issuer (iss) parameter — a protocol-level defense that, according to the announcement, closes an entire class of so-called mix-up attacks, in which a client can be tricked into associating an authorization response with the wrong identity server.
Was anyone actually attacked? No, Delimarsky said — this was preventive engineering, not incident response. “This is not something that is gated in any existing vulnerabilities or active exploitation,” he said. “This is more of us engaging directly with the security community.” The philosophy, he explained, is to borrow rather than invent: “MCP as a protocol is very much establishing the pattern of: we do not want to reinvent the wheel, but we also want to be at the forefront of a lot of the security innovation.”
That posture is most visible in the new Enterprise Managed Authorization extension, developed in close collaboration with identity provider Okta, which lets organizations make their corporate identity provider the authoritative gatekeeper for MCP server access. “If I’m somebody that manages tens, hundreds of MCP servers for my organization, I want to make sure that I enforce some level of common governance, where folks auth with their corporate credentials and not their personal credentials, so that the client doesn’t send data to sources that are unauthorized,” Delimarsky said. Okta bootstrapped the underlying open standard, he noted, and the maintainers then worked “to make sure that it’s adopted ecosystem-wide, and it’s not something that is specific to only one vendor or provider.”
More is coming: Delimarsky said proposals are already on deck for demonstrated proof-of-possession and workload identity federation — capabilities requested by security teams running MCP in production. Gilbert connected the work to a broader maturation: “MCP has now bridged that gap with these authorization protocols, so it’s basically now becoming what we call enterprise ready, versus an open lab sort of experiment.”
Two capabilities graduate to official extension status in this release, taking advantage of a new framework that lets extensions evolve on their own timelines, independent of the core specification — a structural choice that lets the protocol grow without bloating its core.
MCP Apps allows servers to ship rich, interactive, server-rendered user interfaces directly into AI clients — moving agent output beyond walls of text toward dashboards, forms, and visualizations, and dramatically accelerating development of user-facing agentic applications, according to the announcement. MCP Tasks tackles the reality that not every tool call finishes in one round trip. Instead of holding fragile, long-lived connections open while a batch job or heavy computation grinds away, servers now return a durable task handle; clients can disconnect, crash, restart, and resume polling. “You’ve been processing some audio for a podcast or a video — it can notify back the client and say, hey, the task is done. You don’t need to wait and keep the stream open,” Delimarsky said.
A third addition, multi-round-trip requests, lets servers and clients negotiate back and forth within a single logical operation. “It’s not just a one-shot — over the stream, get the input and you’re done,” Delimarsky said. “You can actually interact, server to client, to get the right parameters to execute an action.”
Soria Parra emphasized that these capabilities emerged from the same source as the architectural overhaul: heavyweight production users. “This is a version that came together by some of the best distributed systems experts at Microsoft, Google, and others coming together and working on this for their specific needs — and the needs of the industry at large,” he said.
Anthropic created MCP in November 2024 and donated it to the newly formed AAIF under the Linux Foundation in December 2025, alongside founding projects from Block and OpenAI. Seven months later, the independence question still hangs over the project — and both sides addressed it head-on.
Soria Parra was disarmingly direct about the residual power he holds. As lead maintainer and Anthropic employee, “I do have veto rights, technically,” he acknowledged — “but I think we have never actively used it in any kind of discussion.”
The core maintainer group now spans Anthropic, Microsoft, OpenAI, Google, and Amazon, with contributions from companies like Block, and key decisions “are usually unanimous,” he said. “Technically we have a lot of influence; de facto, we’re not exerting any of it.” He added that governance will progressively broaden: “As the project progresses, we will increasingly move to more different governing structures that include more and more people.”
Gilbert, who has helped stand up multiple foundations during his time working with the Linux Foundation, offered the numbers behind the neutrality claim. The AAIF has grown from roughly 40 members at its December inauguration to 240 today — “the fastest growing foundation” in Linux Foundation history by membership, he said, “signing up one member every day.”
Anthropic’s share of contributions, by his estimate, has fallen below half. “Holding control of a project doesn’t make it an open standard,” Gilbert said. “You have to let go. You have to contribute, and you have to grow the pie and the community. And Anthropic has done an incredible job doing exactly that.”
Notably, the foundation’s membership has expanded well beyond tech vendors into retail, finance, and telecom companies — adopters who, Gilbert says, “are no longer just deploying the protocols. They want a voice, and they want to be at the table to influence the protocol from the get-go, and that’s something we have not seen before.” The roster now includes CERN and, tellingly, Consumer Reports — “because somebody has to defend consumers when this internet of agents comes alive.”
The AAIF is betting that neutrality can hold even amid geopolitical friction. The foundation will host AGNTCon and MCPCon events this fall in Shanghai, Tokyo, Amsterdam, and San Jose, with additional events planned in South Korea, Nairobi, and Toronto, and Gilbert said he is personally investing in growing membership across Asia and India, where he sees underdeveloped growth markets for the foundation.
His answer to the geopolitics question was emphatic model-agnosticism. “We’re completely agnostic to what the model is, whether the model is Kimi, or Gemma, or a frontier model from Anthropic, or from anybody,” he said. “Every model will have to support MCP — whether it is a Chinese model or whether it is a U.S. model, it doesn’t matter. The protocols must be open, standardized.”
The logic is economic as much as diplomatic. Enterprises, Gilbert argued, increasingly pick models “left, right, and center” based on the task at hand — and no model, regardless of national origin, “can provide value to an enterprise 500 customer company unless you have the protocols open, standardized.” In his telling, the foundation exists precisely to provide neutral ground: a place “where competitors who compete furiously during daytime” can “come to a neutral room and debate, converse, align, consolidate, and drive open standards of how the Internet of Agents will evolve.”
That framing echoes his favorite historical analogy. HTTP earned global trust, he said, because of three things: an open standard, stateless scalability, and neutral governance under a standards body. “If I were a Fortune 500 company looking at how I trust the internet, I’d need those three things to fall into place — and they were not in place a year ago. They were not in place even six months ago. But they are in place today.”
The scale of what’s now riding on this specification is difficult to overstate. Soria Parra said SDK downloads have doubled in the past six months, reaching roughly 250 million per week — “which is just insane numbers.”
For context, Anthropic reported 97 million monthly downloads across just the Python and TypeScript SDKs when it donated the protocol in December 2025. Delimarsky pointed to that same adoption curve as his preferred success metric going forward: “There is certainly a certain inflection point where this is no longer just an open source project. This is a substrate for a lot of the agentic workflows that we see across enterprises, across startups, across all sorts of companies.”
Success, the maintainers say, will be measured in server counts on the new specification, in feedback flowing through working groups, GitHub discussions, and the project’s Discord — and in whether the biggest drivers of the changes, Microsoft and Google among them, ship on it. “They are effectively the ones who have been driving a lot of the changes,” Soria Parra said. “Every early indication we have — it looks very, very positive.”
Both maintainers closed on the same note: this release belongs to no single company. “If you look back 18 months ago, when it was an Anthropic-only project, and then 12 months ago, where there was a lot of engagement — now it’s a truly global community,” Soria Parra said. “I’m incredibly proud of what they have worked together.” Delimarsky, “being very unoriginal,” seconded him: the release “would not be possible without a large community of folks that are also volunteering a lot of their own time in making MCP successful.”
Gilbert, meanwhile, is already looking past this release — toward how MCP interlocks with the AAIF’s newly announced Agent Gateway project for traffic management and policy enforcement, and toward agentic commerce, where MCP serves as the discovery layer letting merchants expose products and services to AI agents. The web took thirty years to become invisible infrastructure that billions trust without thinking. By Gilbert’s reckoning, the internet of agents is “in its first, second year” — and as of today, it finally has plumbing built to carry the load.