When enterprise buyers build out their next AI accelerator evaluation list this cycle, they’re more likely to put a non-Nvidia chip on it than Nvidia’s own next-generation GPU. According to VentureBeat’s July VB Pulse survey of 170 AI infrastructure respondents, 39.4% said they’re likely to evaluate non-Nvidia accelerators — AWS Trainium, Google TPU, AMD Instinct, Intel Gaudi or in-house ASICs — over the next 12 months, compared with 25.3% for Nvidia Blackwell (GB300) or other next-generation Nvidia GPUs, a 14-point gap.
Nvidia remains the default in most production environments. But organizations are building real optionality into their accelerator strategy rather than treating Nvidia as the only evaluation worth doing.
The finding sits inside a broader pattern: enterprises are expanding and optimizing the AI infrastructure they already operate before making another major platform change. Greater infrastructure activity did not produce greater urgency to switch platforms. The share of respondents expecting a platform change within three months fell from 38.3% in June to 28.8% in July, even as production adoption, accelerator utilization, and exploration of neoclouds and open-source infrastructure all rose.
The July data shows organizations operating AI infrastructure more intensively and putting more provider platforms into production.
Microsoft Azure posted the largest production adoption growth among the major platforms measured, with the share of respondents reporting Azure in production increasing from 29% in June to 47.1% in July, an 18.1 percentage-point increase. Some of that jump reflects who was surveyed: July’s respondent base skewed more up-market than June’s (57% at organizations above 1,000 employees, versus 37% in June), and Azure adoption rises with company size in both waves. Google’s Gemini was the most-used platform in both waves, with the share of respondents reporting it in production rising from 41.1% in June to 47.6% in July, narrowly ahead of Azure.
The share of respondents reporting OpenAI in production rose from 40.2% to 49.4%. Anthropic production adoption increased from 12.1% to 24.7%.
Among enterprises that operate their own GPUs, the share running at half capacity or less fell from 83% in June (100 respondents) to 69% in July (155 respondents), with the share above 50% utilization rising from 13% to 23%.
The definition of infrastructure effectiveness is also becoming more operational. The share of respondents who selected uptime and reliability as important effectiveness measures increased from 42.1% to 51.2%. The share selecting throughput rose from 21.5% to 24.7%.
Ease of implementation improved from an average rating of 3.84 to 4.04 on a five-point scale. Overall satisfaction moved only slightly, from 4.07 to 4.14, while perceived value was essentially unchanged at approximately 3.9.
That combination is telling. Enterprises are not reporting a dramatic improvement in value simply because they are deploying more infrastructure. They are becoming more capable operators with better architectures, but they are also setting a higher bar for what that infrastructure must deliver, with reliability leading the way.
The strongest counter-signal in the July findings is the declining share of respondents who plan to make an immediate platform change.
The share expecting a change within zero to three months declined by 9.5 percentage points. The share expecting a change within three to six months rose by 4.1 points, while the six-to-12-month window rose by 5.3 points. The share with no planned change remained effectively flat at approximately 40%.
Urgency is shifting outward, with the open-weight-model and open-source-harness debate playing a role in which pieces get enhanced versus fully replaced.
The selection criteria support that interpretation. Integration with existing cloud and data stack was the top factor in both waves, holding steady at 41.1% in June and 40.0% in July. The share of respondents prioritizing performance increased from 24.3% to 35.3%. The share prioritizing cost per million tokens increased from 7.5% to 15.9%, while the share prioritizing access to GPUs rose from 18.7% to 23.5%.
By contrast, the share selecting broad total cost of ownership as a leading factor fell from 34.6% to 21.8%.
The market appears to be moving from general infrastructure planning toward workload-level scrutiny. Buyers increasingly want to know how a platform performs under production inference, how reliably it operates and what each unit of useful work costs.
That 39.4% figure was 31.8% in June, already climbing before this wave. The alternatives enterprises are weighing include AWS Trainium, Google TPU, AMD Instinct, Intel Gaudi and other in-house ASICs.
Interest was even stronger among respondents with strategic purchasing authority, though the C-suite sample is small: the share of C-suite respondents likely to evaluate non-Nvidia accelerators rose from 42.9% (6 of 14) in June to 57.1% (12 of 21) in July. Among final decision-makers, the same interest rose from 35.4% to 50%.
This was especially true for organizations in the small and medium-size business tiers. Among organizations with 251 to 1,000 employees, the share increased from 41.4% to 53.2%. Among organizations with 101 to 250 employees, it rose from 33.3% to 57.7%.
These findings show organizations building optionality into their accelerator strategy.
The increased attention from C-suite respondents and final decision-makers suggests that accelerator diversity is becoming a strategic infrastructure question, not just a technical one for engineering teams.
The infrastructure findings align with a separate VB Pulse survey of agentic context layers. That survey included 101 substantive respondents in June and 101 respondents in July.
The AI harness is the operational layer connecting models to enterprise data, tools, orchestration, evaluation, identity, security, observability and business processes. It determines what an agent can access, which actions it can take and how the organization evaluates its output.
In July, 36.6% of context-layer respondents said they planned to retain best-of-breed standalone tools alongside their models. Another 36.6% expected to mix provider-native runtimes with standalone tools, while only 5.9% intended to build and own the context layer in-house.
Combined, 79.2% of July respondents favored an approach that maintained at least some architectural control outside a single model provider, compared with approximately 65.3% in June. Only 11.9% of July respondents favored consolidating onto a single model provider’s native context stack, down from 20.8% in June.
Most want to preserve provider choice, independent governance or control over critical components around the model.
The need for that control is becoming clearer. In July, 62.4% of context-layer respondents reported that a governed semantic or context layer was either in production or being built. Production adoption alone increased from 24.8% to 31.7%.
At the same time, 68.3% of July respondents reported experiencing at least one confident-but-wrong agent answer caused by missing or incorrect context, compared with 57.4% of June respondents.
The share expecting to use multiple retrieval architectures by use case increased from 12.9% to 28.7%. The share expecting to mix provider-native and standalone context tools increased from 20.8% to 36.6%.
The emerging architecture is a controlled combination of models, infrastructure, retrieval approaches, context systems and operational tooling selected by workload.
Neoclouds are specialized cloud providers focused heavily on AI infrastructure, particularly access to accelerators and supporting services. The July results suggest that these providers are becoming a more credible part of enterprise multi-provider strategies.
The share of respondents expecting to do more with neoclouds increased from 33% in June to 38% in July. At the same time, the share expecting to do less with neoclouds fell from 9.7% to 5.4%.
The movement was especially pronounced among respondents in the technology and software vertical. The share of that July segment expecting to do more with neoclouds reached 57.6%, compared with 44.4% in June.
Current production adoption remains much smaller than broad expansion intent. Across the named providers measured consistently in both waves, such as CoreWeave, Lambda, Crusoe and Nebius, production use increased from 1.9% of June respondents to 5.9% of July respondents.
The difference between 38% expansion intent and 5.9% current named-provider production use may point to a sizable evaluation and adoption pipeline.
The neocloud demand pipeline is not theoretical. CoreWeave reported around $104 billion in revenue backlog at the end of June, excluding more than $25 billion in additional customer commitments secured during early Q3. Nebius does not disclose a directly comparable backlog metric, but said it could sell its entire 2027 capacity under current terms and reported four second-quarter AI cloud agreements, each averaging more than $1 billion in total contract value.
The larger implication is that neoclouds are becoming a viable source of strategic leverage. They give organizations additional options for accelerator availability, software stacks, workload placement and ammunition for negotiations with hyperscale providers.
Neoclouds will still have to demonstrate enterprise-grade reliability, security, support, networking, and data management capabilities. Specialized compute access may open the door, but durable enterprise adoption will depend on the surrounding operational stack.
The most accurate answer is that open-source production usage is growing, while broad platform consideration remains relatively flat.
The share of respondents reporting a custom, self-managed open-source production stack increased from 3.7% in June to 12.9% in July. The stack definition included technologies such as PyTorch, Triton, vLLM, Ray and Kubernetes.
The movement was visible across several segments with July bases above 20 respondents:
Among individual contributors, 23.9% reported production use in July.
Among recommenders and influencers, 13.7% reported production use in July.
Among organizations with 251 to 1,000 employees, 12.8% reported production use in July.
The share of respondents using open-source key-value cache tooling, including LMCache and vLLM prefix caching, increased from 6.5% to 11.8%. Among technology and software respondents, usage increased from effectively 0% to 13.3%.
Open-source platform consideration ticked up slightly but remained essentially unchanged, moving from 5.6% to 6.5%.
This combination suggests that growth is concentrated among organizations moving into implementation rather than across a dramatically larger population of evaluators. Open source appears to be deepening inside an active portion of the market.
Organizations may be turning to open-source components for greater portability, model choice and control over inference optimization. But ownership also transfers responsibility. Teams adopting self-managed stacks must operate upgrades, security, observability, integration and production support themselves.
That combination of more activity, less urgency and more optionality is the throughline across all of it. Enterprises are running more AI infrastructure while deliberately keeping multiple paths open on chips, clouds and the layer that connects models to their own data. The next platform change, when it comes, will be a choice made from a stronger position.
For this article, I compared two independent, cross-sectional infrastructure survey waves: 107 respondents in June 2026 and 170 respondents in July 2026. These waves are not a longitudinal panel, so the findings describe changes between respondent populations rather than changes made by the same organizations. Platform-change timing shares add to slightly more than 100% because a small number of respondents selected more than one window (5 in June, 9 in July).
Sample composition changed between the waves. Respondents selecting the 1–100 employee organization-size category were excluded before calculating results. The remaining wave composition still differed, including a larger July share from organizations with more than 10,000 employees. Month-to-month movements should therefore be treated as directional signals rather than proof of causation. No statistical-significance testing was applied to the comparisons reported here.
The context-layer findings come from a separate survey, with 101 substantive respondents in June and 101 in July. Those results use a different respondent base and are included as supporting evidence, not combined with the infrastructure-survey results.
Enterprises trying to feed PDFs, slides and scanned documents into AI pipelines keep running into the same wall: the tools either miss the structure — tables, charts, layout — or cost too much to run at scale.
Cohere released Parse 5 on Thursday, positioning it on price-to-performance, not raw accuracy — the right cost-capability mix for enterprise scale. Parse 5 is a 2.3-billion-parameter vision language model built to convert PDFs, slides and images into structured Markdown at enterprise scale.
Cohere’s own published benchmark comparison puts Parse 5 behind three larger, general-purpose frontier models on accuracy. GPT-5.5, Opus 4.8 and Gemini 3.5 Flash all score higher than Parse on the three ParseBench dimensions Cohere reports. Cohere is not claiming the top score. It is claiming the best price for a score close to the top.The company priced the model at $1.50 per 1,000 pages through its API, with Model Vault, Cohere’s secure, single-tenant platform for managed inference, available for higher-volume deployment.
“Document parsing isn’t solved because the hard part isn’t reading text, it’s preserving structure and meaning,” Nils Reimers, VP of AI Search at Cohere, told VentureBeat. “Enterprise documents mix tables, diagrams, charts, and formatting that change the interpretation of the data. Most tools still drop structure or hallucinate content, and even frontier models break on layout‑heavy pages.”
Parse 5 takes a page as an image, runs it through a single vision-language model pass and returns structured Markdown, collapsing the OCR-plus-model pipeline most tools run as separate steps.
Architecture. It is a 2.3-billion-parameter vision language model built on Cohere Labs’ North-Micro-Vision-Instruct architecture, with an 8,192-token context window and roughly a 4.6-gigabyte footprint. It accepts a PDF, PowerPoint or JPEG page as a base64-encoded image and returns Markdown in reading order, with tables rendered as HTML, image descriptions and bounding box coordinates for tables and images.
Language coverage. Arabic, English, French, German, Italian, Japanese, Korean, Portuguese and Spanish get stable accuracy, with lower-accuracy zero-shot support elsewhere.
Output modes. The default output returns a Markdown string per page. A blocks mode returns typed elements, where each table carries its own HTML, bounding box and description, the format Cohere positions as what makes citation-level traceability possible for agents.
Availability. Parse 5 is generally available now through the Cohere API, Model Vault, Microsoft Foundry and AWS SageMaker.
ParseBench is a benchmark that scores document-parsing tools against human-verified enterprise pages. Cohere reports Parse 5 scoring 79.2 across three dimensions: tables, content faithfulness and semantic formatting. That puts Parse 5 behind GPT-5.5 (84.4), Opus 4.8 (84.3) and Gemini 3.5 Flash (81.8), and ahead of LlamaParse’s Cost Effective tier (78.3), Mistral OCR 4 (74.5), Databricks AI Parse (72.4) and Azure Document Intelligence (69.3).
Cohere’s table notes two excluded dimensions, Layout and Chart, and attributes both to product scope rather than a performance gap. Parse 5 returns reading-order Markdown instead of per-element bounding boxes for text, and describes charts rather than extracting their underlying data, with chart-data extraction planned for a future version.
Reimers said that design choice reflects where agentic workflows actually break.
“For charts, for example, we provide a general description of the chart together with an indicator, how Agentic AI can visually inspect the chart,” Reimers explained. “Other solutions try to extract the data from the chart, but then miss out critical information (for example, the color or the pattern of a line) that leads to hallucinations in Chat and Agentic AI applications.”
Cost is where Cohere makes its real case. Reimers pointed to a workflow the company modeled for a large financial services firm.
“We ran the numbers for a large financial services workflow that processes 750 million documents a year and showed that choosing Parse 5 over a large general‑purpose model like GPT‑5.5 would reduce costs by more than 98 percent.”
That figure is Cohere’s own estimate for a single modeled workflow, not an audited deployment.
There is no shortage of options for enterprises looking at parsing solutions.
General-purpose frontier models, GPT-5.5, Opus 4.8 and Gemini 3.5 Flash, top the accuracy comparison but carry the cost and latency of a large model on every page.
Then there are specialized parsers, including Mistral OCR 4, LlamaParse and open-weight options like Chandra OCR 2 and RedNote’s dots.mocr.
Hyperscaler document intelligence services, AWS Textract, Google Document AI, Azure Document Intelligence and Databricks AI Parse, compete more on ecosystem convenience than on raw parsing quality, and score lowest in Cohere’s own comparison.
Kevin Petrie, VP of Research at BARC US, said document parsing sits at the center of enterprise AI adoption right now.
“We’re completing a survey now that shows document analysis is by far the #1 use case for AI, with 62% adoption rates among organizations we polled,” Petrie told VentureBeat. “Documents and other unstructured objects, including images and so on, hold the proprietary context that organizations need to differentiate their agentic AI initiatives.”
Petrie added that only time will tell how Cohere’s cost-performance stacks up against frontier models, but strategically his view is that Cohere has the right focus.
Stephanie Walter, Practice Leader for AI Stack at HyperFRAME Research, sees Cohere Parse 5 as sitting in a good spot between legacy OCR and using an expensive frontier model on every page.
“Its potential advantage is delivering structure, spatial provenance and private deployment at a price suitable for high-volume ingestion,” Walter told VentureBeat. “It does not need to win every benchmark. It needs to make reliable enterprise-scale parsing economical.”
“Parsing is the first quality gate in the enterprise AI stack,” Walter said. “If tables, headings, images, or reading order are lost during ingestion, better embeddings and larger models cannot recover that missing structure.”
A benchmark score isn’t the only input that matters here. “Enterprises should test parsers against their own most difficult documents and measure downstream retrieval and task accuracy, not how clean the extracted text looks,” Walter said. “The right question is not ‘Did it read the PDF?’ but ‘Can the agent now use the information correctly?'”
Enterprises that already got burned by an AI agent passing its evals and then failing in production are moving faster toward removing humans from deployment decisions, not slower — even as trust in automated evaluation is rising across the board, new VB Pulse research shows.
In July, 13% of 108 enterprises surveyed said they trust automated evaluation, up from just 5% the month prior. Meanwhile, survey respondents citing poor alignment between tests and real-world results as their biggest concern fell 10 points, from 29% to 19%, month over month.
Yet, 49% of survey respondents said that an AI agent or LLM-powered feature that had cleared company testing subsequently created a problem visible to customers, essentially unchanged from 50% in June. And nearly a quarter, 24%, said this troubling outcome had occurred more than once.
The latest findings from VentureBeat Intelligence uncovered a more troubling phase of the enterprise agent rollout: the gap is no longer only between how much autonomy companies give agents and how well they can verify them. It is increasingly a gap between confidence in the evaluation layer and evidence that the layer is getting better at preventing failures.
The most revealing split appears inside the July data.
Of the enterprises that experienced an AI feature clear testing only to go on to disappoint a customer, 4% placed complete faith in automated checks. Of those that had detected no comparable incident, 24% expressed full confidence — a sixfold difference.
It makes sense: those who experienced test-passing agents failing in live production are, unsurprisingly, more likely to doubt the automated checking process.
Perhaps it makes sense then, that companies geared toward tackling this problem — like automated agent error monitoring and mitigation platform Raindrop.ai — are seeing the market transform wildly from just a few months ago.
“We are seeing the great-decline of evals as we know them,” Raindrop CTO Ben Hylak told VentureBeat in a direct message. “The Fortune 100 are increasingly reducing eval sets and deprioritizing maintenance. As systems grow more complex (MCPs, subagents, etc.) it becomes impossible to fully enumerate the failure cases. Instead, they’re leaning on anomaly and issue detection solutions, both before and after production.”
VentureBeat fielded the July wave among 108 people representing companies with workforces of at least 100. This is down from 157 respondents in June.
Of the 108, 69% described themselves as final AI-buying authorities or people who recommend and influence those purchases. The sample skewed toward midsize organizations: 63% worked at companies with 100 to 2,499 employees.
The findings should be read directionally. The survey is self-selected rather than a probability sample, and the burned-vs.-unburned splits cited throughout this piece rest on groups of 41 to 53 respondents, and other cross-tabs in the report range from 40 to 68.
The industry mix also changed: technology and software participation declined nine points, ending at 14%, while the retail and consumer share added four points and ended at 19%.
The report nevertheless identifies four month-to-month changes worth noticing: more respondents professing complete confidence, fewer naming poor real-world alignment, more choosing integration ease as the decisive buying factor, and Braintrust gaining primary-platform share.
VentureBeat’s June research identified an enterprise evaluation gap: companies were granting agents more authority faster than they were developing reliable ways to test them.
July preserves the key number from that first wave. Across 265 enterprise responses over the two months, the proportion reporting at least one test-approved system that disappointed customers stayed within a single percentage point: 50% in June and 49% in July.
This figure does not mean that 49% of all agent runs fail, or that any particular evaluation product has a 49% failure rate. The survey asks whether an organization experienced at least one customer-facing incident in the previous year after an AI feature passed its internal tests. Companies that deploy far more agents have more opportunities to encounter such an incident.
But that limitation does not make the result less important. An internal evaluation serves as a release gate. If roughly half of surveyed organizations have seen that gate approve a system that later fails in front of customers, a passing score cannot be treated as proof of production reliability.
The cross-tab reinforces the point. Ten of the 41 enterprises with no identified testing miss placed complete faith in automation.
Only two of the 53 previously burned enterprises said the same. Confidence is strongest among respondents with the least evidence that the release gate can fail.
The counterintuitive finding is what companies do after an evaluation miss.
Overall, 67% either let an agent push code or change a system without a person’s approval in certain low-risk cases, or are modifying their pipelines to support that practice during the coming year. That is unchanged from June. In July, 37% already permitted it in limited cases and another 30% were building toward it.
Among enterprises where a test-approved system had disappointed a customer, however, 85% were pursuing that no-approval model, compared with 61% in the group reporting no comparable incident. Only 11% of burned respondents rejected end-to-end deployment automation for the years ahead, versus 24% of unburned respondents.
It would be easy to read that as recklessness, but the data supports another plausible explanation: deployment maturity. Organizations running more agents, at higher volume and across more consequential workflows, are more likely both to encounter failures and to have the engineering infrastructure needed for automated deployment.
The survey cannot establish which explanation dominates. It does establish that a customer-visible incident does not appear to stop the move toward autonomy. Respondents with firsthand proof that testing can miss defects are also moving most aggressively to let those tests authorize production changes.
If their per-deployment failure rate remains constant while deployment volume rises, the total incident count could grow even without the percentage of affected companies increasing. The July data does not measure incident volume, so that remains a risk implied by the pattern rather than a measured outcome.
Pre-deployment evaluation and production monitoring answer different questions. An evaluation asks whether an agent appears ready to ship. Production monitoring asks what the agent is doing after release and whether its live outputs remain correct.
Most companies in the July sample still emphasize whether the system functions, not whether the answer is correct.
Among the 106 valid responses to this question, 26% used inline quality assertions — automated judges or guardrails checking live traffic for output-quality problems. Another 26% focused on transaction traces such as infrastructure spans, token usage and raw inputs and outputs, while 24% mainly tracked gateway metrics such as latency, errors and cost.
Trace and gateway data can reveal outages, slowdowns and broken requests. They may not flag a fluent, fast and confidently wrong answer. Grouped by the report according to what each architecture actually watches, half of respondents monitored whether an agent was functioning, while just over a quarter automatically monitored whether its production output was correct.
The gap is sharpest among the 40 respondents already permitting no-approval deployment in limited cases. Only 28% of that group automatically checked the meaning and correctness of live answers. In other words, most enterprises that have eliminated a person from at least some release decisions have not installed automated semantic-quality monitoring as the production backstop.
This is the clearest operational lesson in the data. A pre-deployment test suite and infrastructure observability are necessary, but they do not cover the same failure mode. Enterprises need a way to detect bad outputs after the agent begins interacting with real users, data and tools — especially when nobody reviews the deployment decision first.
The vendor data offers a more encouraging sign: enterprises are adding dedicated evaluation tools, and specialist platforms are gaining ground.
OpenAI’s native evals and traces narrowly led as the primary platform at 18%, followed by Confident AI’s DeepEval at 17% and Braintrust at 15%.
Anthropic’s Claude Console and Workbench held 12%, tied with organizations reporting no dedicated evaluation platform. Three options each held 6%: internally built tools, Promptfoo and LangSmith.
Braintrust’s primary share increased from 8% in June to 15% in July, the biggest gain by one vendor and the one the report flags as statistically significant. DeepEval rose from 12% to 17%. Use of no purpose-built platform declined five points to 12%, although that smaller change does not by itself confirm a trend.
Because many companies use more than one tool, the broader footprints are larger. OpenAI native evaluation appeared somewhere in 31% of stacks, DeepEval in 27%, Braintrust in 22% and Anthropic’s native tooling in 20%. Custom internal tooling reached 14%, while Weights & Biases Weave and open-source Langfuse each reached 11%.
These are adoption figures, not product-performance scores. The survey does not establish that one vendor produces more reliable agents than another. Still, the results point toward evaluation becoming a distinct enterprise software layer rather than a loose collection of internal scripts or a feature used only inside a model provider’s platform.
Purchasing priorities are changing with that market. The proportion choosing integration ease as the decisive factor climbed 12 points to 39%, displacing cost, which fell from 28% to 23%. Evaluation accuracy ranked second at 28%. The combined average for satisfaction, implementation simplicity and economic value was 3.9 out of five.
The move from price toward integration suggests enterprises increasingly want a tool they can install into existing development and monitoring pipelines now. Yet their leading success metric remains evaluation consistency at 38%, followed by fewer failures and regressions at 20%. Buyers are selecting for fit while still judging results on repeatability.
Switching intent also cooled: 56% still expected to add or replace a platform during the coming year, down from 64% in June.
The proportion staying put increased eight points, reaching 44%. Together with specialist adoption, they suggest some buyers are moving from evaluation to implementation.
The budget data reveals how enterprises are managing the contradiction between greater autonomy and unreliable evaluation.
People-centered review workflows edged narrowly ahead of production observability as the most frequently cited area for increased investment, 31% to 30%.
Automated evaluation pipelines ranked third at 19%, followed by testing for safety and policy compliance at 16%. Only 6% said their reliability and evaluation budget was not increasing.
Among enterprises that had experienced a testing miss, 38% said people-centered review would receive the fastest investment growth, compared with 24% of organizations that had not been burned.
That produces an apparent paradox: the burned group is most likely to remove people from the release checkpoint and most likely to increase spending on people elsewhere in the process. The strategy appears to be automation with a human backstop — allow agents to move faster, then use reviewers to catch what automated evaluation misses.
The open question is whether that model scales. Agent deployments and automated checks can grow with software volume. Reviewer hours do not fall at the same rate. Enterprises may therefore be replacing a human approval step with a larger downstream review function rather than eliminating human oversight.
July’s data does not show that enterprise agent evaluation is failing everywhere, nor does it prove automated judges are getting worse. It shows something more precise: confidence rose before the measured failure incidence improved.
At the same time, the infrastructure around evaluation is maturing. More enterprises are adopting specialist tools, integration has become the leading purchase criterion and companies that have already experienced failures are increasing investment in human review. The market recognizes the problem and is spending against it.
But the central reliability result remains stubborn. Nearly half of surveyed enterprises still report that an AI feature cleared internal checks before disappointing a customer. Respondents with that experience place less faith in automation — and move faster toward deployments with no human approval.
The report frames this as an incomplete verification model: a passing pre-deployment score marks the start of monitoring, not the end of it. For most enterprises, the production quality checks and evaluation testing that would close that gap still aren’t in place.
A company builds a governed context layer specifically to stop its AI agents from confidently giving wrong answers. Once that layer is live, the company is more than twice as likely to report the failure happening — not less.
In the past six months, 68% of enterprises have traced a confident but wrong AI agent answer to missing or inconsistent business context. Thirty-seven percent say it happened more than once, ahead of the 32% who saw it happen only once. The figures come from a VB Pulse July 2026 survey of 101 qualified enterprises with more than 100 employees. That’s up from 57% in a VB Pulse survey conducted in June. Recurring failures climbed too, from 31% then to 37% now.
This is the second time VB Pulse has asked enterprises this exact question, once in June and now in July. The failure rate is climbing, not falling, even as more enterprises report a governed layer in production, up from 25% in June to 32% now.
Every AI agent needs some way to know what the business actually means, whether a metric is defined consistently, whether a document is current. That’s the operation. The challenge is that enterprises hand agents that context in very different ways, and those ways are not equally reliable.
Retrieval over documents remains the most common approach, the primary source for 31% of enterprises. But a real share of enterprises skip a structured approach altogether. Thirteen percent run agents primarily on long-context loading, feeding documents directly into the model’s context window rather than retrieving them. Five percent give agents no structured context at all, just the model’s general knowledge. Between them, nearly one in five enterprises are feeding agents business context by brute force or not feeding it at all.
Even the leading approach can still produce a confidently wrong answer. Retrieval works by matching a question to text that looks similar in meaning. Similar wording doesn’t guarantee the same meaning. Srijith Rajamohan, an AI research leader at Redis, described exactly this gap in an interview with VentureBeat earlier this year.
“If you have a sentence like ‘Rome is closer than Paris’ and another that says ‘Paris is closer than Rome,’ and you do an embedding retrieval followed by a text search, you’re not going to be able to tell the difference,” Rajamohan said. “The same words exist in both sentences.”
The way enterprises choose a retrieval system doesn’t help close the gap. Access control and permissions now tie ease of data ingestion as the top selection criteria, at 24% each. It’s the first time in this survey series that a governance property has led to the buying decision. Retrieval accuracy trails at 15%. The property most directly tied to a confident wrong answer isn’t the property most enterprises are buying for.
Once a system is running, correctness is still how enterprises judge it. Response correctness is the primary success metric for 38% of enterprises, twice the next closest answer, security and access control at 19%. Enterprises are shifting how they buy toward governance. They’re still grading success on whether the answer is right.
A governed context layer is meant to fix this. It’s one shared, agreed-on model of what the business’s data means, that every agent and BI tool references instead of guessing on its own. Adoption is far from settled.
Thirty-two percent of enterprises run one in production. Thirty-one percent are piloting or building one right now. Twenty percent are evaluating one. Fourteen percent have no plans to, and 4% don’t know.
Compare that adoption data against who’s actually had the failure, and the picture inverts. Among the 91 enterprises able to say whether they’d experienced the failure at all, those running or building a governed layer report it recurring at 50%. Those without one report it at 21%.
A governed layer doesn’t cause the failure — it’s what makes the failure visible in the first place. Tracing a bad answer to a broken definition or a stale table requires a shared, governed reference point. A context layer provides that. Without one, the same wrong answer still happens — it just gets chalked up to the model, or never gets traced at all.
The pain point predates AI by decades. Kyle Nesbit, founder of the semantic layer startup Credible Data, described it to VentureBeat last month. “It’s the same pain point people have had for 30 years, the lack of governed data analysis,” Nesbit said. “Now with AI, it’s the same problem, but orders of magnitude more chaos and pain.”
Company size sharpens the same point. Enterprises with more than 1,000 employees report recurring failures at 55%, against 30% for those between 101 and 1,000 employees. That’s despite the bigger companies being less likely to have a layer already in production, 24% against 37%. More instrumentation and more people asking why a number was wrong turns up more failures, not fewer. A clean record is not evidence of a healthy context layer. It’s at least as likely to be evidence that nobody’s checking.
Here’s what this adds up to for enterprises building on this layer.
Retrieval alone will not close the context gap. RAG remains the default context source, and nearly one in five enterprises are running agents on long-context loading or no structured context layer at all. More documents or a bigger index doesn’t fix a definition that means two different things in two different systems.
The budget is moving faster than the infrastructure is shipping. Sixty-three percent of enterprises are already building or running a governed context layer. Only 32% have actually gotten one into production. That gap is where the spend is going, not where the problem has been solved.
A clean failure record is a red flag, not a green one. The 22% of enterprises reporting no context failure at all are not the best-governed group. They’re the group least likely to be checking. The size data backs this up directly. Larger enterprises report recurring failures at nearly twice the rate of mid-market peers, despite being less likely to have a governed layer in production, not more.
No one is planning to hand the layer to a single provider. Seventy-nine percent of enterprises intend to keep at least part of the context layer outside any one vendor’s stack, split between best-of-breed tools and an explicit mix. Just 12% plan to consolidate onto a single provider’s native context stack.
The finding that organizations aren’t likely to hand over control to a single provider is a theme that VentureBeat has reported on consistently this year. Michael Ni, an analyst at Constellation Research, put it bluntly earlier this year when DataHub’s context layer push first landed.
“Whoever controls runtime context, controls the AI decision layer for enterprise data,” Ni said.
Presented by MongoDBBuilding AI that is accurate, secure, and reliable is a major engineering feat for organizations subject to the compliance obligations that govern healthcare, financial services, and transportation. The challenge of delivering AI-dr…
Skan AI, a startup that builds what it calls a “context graph of work” by observing how employees actually perform their jobs across enterprise software, has raised $63 million in Series C funding co-led by Cathay Innovation and Dell Technologies Capital, the company announced Wednesday.
Citi Ventures, Bloomberg Beta, State Farm Ventures, and Wipro Ventures also participated in the round, which brings the seven-year-old company’s total funding to roughly $120 million. Alongside the raise, Skan is announcing the general availability of two new products — Skan AI Blueprint and Skan AI Agents — that, together with its existing Skan AI Intelligence offering, form a complete platform for discovering, modeling, and ultimately automating enterprise workflows.
The announcement lands at a moment of deep frustration in enterprise AI. Companies have poured billions into generative AI pilots, but the results have been dismal: Gartner research cited by the company finds that only 8% of enterprises have AI agents in production, and 95% of early implementations will require a complete redesign. Those figures echo an MIT report last year, covered by Fortune, which found that roughly 95% of enterprise generative AI pilots were failing to deliver measurable returns.
Avinash Misra, Skan’s co-founder and CEO, believes the industry has misdiagnosed the problem. The models are fine, he argues. What they lack is an accurate picture of the businesses they are being dropped into.
“Everyone is obsessed with building a better driver,” Misra told VentureBeat in an exclusive interview ahead of the announcement. “We think the bigger opportunity is building a better navigation system.”
The standard playbook for grounding AI agents — feeding them process documentation, standard operating procedures, and system logs — is built on a fiction, Misra argues. The way work is documented and the way work actually happens inside a large enterprise are two different things, and the gap between them is precisely where agents fail.
That gap is what sent Misra and co-founder Manish Garg down this path seven years ago, long before agents were a boardroom obsession. “Why is it so difficult for an organization, and a large enterprise especially, to understand how its own work actually gets done?” Misra said. “Why does it need to fly in McKinsey consultants for that?”
The question has only grown more consequential as enterprises race to operationalize AI. Frontier models arrive at the company door brilliant but blind, with no knowledge of the exceptions, decisions, handoffs, and institutional habits that define how a claims department or a compliance team actually operates. Every company now stuffing agents with documentation and logs, Skan contends, is discovering the same uncomfortable truth: the source data was never the whole story. And a source data problem cannot be fixed downstream.
Skan’s answer is to go to the source itself. The company deploys observation technology on employee desktops that continuously watches how work moves across applications — the spreadsheet, the CRM, the email client, the 40-year-old mainframe — and abstracts those observations into a living model of the underlying business process.
“Think of it this way: if I were to share my screen here, and you were to observe my screen going from Excel sheet, CRM system, email client, in about two iterations you’d build a model of what I do,” Misra said. “Except you couldn’t do that at scale. You couldn’t do it 24/7, and for 1,500 people like me. Now replace yourself with our technology.”
That framing also explains how Skan positions itself against process mining vendors like Celonis, which reconstruct workflows from the data trails left in backend systems. System logs, Misra argues, only capture completed transactions — not the messy human work that produced them.
“All backend data, by definition, is a committed state of work. Work is really what happens between those committed states,” he said. “Eighty percent of what you’re interested in, from an AI point of view, in execution of work, actually lies between those systems.”
The screen, in Skan’s view, is the one place where everything converges. “It brings together human agency, it brings together the entire application landscape, and it brings together the data that matters,” Misra said. Two decades of user interface design have quietly buried enormous amounts of process knowledge in the space between a worker’s eyes and their monitor; Skan’s pitch is to bring that hidden layer back to the surface.
But watching, he insists, was never the hard part — a point aimed squarely at the incumbents who might be tempted to copy the approach. “The hard problem is not screen observation,” Misra said. “The hard problem is abstraction of what you see on the screen — the intent extraction.” A human watching a colleague’s screen can instantly tell whether a jump back to step one means a new case or rework on an old one, because humans understand the signature of the work. Teaching a model to make that same judgment, statefully and at enterprise scale, is where Skan believes its seven-year head start lives.
The result is a context model that AI can reason over and act on — the raw material for the agents that now sit at the top of the company’s product stack, and the foundation for everything else the platform does.
An approach built on continuously watching employee screens invites an obvious objection, and it is not a hypothetical one. In June, Reuters reported that Meta scaled back an internal tool that tracked employee mouse clicks after workers raised concerns — a sign that even AI-forward companies are wary of the line between operational telemetry and surveillance.
Misra says he heard the objection before he wrote a line of code. When he first pitched the concept to Delphine Icart, then chief transformation officer at AXA Mexico, her reaction was blunt. “Delphine’s first words to me were, ‘This sounds like a great idea, but you are dead on arrival,'” Misra recalled. “‘You are observing things that you shouldn’t be observing — the privacy of my operators, and the sovereignty of my data on those screens.'”
That conversation, he says, shaped the architecture. Skan aggregates rather than individuates: the system surfaces statistical patterns across hundreds of workers performing the same process, not the behavior of any one of them. “We’re not interested in what John is doing at 10 hours and 43 seconds,” Misra said. “We are interested in what hundreds of Johns put together — what are the statistical and the semantic decisions that they are making in that business process?”
Organizations control what the technology can see through an opt-in scoping model — specific applications and URLs, nothing else — and the data Skan produces never leaves the enterprise firewall. A three-tier architecture sends only anonymized metadata to the cloud. Misra points to deployments approved by European works councils, among the most privacy-protective labor bodies in the world, as evidence the model holds up under scrutiny — and credits it for clearing security review at institutions where most AI tools cannot operate.
Whether aggregation fully defuses the concern is likely to remain contested. The same telemetry that reveals a broken process can, in principle, reveal an underperforming team, and Misra acknowledged that the technology has led some customers to reduce headcount in certain processes.
Skan claims more than $500 million in cumulative customer value to date, a figure worth unpacking. Pressed on whether that represents realized savings or projections, Misra was direct that it is an envelope, not a bank balance.
“The number comes from the cumulative, across all our customers, of the quantified savings that we have brought to them — the savings that they have expected they would save,” he said. “Now they are on the roadmap of recouping those savings through a variety of interventions,” including process redesign, technology changes, and, increasingly, AI agents. In other words, $500 million is identified opportunity, some portion of which has been captured.
The more concrete evidence comes from individual deployments. At one top U.S. bank, according to the company, Skan observed 11.2 million context switches across 1,500 finance professionals and uncovered $37 million in operational friction. Turning those observations into agent-executable context cut cost per transaction by 32%, lifted throughput by 41%, and delivered $18 million in annualized savings.
Misra pointed to an anti-money-laundering operation at one bank where “60% of the cases are now being run by AI agents,” adding that the results surprised even him: “The accuracy of those agents surpasses many times over the accuracy of humans. It’s not just an argument of efficiency; it has also become an argument of quality.” Among insurers, he said, Skan typically delivers roughly 25% productivity uplift in core claims processes; one customer doubled its case volume over the past year without adding a single claims specialist.
Skan’s publicly referenceable customers include Unum, the $13.8 billion employee benefits provider, and Mitie, the U.K. facilities management company, whose chief technology and digital officer, Cijo Joseph, said Skan’s technology “gives us unprecedented operational visibility that has dramatically accelerated our AI transformation.” The company declined to share revenue but said it grew more than 300% year over year — for the second consecutive year — with net dollar retention around 150%, and now counts seven of the ten largest U.S. banks and a quarter of the Fortune 50 as customers.
Skan’s thesis rests on observing how work actually gets done — which raises an uncomfortable question. Real employees make mistakes, take shortcuts, and entrench inefficiencies. What happens when the context graph faithfully encodes bad process?
Misra’s answer reaches for the most famous precedent in modern AI. “Think for a moment what OpenAI did,” he said. “OpenAI took the totality of the world’s text and fed it into a transformer architecture, and semantic understanding emerged. OpenAI’s model has seen bad language and has seen good language, and yet it is able to have semantic understanding.”
Skan, he argues, does the analogous thing with work: treat business process execution as a language, where process steps, screen features, and handoffs stand in for words and sentences. Fed enough end-to-end executions, the model learns the full distribution of paths — efficient ones, slow ones, compliant ones — without assuming any single path is best. “The longest path may be the best path, because it is more compliant,” Misra said. An organization then constrains the model along the axes it cares about, and the model returns the path that satisfies them.
“It is not record and play — and that’s the fundamental difference between us and a lot of our competition, UiPath and so on,” he said. “It is fundamentally creating an AI model that understands work, and then constraining that model.”
He offered a concrete illustration of what that unlocks: at one large bank, Skan’s telemetry continuously compares live case execution against a 600-page controls inventory, with agents that trigger alerts when cases miss required compliance steps — turning a document no human could hold in their head into a real-time enforcement layer. It is the kind of application that only becomes possible, Misra argues, once a model genuinely understands the work rather than merely replaying it.
Skan sits at the intersection of several crowded categories, and its answer to each competitor is a variation on the same theme: scope. Process mining vendors see only what the logs record. RPA incumbents replay tasks without understanding them. And the platform giants — ServiceNow, Salesforce, Microsoft — are shipping capable agents whose vision ends at their own walls.
“The context that these agents have access to is limited to ServiceNow, limited to Salesforce, whereas work spans processes across the board,” Misra said. “Creating a customer entry is a task. To receive an email and decide whether a customer entry has to be created, or something else — that is the process, and that’s what we are after.”
The deeper strategic argument, and the one that seems to resonate with Skan’s regulated customer base, is about differentiation in a world where every enterprise has access to the same frontier models. “If every insurance company, every bank had access to the same models, then the outcomes will asymptotically decay to the outcome of the model,” Misra said. “Historically, you have competed and differentiated in the way you have organized work. That old word — process — now comes back as context for AI. But that context is protected by you. It’s not part of the model.”
That logic explains both the company’s posture toward the model makers — “the more they are successful, the more power we have,” Misra said, disclaiming any ambition to compete with them — and the Nvidia partnership featured prominently in the announcement. Skan runs on Nvidia AI Enterprise and NIM microservices, and Misra described growing demand for private appliances that can observe work, hold the context model, and execute agents entirely inside a customer’s own infrastructure. It also fits the market’s direction: venture investors surveyed by TechCrunch at the end of last year predicted enterprises would spend more on AI in 2026 but through fewer vendors — a consolidation that favors Skan’s decision to ship discovery, intelligence, and agents as a single closed loop.
Misra argues that loop matters more, not less, as automation scales, because agents demand oversight in a way humans never did. “It is an irony of sorts,” he said, “that you’ll probably need much more observation and much more understanding of work in an automated way than you would with humans.” The bet embedded in this round is that work context becomes foundational infrastructure for enterprise AI the way CRM became the system of record for customers — a comparison Cathay Innovation partner Simon Wu made explicitly, calling Skan “one of the defining platform companies of the next decade.”
Misra put the stakes more simply. “You cannot retrieve context that you do not capture,” he said. “The battleground is shifting from the smartest model to knowing how your company actually works — because everyone will have access to the smartest model.”
The frontier labs, in other words, can keep their arms race for the better driver. Skan just raised $63 million on the conviction that the money is in the map.
Presented by MongoDB We have been building databases as an industry for roughly 60 years. We have been building AI agents, in the form most people mean when they say the word today, for about 18 months.Sit with that ratio for a second, because it expla…
A VB Pulse survey this June found that 57% of enterprises had traced a confidently wrong agent answer back to missing or inconsistent context — the latest sign of how central context has become to whether AI agents can be trusted to act on their own.
Most of the fixes so far have solved a narrower version of that problem: one agent remembering more, in one session. What’s been missing is a way for a team of agents to draw on the same context at once, and that gap is where a newer problem is surfacing. Once an agent’s context is shared across a whole team, a wrong fact doesn’t cost one person a repeated explanation. It costs the whole team.
Tencent’s answer to that gap is Agent Memory, an open-source project the team said grew out of six months spent fixing a narrower problem: agents losing context in long sessions. Part of that system is a persona layer, a stable, distilled picture of who a user is and how they work, built up over many conversations rather than reconstructed each time. On Tencent’s own benchmark for whether an agent still applies that picture correctly after extended use, accuracy rose from 48% to 76%, a 59% relative improvement, once the persona layer was added. This week, Tencent extended that project with the beta launch of Team Memory, which opens the same approach up to a whole team instead of one agent. Tencent said the repo hit No. 1 on GitHub’s TypeScript trending list this week.
Agents on a team can now read from a shared memory hub instead of keeping separate, siloed context, governed through an access control layer that determines who can read what.
The core idea is a shared hub rather than a shared prompt. Instead of pasting one large context block into every agent’s window, Team Memory registers four kinds of reusable assets and equips each agent with only the ones it needs.
Chat Memory. Retains preferences, facts, decisions, and interaction history, distilled through four layers, from raw conversation up to a stable long-term persona, so an agent does not need to be reintroduced to a user it has already worked with.
Skill. Captures procedures pulled from completed work, versioned and reviewed before they are shared rather than dropped into a folder as-is.
LLM-Wiki. Turns documents and specs into structured, linked pages.
Code-Graph. Indexes a codebase’s symbols, files, and call relationships so an agent can check what a change might affect before making it.
Tencent’s documentation draws the distinction directly: “RAG answers ‘what can be found?’ Team Memory also answers ‘who can use it, which version is valid, and which Agent should receive it.'”
In practice, that’s what Tencent calls an “Agent Loadout”: a Scout agent doing research can be equipped with market research and competitive analysis assets, while a Builder agent gets the code graph and product docs it needs instead, rather than every agent getting access to everything.
Which assets an agent gets equipped with is governed through four visibility tiers:
Private. Readable only by the asset’s owner.
Team. Readable by anyone on the team.
Restricted. Gated by user, role, or agent-level access control.
Agent. Equipped to one specific agent within a team.
New assets default to private, so sharing has to be a deliberate action rather than something that happens automatically.
That access model answers a real question, who is allowed to read a given memory asset. It does not answer a second one, which is what happens once a memory asset turns out to be wrong. Tencent’s own documentation lays out ownership, versioning, and status tracking for each asset, but nothing in the documentation describes a correction or expiry process for a fact that’s already been read and reused by other agents on a team, or a way to resolve it when two agents’ memories of the same thing disagree.
That gap is what practitioners flagged within hours of the launch post.
“Shared memory makes the write path the interesting problem. Retrieval gets most of the attention, but a wrong fact written once now propagates to every teammate’s agent instead of just yours. Curious how the governance layer handles correction and expiry,” Blake Murphy wrote on X.
The concern wasn’t only about fixing a bad fact after the fact. It was about the decision to leave something out of the record in the first place. “the governed part is the hard part. once teammates’ agents can read each other’s context, someone has to decide what never gets written down,” Virgil Maro wrote on X.
Others pushed further into what happens once two agents’ memories actively contradict each other, not just go stale.
“The Code-Graph plus LLM-Wiki split is the right call. The part I’d want to see benchmarked: in shared mode, whose memory wins when two teammates’ agents have written contradicting facts about the same module? Single-agent memory drifts slowly. Shared memory drifts fast, because one stale write propagates to people who never saw the session that produced it,” Austin Green wrote on X.
The reaction wasn’t uniformly critical. “Interesting shift: making memory a shared service turns agents into a real team rather than isolated bots. Governance will be the trickiest part, especially when facts conflict,” Moez Zhioua wrote on X.
None of these are edge cases specific to Tencent’s implementation. A March 2026 paper on production multi-agent memory architecture, “Governed Memory: A Production Architecture for Multi-Agent Workflows,” published independently of any single vendor, identifies governance fragmentation and silent quality degradation without feedback loops as structural risks in shared multi-agent memory generally. The pattern the paper describes matches what the commenters above pointed at directly: a wrong fact in a single-agent memory system costs one user a repeated correction, while the same wrong fact in a shared, team-wide memory system propagates to every agent that inherited it before anyone catches it.
AI agent memory work in 2026 has mostly focused on a single agent remembering more, in one session, about one user: LangChain’s LangMem SDK, Google’s Always On Memory Agent, and Anthropic’s work inside the Claude Agent SDK all work this way. A different line of work has focused on giving agents access to a shared model of business data. VB’s own June survey found only 25% of enterprises had that kind of governed context layer in production, while vendors including AWS, Couchbase, Oracle, Redis, and Pinecone have all shipped versions of it this year.
Team Memory’s closest existing comparison is likely Asana, which built shared memory across a company’s AI teammates so an agent doesn’t need to be re-briefed on context another agent already has. Asana’s CPO described the same tradeoff Tencent’s practitioners are now raising, an access control system built specifically to stop one agent’s memory from leaking into a project another agent isn’t cleared to see. Tencent’s version is open-source and portable across frameworks rather than scoped to one platform, but it’s answering a question Asana’s team already ran into while building a closed one.
For teams evaluating this category, the upside is real: agents stop relearning what the team already knows. The tradeoff is just as real: one bad write is no longer contained to one agent — it’s inherited by every agent that reads from the shared pool, with no correction or expiry process yet in place to catch it.
The AI agent observability space is taking off — but how can enterprises be sure what observability products and solutions they need?
Observability startup groudcover (lower case “g” intentional) announced this week that it raised $100 million in a round led by One Peak, bringing its total funding to $160 million.
The company says it has more than 250 paying customers, tripled annual recurring revenue over the past year and is increasingly replacing established observability platforms inside enterprise environments. Those are company-reported figures, but together they point to growing momentum in one of enterprise software’s most competitive markets.
That market has long been dominated by companies including Datadog, Dynatrace, New Relic, Splunk and Grafana. Between them, they represent billions of dollars in annual revenue and years of product maturity. Breaking into that group has never been easy.
groundcover’s argument is that artificial intelligence has fundamentally changed the assumptions those platforms were built on.
Rather than competing feature for feature, the four-year-old company is trying to convince enterprises that the architecture underpinning observability itself needs to change as AI systems become more autonomous, produce vastly more telemetry and increasingly participate in software operations. Whether that thesis proves correct remains an open question, but it offers a compelling lens through which to examine how observability is evolving alongside enterprise AI.
Observability has traditionally been viewed as a post-production discipline. Engineers deploy applications, monitor logs, metrics and traces, investigate incidents, and improve reliability over time.
That workflow is changing.
AI-assisted software development has dramatically accelerated deployment cycles. Coding assistants generate more code, infrastructure evolves more rapidly, and organizations are deploying increasingly complex distributed systems that combine microservices, Kubernetes clusters, APIs and large language models. At the same time, enterprises are beginning to operate AI agents that execute multi-step workflows, call external tools and interact with production systems.
Each of those activities generates telemetry.
The result is an explosion of operational data that organizations increasingly want to retain rather than discard. AI applications introduce additional layers of observability beyond traditional infrastructure monitoring, including prompt execution, model latency, token consumption, retrieval pipelines, tool invocations and agent behavior. As enterprises experiment with autonomous systems, that telemetry becomes increasingly valuable because it provides the context needed to understand what an AI system actually did and why.
For many organizations, this creates tension with pricing models that charge according to the amount of data ingested.
Historically, engineers have often responded by sampling traces, shortening retention periods or limiting which data is collected. Those approaches reduce costs, but they also reduce visibility precisely when AI-driven systems demand more complete operational context.
“We’ve seen telemetry exploding,” groundcover co-founder and CEO Shahar Azulay said during a recent media briefing. “Users are frustrated by not getting all the value from Datadog and similar platforms. They’re limiting the data, siloing it, sampling it.”
Whether that frustration is widespread enough to reshape the market remains to be seen, but the underlying trend is difficult to ignore. AI is making observability less about collecting enough data and more about collecting everything organizations may eventually need.
Many observability vendors have introduced AI assistants, AI-powered root cause analysis and AI observability features over the past two years. Datadog, Dynatrace, New Relic and Grafana have all announced products aimed at helping enterprises monitor AI applications or automate operational tasks.
groundcover acknowledges those developments but argues they do not address what it sees as the more fundamental issue: where telemetry lives and how customers pay for it.
Instead of operating a conventional SaaS platform that stores customer telemetry in vendor-managed infrastructure, groundcover uses what it calls a bring-your-own-cloud (BYOC) architecture.
Customers keep the data plane—including telemetry storage and processing—inside their own AWS, Microsoft Azure or Google Cloud environments, while groundcover provides a managed control plane and user experience. A fully self-hosted deployment option is also available.
While some competitors, including Datadog and a few other observability vendors, do offer limited hybrid or customer-controlled data residency options, these are generally not equivalent to a full BYOC model. In most cases, telemetry is still processed and stored within the vendor’s managed infrastructure, with only partial controls (such as regional data residency, private links, or selective log forwarding) available.
That architectural decision influences nearly every aspect of the company’s strategy.
Because customers already pay for their own cloud infrastructure, groundcover argues it can avoid charging based on telemetry ingestion. Instead, pricing is based primarily on monitored hosts, regardless of telemetry volume.
The company believes this changes customer behavior.
Rather than deciding which logs or traces are too expensive to keep, organizations can theoretically retain complete telemetry and use it for operational analysis, compliance and AI-assisted troubleshooting.
“We don’t price by data volume,” Azulay said. “We price by the size of the infrastructure.”
The distinction matters because AI workloads tend to increase telemetry far faster than infrastructure itself.
That does not necessarily make host-based pricing universally cheaper. Organizations with relatively light workloads spread across many hosts may find different economics than dense Kubernetes environments generating enormous amounts of telemetry. The company’s own briefing notes that per-host pricing is most advantageous for organizations with high telemetry density and may be less compelling for lightly utilized fleets.
Still, the broader argument is less about cost alone than predictability. Enterprise infrastructure teams often struggle with observability bills that fluctuate alongside application growth. groundcover’s model attempts to align pricing more closely with infrastructure planning rather than data generation.
The second pillar of groundcover’s strategy is eBPF, a Linux kernel technology that has rapidly become one of the most important building blocks for modern cloud observability.
Instead of requiring developers to manually instrument applications, eBPF allows software running inside the operating system kernel to observe network traffic, system calls and application behavior with minimal code changes.
That enables faster deployment and broader visibility across infrastructure.
For organizations operating Kubernetes clusters and cloud-native applications, reducing instrumentation complexity can significantly shorten deployment times while increasing telemetry coverage.
Azulay argues this becomes especially important as AI systems generate increasingly complex interactions across services.
“Our sensor allows us to observe systems very deeply from infrastructure to application to AI workloads without developers needing to instrument code,” he said during the briefing.
eBPF itself is hardly unique. Many observability vendors now incorporate it into their platforms.
What groundcover argues differentiates its approach is combining automatic eBPF collection with customer-controlled storage, OpenTelemetry compatibility and unified pricing inside a single platform.
The company’s own research briefing acknowledges that none of these technologies individually represents a competitive moat. The claimed differentiation lies in the combination of eBPF-first collection, managed BYOC architecture, host-based economics and full-stack observability delivered together.
Perhaps the most interesting aspect of groundcover’s strategy extends beyond traditional monitoring.
The company increasingly describes observability as infrastructure for autonomous software development.
Historically, observability platforms have served human operators investigating production incidents.
groundcover believes future observability platforms will increasingly serve AI agents as well.
Its Agent Mode product allows engineers to investigate incidents using natural language across logs, metrics, traces and Kubernetes events. More importantly, Azulay envisions observability becoming the feedback mechanism that informs coding agents about what actually happened in production.
Rather than simply detecting failures after deployment, observability becomes continuous operational context that autonomous systems can use to evaluate changes, identify regressions and eventually recommend or implement fixes.
“We’re seeing observability moving from being a post-production tool… to people taking context from production and feeding it back to their coding agents so they can write code better,” Azulay said.
Today, the company emphasizes that humans remain in the loop.
Agent Mode investigates incidents and surfaces recommendations, but production changes still require human approval. Azulay expects autonomy to increase gradually as organizations become more comfortable allowing AI systems to participate in operational workflows.
That vision reflects a broader trend emerging across enterprise software, where AI agents increasingly span development, testing, deployment and operations rather than functioning as isolated assistants.
groundcover is entering an intensely competitive market populated by vendors with decades of enterprise experience.
Datadog alone generated more than $3 billion in annual revenue in 2025. Dynatrace, Cisco’s Splunk business, Grafana Labs and New Relic all maintain extensive partner ecosystems, mature integrations and enterprise support organizations that newer entrants cannot easily replicate.
groundcover is not attempting to outscale those incumbents overnight.
Instead, it argues that AI creates an architectural inflection point similar to previous transitions from on-premises infrastructure to cloud-native computing.
According to Azulay, many customers initially adopt groundcover to reduce observability costs but increasingly remain because they want unrestricted access to richer telemetry and AI-native workflows.
He says deployments typically replace incumbent platforms rather than operate alongside them, although the company has not publicly disclosed customer migration data or independent studies validating that claim.
The company’s journalist briefing also urges caution around some performance claims.
Revenue growth, customer counts and enterprise adoption figures originate from groundcover itself. Published customer case studies reporting significant cost savings are vendor-authored and should not be treated as independent validation without additional evidence. The briefing also recommends scrutinizing exactly what metadata leaves customer environments in standard BYOC deployments, rather than assuming that no operational data ever reaches vendor infrastructure.
Those caveats are important because the observability market has become crowded. Gartner currently tracks more than one hundred observability products, and nearly every major vendor now markets AI-powered operational capabilities.
Success will likely depend less on whether AI matters—which increasingly appears inevitable—and more on whether enterprises conclude that existing architectures remain sufficient.
Viewed narrowly, groundcover’s Series C is another large infrastructure funding round.
Viewed more broadly, it reflects a growing debate about what observability becomes in an era where software increasingly writes, tests and operates itself.
If AI continues generating exponentially larger volumes of operational data, traditional assumptions about telemetry collection, pricing and storage may come under increasing pressure. Vendors that built businesses around charging for data ingestion may need to evolve their economics alongside customer expectations. New entrants, meanwhile, have an opportunity to design around those changing assumptions from the outset.
groundcover believes that opportunity lies in combining customer-controlled infrastructure, automatic telemetry collection and AI-assisted operations into a platform designed for autonomous software rather than simply adding AI features to existing observability products.
Whether that architectural bet proves durable will depend on enterprise adoption over the next several years.
But the company’s latest funding round suggests at least some investors believe the next battle in observability will not be fought over dashboards or alerts. It will be fought over who builds the operational data layer that increasingly intelligent software relies upon to understand—and eventually manage—the systems it runs.