Aqara launches the Thermostat Hub W200, Camera Hub G350, U400 next-gen smart lock, and new spatial sensors at CES 2026
The OnePlus Pad Go 2 brings a 120Hz display, an upgraded battery, and smooth multitasking to a competitive Android tablet market. Can OnePlus challenge the Apple iPad?
The arms race to build smarter AI models has a measurement problem: the tests used to rank them are becoming obsolete almost as quickly as the models improve. On Monday, Artificial Analysis, an independent AI benchmarking organization whose rankings are closely watched by developers and enterprise buyers, released a major overhaul to its Intelligence Index that fundamentally changes how the industry measures AI progress.
The new Intelligence Index v4.0 incorporates 10 evaluations spanning agents, coding, scientific reasoning, and general knowledge. But the changes go far deeper than shuffling test names. The organization removed three staple benchmarks — MMLU-Pro, AIME 2025, and LiveCodeBench — that have long been cited by AI companies in their marketing materials. In their place, the new index introduces evaluations designed to measure whether AI systems can complete the kind of work that people actually get paid to do.
type: embedded-entry-inline id: 1bCmRrroGCdUb07IuaHysL
“This index shift reflects a broader transition: intelligence is being measured less by recall and more by economically useful action,” observed Aravind Sundar, a researcher who responded to the announcement on X (formerly Twitter).
The benchmark overhaul addresses a growing crisis in AI evaluation: the leading models have become so capable that traditional tests can no longer meaningfully differentiate between them. The new index deliberately makes the curve harder to climb. According to Artificial Analysis, top models now score 50 or below on the new v4.0 scale, compared to 73 on the previous version — a recalibration designed to restore headroom for future improvement.
This saturation problem has plagued the industry for months. When every frontier model scores in the 90th percentile on a given test, the test loses its usefulness as a decision-making tool for enterprises trying to choose which AI system to deploy. The new methodology attempts to solve this by weighting four categories equally — Agents, Coding, Scientific Reasoning, and Genera l— while introducing evaluations where even the most advanced systems still struggle.
The results under the new framework show OpenAI’s GPT-5.2 with extended reasoning effort claiming the top spot, followed closely by Anthropic’s Claude Opus 4.5 and Google’s Gemini 3 Pro. OpenAI describes GPT-5.2 as “the most capable model series yet for professional knowledge work,” while Anthropic’s Claude Opus 4.5 scores higher than GPT-5.2 on SWE-Bench Verified, a test set evaluating software coding abilities.
The most significant addition to the new index is GDPval-AA, an evaluation based on OpenAI’s GDPval dataset that tests AI models on real-world economically valuable tasks across 44 occupations and 9 major industries. Unlike traditional benchmarks that ask models to solve abstract math problems or answer multiple-choice trivia, GDPval-AA measures whether AI can produce the deliverables that professionals actually create: documents, slides, diagrams, spreadsheets, and multimedia content.
Models receive shell access and web browsing capabilities through what Artificial Analysis calls “Stirrup,” its reference agentic harness. Scores are derived from blind pairwise comparisons, with ELO ratings frozen at the time of evaluation to ensure index stability.
Under this framework, OpenAI’s GPT-5.2 with extended reasoning leads with an ELO score of 1442, while Anthropic’s Claude Opus 4.5 non-thinking variant follows at 1403. Claude Sonnet 4.5 trails at 1259.
On the original GDPval evaluation, GPT-5.2 beat or tied top industry professionals on 70.9% of well-specified tasks, according to OpenAI. The company claims GPT-5.2 “outperforms industry professionals at well-specified knowledge work tasks spanning 44 occupations,” with companies including Notion, Box, Shopify, Harvey, and Zoom observing “state-of-the-art long-horizon reasoning and tool-calling performance.”
The emphasis on economically measurable output is a philosophical shift in how the industry thinks about AI capability. Rather than asking whether a model can pass a bar exam or solve competition math problems — achievements that generate headlines but don’t necessarily translate to workplace productivity — the new benchmarks ask whether AI can actually do jobs.
While GDPval-AA measures practical productivity, another new evaluation called CritPT reveals just how far AI systems remain from true scientific reasoning. The benchmark tests language models on unpublished, research-level reasoning tasks across modern physics, including condensed matter, quantum physics, and astrophysics.
CritPT was developed by more than 50 active physics researchers from over 30 leading institutions. Its 71 composite research challenges simulate full-scale research projects at the entry level — comparable to the warm-up exercises a hands-on principal investigator might assign to junior graduate students. Every problem is hand-curated to produce a guess-resistant, machine-verifiable answer.
The results are sobering. Current state-of-the-art models remain far from reliably solving full research-scale challenges. GPT-5.2 with extended reasoning leads the CritPT leaderboard with a score of just 11.5%, followed by Google’s Gemini 3 Pro Preview and Anthropic’s Claude 4.5 Opus Thinking variant. These scores suggest that despite remarkable progress on consumer-facing tasks, AI systems still struggle with the kind of deep reasoning required for scientific discovery.
Perhaps the most revealing new evaluation is AA-Omniscience, which measures factual recall and hallucination across 6,000 questions covering 42 economically relevant topics within six domains: Business, Health, Law, Software Engineering, Humanities & Social Sciences, and Science/Engineering/Mathematics.
The evaluation produces an Omniscience Index that rewards precise knowledge while penalizing hallucinated responses — providing insight into whether a model can distinguish what it knows from what it doesn’t. The findings expose an uncomfortable truth: high accuracy does not guarantee low hallucination. Models with the highest accuracy often fail to lead on the Omniscience Index because they tend to guess rather than abstain when uncertain.
Google’s Gemini 3 Pro Preview leads the Omniscience Index with a score of 13, followed by Claude Opus 4.5 Thinking and Gemini 3 Flash Reasoning, both at 10. However, the breakdown between accuracy and hallucination rates reveals a more complex picture.
On raw accuracy, Google’s two models lead with scores of 54% and 51% respectively, followed by Claude 4.5 Opus Thinking at 43%. But Google’s models also demonstrate higher hallucination rates than peer models, scoring 88% and 85%. Anthropic’s Claude 4.5 Sonnet Thinking and Claude Opus 4.5 Thinking show hallucination rates of 48% and 58% respectively, while GPT-5.1 with high reasoning effort achieves 51%—the second-lowest hallucination rate tested.
Both Omniscience Accuracy and Hallucination Rate contribute 6.25% weighting each to the overall Intelligence Index v4.
The benchmark reshuffling arrives at an especially turbulent moment in the AI industry. All three leading frontier model developers have launched major new models within just a few weeks — and Gemini 3 still holds the top spot on much of the leaderboards on LMArena, a widely cited benchmarking tool used to compare LLMs.
Google’s November release of Gemini 3 prompted OpenAI to declare a “code red” effort to improve ChatGPT. OpenAI is counting on its GPT family of models to justify its $500 billion valuation and over $1.4 trillion in planned spending. “We announced this code red to really signal to the company that we want to marshal resources in one particular area,” said Fidji Simo, CEO of applications at OpenAI. Altman told CNBC he expected OpenAI to exit its code red by January.
Anthropic responded with Claude Opus 4.5 on November 24, achieving an SWE-Bench Verified accuracy score of 80.9% — reclaiming the coding crown from both GPT-5.1-Codex-Max and Gemini 3. The launch marked Anthropic’s third major model release in two months. Microsoft and Nvidia have since announced multi-billion-dollar investments in Anthropic, boosting its valuation to about $350 billion.
Artificial Analysis emphasizes that all evaluations are run independently using a standardized methodology. The organization states that its “methodology emphasizes fairness and real-world applicability,” estimating a 95% confidence interval for the Intelligence Index of less than ±1% based on experiments with more than 10 repeats on certain models.
The organization’s published methodology defines key terms that enterprise buyers should understand. According to the methodology documentation, Artificial Analysis considers an “endpoint” to be a hosted instance of a model accessible via an API — meaning a single model may have multiple endpoints across different providers. A “provider” is a company that hosts and provides access to one or more model endpoints or systems. Critically, Artificial Analysis distinguishes between “open weights” models, whose weights have been released publicly, and truly open-source models—noting that many open LLMs have been released with licenses that do not meet the full definition of open-source software.
The methodology also clarifies how the organization standardizes token measurement: it uses OpenAI tokens as measured with OpenAI’s tiktoken package as a standard unit across all providers to enable fair comparisons.
For technical decision-makers evaluating AI systems, the Intelligence Index v4.0 provides a more nuanced picture of capability than previous benchmark compilations. The equal weighting across agents, coding, scientific reasoning, and general knowledge means that enterprises with specific use cases may want to examine category-specific scores rather than relying solely on the aggregate index.
The introduction of hallucination measurement as a distinct, weighted factor addresses one of the most persistent concerns in enterprise AI adoption. A model that appears highly accurate but frequently hallucinates when uncertain poses significant risks in regulated industries like healthcare, finance, and law.
The Artificial Analysis Intelligence Index is described as “a text-only, English language evaluation suite.” The organization benchmarks models for image inputs, speech inputs, and multilingual performance separately.
The response to the announcement has been largely positive. “It is great to see the index evolving to reduce saturation and focus more on agentic performance,” wrote one commenter in an X.com post. “Including real-world tasks like GDPval-AA makes the scores much more relevant for practical use.”
Others struck a more ambitious note. “The new wave of models that is just about to come will leave them all behind,” predicted one observer. “By the end of the year the singularity will be undeniable.”
But whether that prediction proves prophetic or premature, one thing is already clear: the era of judging AI by how well it answers test questions is ending. The new standard is simpler and far more consequential — can it do the work?
The new Cambridge L/R Series of active speakers consist of three models: L/R X, L/R M and L/R S. Each stereo speaker pair has been designed to reveal depth and nuance.
Ugreen has announced a slew of new products at CES 2026, including intelligent AI NAS storage, a proactive home security and an advanced 300W charging station,
The Eve Thermostat has been announced at CES 2026. Controls, schedules and automations can all run locally, with no subscriptions or cloud dependency required.
Multiple elements of the streamer’s 2026 output set to enjoy unprecedented picture and sound quality.
While most brands are still getting to grips with simple RGB implementations of these next-gen screen technologies, Hisense is already adding a fourth color dimension.
The Samsung Galaxy Z Fold 8 “Wide” model is coming sooner than leaked — likely landing before Apple’s iPhone Fold.
For the last two years, the prevailing logic in generative AI has been one of brute force: if you want better reasoning, you need a bigger model.
While “small” models (under 10 billion parameters) have become capable conversationalists, they have historically crumbled when asked to perform multi-step logical deduction or complex mathematical proofs.
Today, the Technology Innovation Institute (TII) in Abu Dhabi is challenging that scaling law with the release of Falcon H1R 7B.
By abandoning the pure Transformer orthodoxy in favor of a hybrid architecture, TII claims to have built a 7-billion parameter model that not only rivals but outperforms competitors nearly 7X its size — including the 32B and 47B variants of Alibaba’s Qwen and Nvidia’s Nemotron.
The release marks a significant shift in the open-weight ecosystem, moving the battleground from raw parameter count to architectural efficiency and inference-time scaling.
The full model code is available now at Hugging Face and can be tested by individuals in a live demo inference on Falcon Chat (a chatbot experience). TII further released a seemingly quite comprehensive technical report on the approach and training methodology for Falcon H1 7B, as well.
The defining feature of Falcon H1R 7B is its “hybrid” backbone. Most modern LLMs rely exclusively on the Transformer architecture, which scales predictably but suffers from high memory costs when processing long sequences.
Falcon H1R 7B integrates Mamba, a state-space model (SSM) architecture, alongside standard Transformer attention layers.
Originally developed by researchers Albert Gu and Tri Dao at Carnegie Mellon University and Princeton University, Mamba was first introduced in the paper “Mamba: Linear-Time Sequence Modeling with Selective State Spaces” published on December 1, 2023.
The architecture processes data sequences differently than Transformers: while Transformers compare every piece of data to every other piece (quadratic scaling), Mamba processes tokens sequentially, allowing it to handle vast amounts of information with linear scaling and significantly reduced compute costs.
This combination addresses one of the most persistent bottlenecks in deploying reasoning models: the cost of “thinking.” Reasoning models require generating long “chains of thought”—step-by-step internal monologues—before arriving at an answer. For standard Transformers, these long contexts explode computational costs.
According to TII’s technical report, the hybrid approach allows Falcon H1R 7B to maintain high throughput even as response lengths grow. At a batch size of 64, the model processes approximately 1,500 tokens per second per GPU—nearly double the speed of the competing Qwen3 8B model.
In the benchmarks released by TII, the disparity between Falcon H1R 7B’s size and its performance is stark. On the AIME 2025 leaderboard—a rigorous test of mathematical reasoning—Falcon H1R 7B scored 83.1%, a result that disrupts the traditional hierarchy of model sizing.
While the 7B model naturally trails massive proprietary frontiers like GPT-5.2 (99.0%) and Gemini 3 Flash (97.0%) on the separate Artificial Analysis index (run by the independent organization of the same name, which has not yet benchmarked Falcon H1R 7B yet), it has effectively collapsed the gap between “efficient” open weights and mid-tier proprietary systems.
Beating Larger “Thinkers”: Falcon H1R 7B (83.1%) outperforms the 15-billion parameter Apriel-v1.6-Thinker (82.7%) and the 32-billion parameter OLMo 3 Think (73.7%), validating TII’s claim that hybrid architectures can out-reason larger Transformers.
Chasing Proprietary Leaders: It sits within striking distance of Claude 4.5 Sonnet (88.0%) and Amazon Nova 2.0 Lite (88.7%), suggesting that for specific math-heavy workflows, this 7B model is a viable, low-latency alternative to expensive commercial APIs.
Outperforming Legacy Giants: On this specific reasoning metric, it decisively beats broadly capable but older architectures like Mistral Large 3 (38.0%) and Llama 4 Maverick (19.3%), highlighting how specialized reasoning training (“Deep Think”) has become more critical than raw scale for logic tasks.
Other key domain wins include:
Coding: The model achieved 68.6% on the LCB v6 benchmark, a score TII claims is the highest among all tested models, including those four times its size.
General Reasoning: While it dominates in math and code, its general reasoning score (49.48%) remains competitive, sitting just below the 14B and 15B parameter models but comfortably ahead of comparable 8B models.
Falcon H1R 7B’s performance is not just architectural; it stems from a rigorous, two-stage training pipeline designed to maximize reasoning density without inflating parameter count, according to TII’s technical report on the model.
Stage 1: Cold-Start Supervised Fine-Tuning (SFT). The model underwent “cold-start” SFT on a curated dataset dominated by mathematics (56.8% of tokens) and code (29.8%), with response lengths stretching up to 48,000 tokens.
Difficulty-Aware Weighting: TII rejected the standard practice of treating all data equally. Instead, they applied a weighting scheme where “hard” problems were up-weighted by 1.25x to 1.75x, while easy problems were down-weighted or removed entirely to prevent overfitting to trivial tasks.
Single-Teacher Consistency: Ablation studies revealed that mixing reasoning traces from multiple “teacher” models actually degraded performance due to conflicting reasoning styles. Consequently, TII opted for a single-teacher approach to maintain coherent internal logic.
Balanced Token Normalization: To handle the massive variance in sequence lengths (short instructions vs. massive reasoning chains), the team introduced a Balanced Data-Parallel Token Normalization strategy. This technique equalizes the gradient contribution of each token across GPUs, preventing ranks with shorter sequences from destabilizing the loss—a change that yielded a consistent 4-10% accuracy boost during training.
Stage 2: Reinforcement Learning via Group Relative Policy Optimization (GRPO). Following SFT, the model was refined using GRPO a reinforcement learning algorithm that rewards correct outcomes without needing a separate value model.
The “No-KL” Shift: In a deviation from standard RLHF, TII removed the KL-divergence penalty (beta=0) entirely. This allowed the model to drift significantly from its base SFT policy, encouraging aggressive exploration of novel reasoning paths.
Math-Only Curriculum: Surprisingly, TII found that training exclusively on math problems during the RL stage yielded better generalization across all domains—including code and science—than mixed strategies. Ablations showed that “code-only” training improved coding scores but harmed general reasoning, whereas math-focused RL lifted performance globally.
TII optimized the model specifically for Test-Time Scaling (TTS), a technique where a model generates multiple reasoning paths in parallel to find the best solution.
The model utilizes Deep Think with Confidence (DeepConf), which leverages the model’s internal confidence scores to dynamically prune low-quality reasoning traces.
Adaptive Pruning: During generation, the system initiates a “warm-up” phase with 16 traces to establish a confidence baseline. It then aggressively filters subsequent traces, terminating any chain that falls below the 10th percentile of the baseline confidence.
Efficiency Gains: This method creates a new Pareto frontier for deployment. In benchmark tests, Falcon H1R 7B achieved 96.7% accuracy on AIME 25 while reducing token usage by 38% compared to the DeepSeek-R1-0528-Qwen3-8B baseline.
TII has released Falcon H1R 7B under the custom Falcon LLM License 1.0 based on Apache 2.0 — but with notable modifications — chiefly among them: not to litigate against TII, and also to always credit it.
For developers and startups, the license is largely permissive:
Royalty-Free: Users can run, modify, and distribute the model commercially without paying TII.
Attribution: Any derivative work (including fine-tunes) must prominently state: “[Name of work] is built using Falcon LLM technology from the Technology Innovation Institute”.
However, unlike a pure Open Source Initiative (OSI) license, the Falcon license includes a strict Acceptable Use Policy (AUP).
The license terminates automatically if the model is used to create work that conflicts with the AUP or if the user initiates patent litigation against TII.
Specifically, the AUP prohibits using Falcon H1R 7B or its derivatives for:
Violating Laws: Any use that violates applicable national, federal, state, local, or international laws or regulations.
Harm to Minors or Living Beings: Exploiting, harming, or attempting to exploit or harm minors or any living beings.
Disinformation: Generating or disseminating verifiably false information with the purpose of harming others.
Harassment: Defaming, disparaging, or otherwise harassing others.
TII is not alone in betting on this hybrid future; the industry is increasingly moving toward architectures that blend the strengths of SSMs and Transformers.
Nvidia recently debuted the Nemotron 3 family on December 15, 2025, which utilizes a hybrid mixture-of-experts (MoE) and Mamba-Transformer design to drive efficient agentic AI.
IBM launched its Granite 4.0 family on October 2, 2025, using a hybrid Mamba-Transformer architecture to cut memory requirements by over 70% while maintaining high performance on enterprise benchmarks.
AI21 has pursued this path with its Jamba (Joint Attention and Mamba) models, releasing the Jamba 1.5 family on August 22, 2024, to boost agentic AI capabilities through a hybrid SSM-Transformer approach.
Mistral entered the space early with Codestral Mamba on July 16, 2024, a model specifically optimized for faster, longer code generation.
Falcon H1R 7B represents the latest evolution in this trend, specifically targeting dense reasoning tasks in a compact form factor.