Full-Stack Gains, Not Moore’s Law, Power AI in 2026
AMD and Cerebras, Moonshot AI and Nvidia rolled out new inference, model and security offerings in 2026, reflecting efficiency from combined hardware, model and software improvements.
AMD, Cerebras, Moonshot AI and Nvidia released new products and initiatives in 2026 that industry participants say change where AI performance and cost gains come from. The announcements cover inference systems, model design and operational security and were made at events and product launches this year, with activity in the U.S. and China.
At AMD’s Advancing AI 2026 event, the company said it will pair its Helios platform with Cerebras’ Wafer-Scale Engine in a single inference workflow. In that configuration, Helios handles prompt processing and high-volume throughput while Cerebras performs the generative compute. Company modeling projects up to five times more tokens per second per watt compared with a Cerebras-only setup.
Moonshot AI introduced Kimi K3, an open-weight model released in China that activates 16 of 896 specialist components per token. Moonshot estimates the design improves scaling efficiency about 2.5 times over its prior Kimi K2 model.
Nvidia launched the Open Secure AI Alliance, which brings together firms across cloud computing, cybersecurity, enterprise software and AI research to address identity, permissions, guardrails, logging and evaluation for agents. The alliance extends Nvidia’s prior coalition work at the model level to operational security and enterprise deployment.
Investment and academic groups have tracked rapid declines in the cost of running AI models. Venture firm Andreessen Horowitz reported that the cost of AI inference at a fixed performance level fell roughly tenfold per year between 2021 and 2024. Recent academic work places the recent frontier rate at five- to tenfold annual improvement. Analysts describe these changes as the result of multiple technical trends compounding: specialized chip designs, memory and networking improvements, serving software and changes in model architecture.
Industry participants are shifting measurement from cost per token to cost per completed, qualified task. Benchmarks that combine quality and cost capture whether a model completes a task correctly in one pass or needs multiple attempts or human review. One index that measures agentic work, coding, scientific reasoning and general knowledge reports Claude Opus 5 at a score of 61 at $2.03 per benchmark task, and DeepSeek V4 Flash at 44 at $0.04 per task.
Companies that provide end-to-end stacks reported partnerships and deployments tied to the new releases. Nebius Group, an AI cloud provider, cited a partnership with Nvidia on AI-factory design, inference software, hardware deployment and fleet management and was an early launch partner for Kimi K3. Vendors and service providers are positioning to support broader deployment as task-level costs decline.
Security and monitoring were a recurring theme in the announcements. Nvidia and other participants describe agent security as a system property that includes evaluation and observability to track identity, permissions and behavior at scale. The industry materials note that agent deployments will increase demands on logging, permissioning and continuous evaluation.
Market participants and indexes tracking AI firms say decision-making now requires monitoring multiple technology curves-chip architecture, model design, software stacks, data and security-and how they combine to lower the cost of useful work. Indexes such as the ROBO Global Artificial Intelligence Index include companies across those layers that are commercializing the combined advances.








