NVIDIA Unveils Next-Generation AI Factory Infrastructure to Drive Agentic Efficiency at AI Infra Summit

The Santa Clara Convention Center transformed into the global epicenter of computing infrastructure this week as the AI Infra Summit convened a record-breaking 8,000 attendees, more than double the attendance of the previous year. Amidst the rapid maturation of agentic AI—a class of autonomous workloads characterized by complex reasoning chains and iterative tool calls—NVIDIA vice president of hyperscale and high-performance computing, Ian Buck, took the stage to redefine the operational metrics of the modern AI factory. The central theme of the summit was clear: as energy constraints tighten, the industry must pivot from measuring raw peak performance to quantifying "validated agentic tokens per megawatt."
This shift underscores a critical evolution in how data centers are designed. Modern AI factories are no longer simple collections of servers; they are intricate, co-designed systems spanning silicon, networking, software, and grid-level power management. To address this, NVIDIA showcased a comprehensive stack, including its Vera Rubin systems, Dynamo inference software, NeMo libraries, and a robust networking suite anchored by NVLink for high-speed scale-up, Spectrum-X Ethernet, and BlueField DPUs for infrastructure security and context-memory management.

The New Frontier of Power Optimization
The imperative for energy efficiency stems from the physical limits of power delivery. With power-constrained deployments becoming the norm, the economic viability of AI operations depends heavily on maximizing throughput within a fixed energy budget. NVIDIA introduced the DSX MaxLPS (Load Power Smoothing) platform as a foundational solution to this bottleneck. By continuously monitoring power consumption across entire GPU clusters and racks, the software dynamically reallocates energy in real-time, reclaiming capacity that would otherwise remain stranded under traditional static provisioning.
The tangible benefits of this approach were validated by AI cloud provider Lambda. In a benchmark study presented at the summit, Lambda demonstrated that by integrating DSX MaxLPS, they could operate 19 nodes within the same power envelope typically reserved for 16 nodes. This transition resulted in a 24% increase in cluster-wide token throughput—climbing from 4 million to 5 million tokens per second—while simultaneously improving performance per watt by 23%. Projections for the upcoming Vera Rubin NVL72 systems suggest that this efficiency gain could translate into a 40% increase in GPU capacity within the same megawatt footprint, fundamentally altering the return-on-investment calculations for massive AI infrastructure projects.
Grid-Responsive AI: The Emerald AI Partnership
Beyond internal rack optimization, the summit highlighted the role of AI factories as active participants in the electrical grid. As data centers consume increasing shares of regional power, the ability to engage in demand-response programs is becoming a competitive necessity. Through a strategic collaboration with Silicon Valley Power, Emerald AI demonstrated the efficacy of the NVIDIA DSX Flex platform.

This software allows AI factories to act as flexible grid resources, automatically responding to real-time electricity grid signals. During the demonstration, the system successfully navigated hundreds of demand-response events. When the grid faced stress, the software autonomously throttled power to low-priority AI tasks, preserving electricity for critical operations and maintaining performance integrity. This ability to "load-shed" effectively transforms the AI factory from a static consumer into a dynamic asset, potentially unlocking new grid capacity for infrastructure expansion in otherwise constrained regions.
Performance Benchmarks and the Rise of AgentX
The emergence of agentic AI has rendered traditional request-response benchmarks insufficient. Unlike standard chat applications, agentic workflows involve "sub-agent spawning," extensive context accumulation, and long-tail tool calls. To measure this accurately, the industry is increasingly turning to the SemiAnalysis AgentX framework, which simulates real-world coding and reasoning sessions rather than isolated queries.
Data presented on the AgentX dashboard confirms that the Vera Rubin NVL72 architecture provides a significant leap in productivity. On the DeepSeek V4 Pro model, the platform delivered up to 30x higher throughput per megawatt compared to the previous-generation GB300 NVL72. Furthermore, the cost-per-million-tokens is estimated to be up to 45x lower, a metric that will likely dictate the profit margins for hyperscale AI providers. This performance is attributed to the deep co-design of the system, utilizing sixth-generation NVLink interconnects and the NVFP4 precision supported by fifth-generation Tensor Cores.

The Vera CPU Ecosystem and Developer Adoption
While the GPU remains the primary engine for AI training and inference, the NVIDIA Vera CPU has emerged as a critical component for the orchestration and management of agentic workflows. A broad range of startups and enterprise players shared benchmark data at the summit, highlighting the CPU’s ability to handle data-intensive tasks.
Perplexity reported that the Vera CPU accelerated its "SPACE" secure sandbox platform by 1.9x, while Redpanda noted a 5.5x improvement in latency and a 73% increase in throughput compared to legacy CPU architectures. Other notable results included Starburst’s 3x improvement in query throughput and DeepInfra’s observation of 2.2x faster orchestration latency. These gains are crucial for agentic AI, where the speed of the "control plane"—the logic that decides which tools to call and how to organize data—often becomes the primary bottleneck in system latency.
Resiliency at Scale: NVLink 6
As clusters scale to include hundreds of thousands of GPUs, the probability of hardware faults, signal degradation, and transient errors increases proportionally. The summit also served as the debut for the architectural details of NVLink 6, which introduces a multi-layered resiliency framework.

To ensure continuous uptime, the system employs custom forward error correction and universal physical layer recovery at the hardware level. At the network level, credit-based flow control and dynamic routing work in tandem to isolate faults and prevent the "cascading stalls" that often plague large-scale distributed systems. By detecting and containing errors locally, NVLink 6 ensures that the high-performance interconnect fabric remains lossless, maintaining the throughput required for training models with trillions of parameters.
Implications for the Future of AI Infrastructure
The developments discussed at the AI Infra Summit mark the end of the "brute force" era of AI infrastructure. As the industry moves toward 2026 and beyond, the focus is clearly pivoting toward high-density, power-aware, and highly resilient systems. The ability to generate more tokens per megawatt is no longer just a technical achievement; it is the fundamental driver of the AI economy.
The integration of deterministic, ultra-low-latency inference via the Groq 3 LPX, combined with the power-management capabilities of DSX MaxLPS, provides a roadmap for sustainable growth. For enterprises, the takeaway is clear: the next generation of AI will not be built on raw power alone, but on the intelligent orchestration of energy, compute, and networking. As NVIDIA and its partners continue to deploy these technologies, the benchmark for "success" in AI will be measured by how effectively these factories can convert limited electrical resources into intelligence, reasoning, and actionable output at an unprecedented scale.







