✦ Free GPU infrastructure assessment — get a complimentary RA:X assessment

Get started
Industry Insights

AI Factory Era: Why NVIDIA Is Betting on Operating Software Over GPU Performance

July 27, 2026

AI Factory Era: Why NVIDIA Is Betting on Operating Software Over GPU Performance

AI Factory Era: Why NVIDIA Is Betting on Operating Software Over GPU Performance

"NVIDIA used to sell GPUs. Now it sells the software that runs them."

In the first half of 2026, NVIDIA delivered a consistent message at GTC and Computex. It wasn't a faster GPU. It was DSX — a software platform that defines how GPUs are designed, deployed, and operated.

As Jensen Huang put it: "Intelligence tokens are the new currency of the AI era, and AI factories are the infrastructure that produces that currency."

The message: GPU clusters are becoming factories that produce tokens, and what determines that factory's efficiency isn't hardware — it's operating software. What does this shift actually mean for the rest of us?

What Is an AI Factory?

Diagram showing GPU clusters evolving into AI factories with an integrated operations software layer

An AI factory treats a GPU cluster not as a pool of compute, but as an industrial facility that manufactures tokens.

Just as a factory feeds in raw materials to produce goods, an AI factory feeds in power and data to produce intelligence — tokens. And just as a factory's productivity is decided by the efficiency of the whole production line rather than any single machine's specs, an AI factory's competitiveness is decided by the efficiency of the entire system rather than any single GPU's specs.

From this view, GPUs are the machines on the factory floor, and operating software is the MES (Manufacturing Execution System) that runs the factory.

NVIDIA DSX: From Chip Company to AI Factory Orchestrator

Through the DSX platform, NVIDIA is trying to define the design, simulation, and operation of the entire AI factory — not just the GPU chip.

DSX isn't a single product. It's a set of software modules:

DSX MaxLPS — Maximizes token performance within a fixed power budget. NVIDIA says it enables running up to 40% more GPUs on the same power budget, combining 45°C liquid cooling with in-rack optimization to run GPUs at their most energy-efficient point.

DSX OS — Modular software for running AI factory operations: lifecycle management, intelligent scheduling, health automation, resilience, and multi-tenant operations. It's open source, so partners can freely build on it.

DSX Flex — Connects the AI factory to grid services. It dynamically adjusts power demand and works with hybrid on-site generation to cut energy costs while keeping the grid stable.

DSX Sim — Simulates the entire AI factory as a digital twin before it's ever built. As Jensen Huang said, you can "simulate the whole factory before spending a single dollar."

Here's the core point: what NVIDIA is building isn't a faster GPU. It's a software platform that runs the entire factory built out of GPUs.

Cloud partners like CoreWeave, Lambda, and Nebius are already deploying DSX components, and Dell, HPE, Lenovo, and Supermicro are building DSX-compatible systems. The entire GPU industry is shifting its center of gravity from "chip performance" to "factory operations."

Tokens per Watt: A New Efficiency Metric Takes Over

Stats showing 30-50% capacity delays, 40% power loss, and 40% more GPUs with DSX MaxLPS

In the AI factory era, the metric that matters most isn't FLOPS. It's Tokens per Watt.

Why has this metric become so important? Because power has become the biggest bottleneck to AI scaling.

Industry analysis suggests that 30–50% of data center capacity planned for 2026 could be pushed back to 2028 or later due to power shortages. Hyperscaler AI capex is projected to exceed $750 billion in 2026 and $1 trillion in 2027 — but grid expansion isn't keeping pace.

In a typical AI factory, only about 60% of the power drawn from the grid actually goes toward AI compute. The remaining 40% is lost to cooling, power distribution losses, and rack-level inefficiency.

NVIDIA's DSX MaxLPS treats that 40% as recoverable compute capacity — the goal is to run 40% more GPUs on the same power budget while minimizing the impact on workload performance.

And this efficiency gap shows up dramatically across hardware generations, too. GB300 NVL72's Tokens per Watt is up to 25x that of Hopper on DeepSeek V4 Pro. Even more striking: on the same hardware, software optimization alone delivered a 5x improvement in a single month.

If you can get a 5x improvement through software alone — without changing the GPU — the real competitive edge isn't hardware. It's operating software.

What This Shift Means for Enterprise GPU Infrastructure

NVIDIA DSX is built for hyperscale AI factories running thousands to tens of thousands of GPUs. But the core of NVIDIA's message applies at every scale of GPU infrastructure.

"GPU infrastructure competitiveness is decided by operating software, not GPU performance."

Applied to the enterprise, this message points to four challenges:

First, utilization matters more than acquisition. GPU prices have doubled and lead times run a year. When buying more isn't an option and average enterprise utilization sits at just 5%, raising operational efficiency has effectively the same impact as acquiring more GPUs.

Second, you need a structure for multiple teams to share GPUs without conflict. Multi-tenant operation is one of the core capabilities of NVIDIA DSX OS. In the enterprise, letting AI, data, and DevOps teams safely share the same cluster requires role-based access control and resource isolation.

Third, power efficiency isn't optional. If Tokens per Watt determines an AI factory's revenue, then in enterprise GPU infrastructure, power optimization decides not just cost savings but how many AI projects you can actually run.

Fourth, you need to measure GPU usage and attribute cost. Just as AI factories measure revenue per token, enterprises need to track GPU usage by team and project and demonstrate cost efficiency to keep securing investment.

An Operating Layer for the Enterprise AI Factory

Comparison of NVIDIA DSX for hyperscale and AIPub for enterprise GPU operations layers

If NVIDIA DSX is the AI factory operating platform for hyperscale, what operating layer does an enterprise GPU cluster need?

The scale is different, but the challenges are the same: resource partitioning, multi-tenant isolation, scheduling, power control, monitoring, and cost attribution. Bringing these six together on a single platform is exactly where AIPub has focused.

AIPub partitions a single GPU into up to 100 blocks to raise utilization, unifies isolation of GPUs, images, and storage by team, project, and role, automatically detects and reallocates idle resources, controls power independently per card, tracks the entire stack in real time across 40+ proprietary metrics, and automatically generates block-level, usage-based chargeback.

Now that NVIDIA has declared "the software that runs GPUs is the core of the AI factory," this operating layer is no longer a nice-to-have. It's the infrastructure that determines the ROI of your GPU investment.

AI Factory Readiness Checklist

Check whether your organization's GPU infrastructure is ready for the AI factory era:

Are you measuring GPU utilization in real time and automatically reallocating idle resources?

Do you have an isolation structure that lets multiple teams share one GPU cluster without conflict?

Are you controlling GPU power per card based on workload characteristics?

Can you track GPU usage by team and project and attribute cost?

Do you have a software layer that controls your entire GPU infrastructure?

Before buying more GPUs, run the GPUs you have at AI-factory-level efficiency.

You don't have to figure this out alone. In the AI factory era — now validated by NVIDIA itself

— let's design your GPU infrastructure operating strategy together with TEN's experts.

👉See AIPub in detail

👉 Learn more about RA:X infrastructure assessment

📩 Talk to an expert

Reference

  • NVIDIA, "DSX Gives Infrastructure Builders the Playbook for AI Factories" (May 2026)
  • NVIDIA, "Vera Rubin DSX AI Factory Reference Design" (March 2026)
  • NVIDIA Technical Blog, "Scaling Token Factory Revenue by Maximizing Performance per Watt" (April 2026)
  • NVIDIA Technical Blog, "DSX OS Delivers Open, Modular Software for Operating AI Factories at Scale" (June 2026)
  • TechTimes, "Tokens per Watt Determines AI Factory Revenue as Power Constraints Tighten" (July 2026)
  • Help Net Security / NTT Data, "Power shortages could slow AI data center expansion" (July 2026)

Related reading