If the keyword for AI infrastructure over the past few years was simply more compute, the large-scale deployment of Agentic AI is forcing the industry to confront a more practical question: How efficiently can all that compute actually be used?
At the 2026 Apsara Conference, H3C put the question in terms of Tokens. As AI agents continuously call on foundation models, trillion-Token-scale services are emerging as a new infrastructure requirement.
Simply adding more GPUs does not necessarily deliver proportional performance gains. Idle compute, network congestion, insufficient data supply, and the operational and failure challenges of large-scale clusters can all erode the value of additional compute.
That is why H3Cโs product lineup at the event, while spanning a wide range of technologies, follows a relatively clear logic: from compute to networking and storage, and from software to operations, AI infrastructure is moving beyond hardware stacking toward system-level optimization.

More GPUs do not necessarily mean more performance
Zhu Shiyin, general manager of H3Cโs Advanced Technology Research Department, described AI clusters as highly integrated systems that require close hardware-software coordination.
As cluster sizes grow from hundreds of GPUs to thousands or even tens of thousands, GPUs themselves become just one part of the equation. Chip-to-chip communication, data transfer, task scheduling, power supply, and cooling can all directly affect overall compute utilization.
H3Cโs UniPoD S80000 Series SuperPod, showcased at the event, reflects this approach. It supports configurations ranging from 32 to 1,024 GPUs and can scale further to 16,384 GPUs, while supporting heterogeneous computing resources including CPUs, GPUs, NPUs, and DPUs and mainstream frameworks.
The focus is not simply on putting more chips into a single system, but on coordinating compute, networking, storage, cloud, security, and operations so that different types of computing resources can be managed as a unified system.
In other words, the competition in AI clusters is shifting from how many GPUs a system has to how many useful Tokens each GPU can actually produce.

As GPUs get faster, the network can become the bottleneck
Zhu also highlighted high-speed interconnects, an increasingly important part of AI infrastructure. The reason is straightforward: as GPU performance improves, the amount of data exchanged between GPUs also grows. If the network cannot keep up, GPUs are forced to wait.
This is why H3C showcased three interconnect scenarios at the event: Scale-Up, Scale-Out, and Scale-Across.
Within a node, H3C introduced the S9828-128EO, a 102.4T NPO silicon-photonics intelligent computing switch that uses co-packaged optics to reduce end-to-end latencyย by 15%. Between nodes, the companyโs 1.6T intelligent computing switch S9828-64FP, adoptsย 224G SerDes technologies, it cuts power consumption by more than 20% in high-density environments to supportย large-scale cluster expansion.
Across data centers, the 800G intelligent computing DCI switch S12500R-64EP is designed to support intra-city cross-data-center compute scheduling and gradient synchronization.
The logic behind these three layers comes down to a core question: As AI clusters grow, how can interconnects avoid becoming the brakes on compute performance?
Interconnects are also no longer simply a matter of bandwidth. As clusters expand, protocols, congestion control, link reliability, and fault diagnosis become equally important. If vendors develop closed technical ecosystems independently, they could create new technology silos. Open protocols and industry standards will therefore become increasingly important.

Beyond compute, there is another frequently underestimated challenge: data
GPUs cannot run without data. In large-model training, data preparation, model training, and parameter exchange all involve massive amounts of data movement. During inference, KV Cache adds further pressure on storage and caching resources.
H3C therefore also positioned high-performance storage as part of the broader Token production pipeline. Its UniStor X20000 series X20836 can deliver up to 200GB/s of bandwidth and 3 million IOPS per node, while supporting interoperability across block, file, object, and HDFS protocols.
According to H3C, its full-speed engine can reduce GPU waiting time by 30%, while XCache inference acceleration can cut time-to-first-token latency by up to 90% in inference scenarios.
The takeaway is fairly practical: AI infrastructure is evolving from a compute center into a comprehensive data processing system. From data preparation and training to inference and Token generation, every period of waiting ultimately translates into cost.

Software may be the next battleground
Once hardware reaches a certain scale, software optimization becomes increasingly important. Zhu noted that everything from the operating system and communication libraries to resource management and scheduling platforms needs to work closely with the underlying hardware.
Communication optimization, overlapping computation with communication, task scheduling, dynamic load balancing, and unified resource management can all help reduce GPU idle time and communication conflicts.
This is also why modern AI infrastructure increasingly resembles a super-system. Chips provide compute, networks connect resources, storage supplies data, and power and cooling keep the system running. Software, meanwhile, is responsible for orchestrating these resources into a functioning whole.
If any one of these components becomes a bottleneck, the impact is no longer limited to a single hardware metric. It ultimately affects how many useful Tokens the system can generate per second and how much each Token costs to produce.
The next phase of AI infrastructure is a competition over compute utilization
From H3Cโs exhibition booth to its forum sessions at the Apsara Conference, the company was essentially making the same point from different angles.
Supernodes address how compute resources can be organized at scale. High-speed networking addresses how those resources can communicate efficiently. High-performance storage ensures that data can be delivered in time. Software, scheduling, operations, and unified management then turn these resources into AI capabilities that can be used continuously and efficiently.
The next phase of competition in AI infrastructure, therefore, may not be determined simply by who has more GPUs. As model sizes reach the trillion-parameter scale and AI agents increasingly call foundation models, metrics such as compute utilization, Token throughput, latency, energy consumption, and cost per Token are becoming increasingly important.
That is the deeper meaning behind the pursuit of extreme Token cost efficiency: it is not about putting more GPUs together, but about making every GPU, every network link, and every unit of storage in the system spend less time waiting and more time working.
As AI enters the stage of large-scale production, compute is no longer just a chip problem. It is becoming a competition over the efficiency of the entire infrastructure system.
