Source: H3C |Original Link


With the exponential surge in global token invocation volumes, many enterprises face significant challenges: indiscriminate GPU accumulation leads to idle computing power, network congestion, and frequent service disruptions caused by system failures. Consequently, the unit cost per token remains persistently high, trapping organizations in a vicious cycle where greater scale results in higher costs.

During WAIC 2026, H3C, a subsidiary of Unisplendor Corporation, prominently showcased its Lingxi Intelligent Computing Solution. Abandoning fragmented delivery models, this solution leverages full-stack synergy across computing, networking, storage, cloud, security, and operations to cover the entire token lifecycle—from infrastructure construction and operation to governance and optimization—establishing a comprehensive cost-reduction closed loop. Through deep collaborative optimization of full-stack software and hardware, the solution has achieved system-level results, reducing total cost of ownership (TCO) across the entire token generation chain by over 30% and energy costs by over 40%. This enables enterprises to enter a virtuous growth curve where expanded token output significantly dilutes unit costs.

Four Core Capabilities

1. Construction: Elastic Rapid Delivery with Lightweight Flexible Deployment

The Lingxi Intelligent Computing Solution adopts an integrated delivery model that effectively avoids prolonged deployment cycles caused by multi-vendor hardware compatibility issues and cross-system integration. It supports standardized batch automated provisioning and unified onboarding of software and hardware platforms. Architecturally, it enables elastic layer-by-layer scaling, expanding seamlessly from single clusters to cross-domain deployments of tens of thousands or even hundreds of thousands of cards. Simultaneously, it accommodates three deployment modes: centralized intelligent computing, edge computing, and hybrid cloud. Existing heterogeneous GPUs can be uniformly integrated and reused without requiring one-time replacement of all hardware. For small-to-medium projects, AI services can go live within days, significantly compressing infrastructure construction cycles for token computing foundations and enabling clients to commence token production more rapidly.

2. Operation: Highly Stable Self-Healing Foundation Ensuring Uninterrupted Token Output

Training interruptions and inference downtime directly invalidate generated tokens, while task reruns incur additional computing resource consumption and time costs. The Lingxi Intelligent Computing Solution features a full-link autonomous operation system capable of detecting hardware anomalies within three seconds and completing isolation and self-healing within minutes. Network flash interruptions result only in minor speed reductions rather than complete service disruption. A dual-plane lossless network achieves millisecond-level link switching with near-zero packet loss in long-distance cross-domain clusters. Integrated with the Lingxi O&M Intelligent Agent, the system proactively conducts fault inspection and localization, maintaining effective training uptime at 99% and ensuring stable, round-the-clock token output for model iteration and high-frequency agent invocations.

3. Governance: Unified Full-Link Management for Secure and Compliant Token Control

The Lingxi Intelligent Computing Solution provides multi-tiered, granular governance across computing resources, data, models, and agents, addressing pain points related to fragmented management and elusive security risks. A unified resource backend supports tenant-level isolation and visualized scheduling of computing tasks, with token usage tracked at API key granularity to facilitate internal and external computing resource operations for government and enterprise clients. At the underlying level, micro-VM sandboxes enforce tenant isolation, while an AI gateway performs content auditing and restricts token invocation frequency per API key. These measures effectively mitigate risks such as prompt injection and malicious computing resource scraping, ensuring end-to-end traceability, security, and compliance throughout the token generation and delivery process.

4. Optimization: Full-Stack Software-Hardware Synergy to Continuously Reduce Token Production TCO

Through deep collaboration between hardware engineering and software optimization, the Lingxi Intelligent Computing Solution precisely converts every unit of computing power into token output.

Hardware Engineering System Capabilities:

  • Liquid Cooling:High-power components utilize liquid cooling, with system-level liquid cooling coverage reaching up to 80%. The system supports innovative liquid cooling technologies such as two-phase cold plates and full liquid cooling, reducing the Power Usage Effectiveness (PUE) to below 1.04 and significantly lowering energy consumption costs during sustained large-scale operations;
  • High-Density Architecture:The UniPoD S80000 series supernode enables the deployment of one CPU and four AI accelerator cards within a single compute node, supporting high-power deployments of hundreds of kilowatts per cabinet;
  • Lossless Network:Featuring the industry's first single-chip 102.4T intelligent computing switch with support for 64×1.6T high-density ports, the system achieves congestion-free forwarding through millisecond-level micro-switching detection and DDC cell-level scheduling. Wide-area lossless transmission delivers latency below 1 ms over distances of 1,500 km and enables long-distance computing power coordination across 2,000 km;
  • High-Performance Storage:The X20836 all-flash storage system achieves an industry-leading density of 18 drive bays per U, delivering 200 GB/s bandwidth and 3 million IOPS per node. The XCache inference acceleration system reduces Time To First Token (TTFT) latency by 90% and increases user concurrency tenfold.

Software-Optimized System Capabilities:

  • Training Optimization:Leveraging Flash Checkpoint, long-sequence sharding, computation-communication overlap, and GDS storage acceleration, the system overcomes bottlenecks in GPU memory, communication, and disk I/O, achieving a maximum effective cluster computing utilization rate of over 82%;
  • Inference Optimization:By implementing PD disaggregation, model splitting, multi-level KV cache offloading, and CXL memory pool acceleration, overall inference performance has been more than doubled, enabling support for higher concurrency requests;
  • Communication Optimization:Leveraging global traffic navigation, collective communication operators, and GDA/PXN communication optimization, cross-node transmission bottlenecks in 10,000-GPU clusters have been resolved, achieving a scaling efficiency of 91% for 1,000-GPU clusters;
  • Scheduling Optimization:Utilizing KV cache affinity awareness, topology-aware computing power scheduling, token load balancing, and busy-idle queue scheduling for training and inference workloads, resource contention between training and inference tasks is effectively mitigated, maximizing cluster resource utilization.

In the era of the token economy, the core of computing power competition has shifted from total GPU volume to systemic capability depth. The H3C Lingxi Intelligent Computing Solution establishes a comprehensive closed-loop capability encompassing construction, operation, governance, and tuning. Through full-stack software-hardware synergy, it continuously reduces the comprehensive cost of token production, enabling enterprises to transcend inefficient computing power investment cycles and steadily advance the deployment of diverse AI business scenarios.