Infinigence Releases PDD Cross-Cluster Inference Architecture: First-Token Latency Down 51.5%, Cost Down 37.5%
Infinigence has released the full technical report for its PDD cross-cluster heterogeneous inference architecture, first disclosed at WAIC 2026. The architecture deconstructs the traditional PD separation link into three layers, achieving 51.5% reduction in first-token latency and 37.5% cost reduction per token on an 80Gbps WAN Ethernet environment.














