AI traffic pushes network testing toward system-level validation, Viavi says

0
1
AI traffic pushes network testing toward system-level validation, Viavi says


Mike Jack, director of product marketing at Viavi, told RCR that the most acute testing and assurance challenges are concentrated within data-center fabrics and data-center interconnect, where AI workloads are most tightly coupled.

In sum – what to know

AI traffic changes – AI training creates sustained, synchronized all-to-all traffic that makes latency, packet loss and link stability critical requirements.

Testing goes holistic – Validation is shifting from individual components toward entire systems operating under realistic traffic, congestion and failure scenarios.

Observability deepens – AI workloads require fine-grained, real-time visibility into network behavior and its effect on application outcomes.

AI workloads are changing network behavior by introducing traffic patterns that are more intense and synchronized than those of traditional cloud environments, according to Mike Jack, director of product marketing at Viavi.

Jack said in an interview with RCR Wireless News that conventional applications generate short, bursty flows tied to user transactions, while AI training produces sustained, high-bandwidth “all-to-all” communication between thousands of accelerators.

“These synchronized exchanges create bursty but coordinated traffic that must complete successfully across all nodes before computation can proceed. Any delay or packet loss can block an entire training cycle, making low latency, lossless performance, and tight synchronization critical requirements,” he said.

AI systems are also expanding rapidly, with clusters growing from hundreds to tens of thousands of accelerators. Jack said this is driving a shift from traditional front-end networks to specialized AI back-end fabrics. These architectures must deliver significantly higher bandwidth per node, moving toward 800G and 1.6T, and support east-west traffic at unprecedented scale.

Jack said the most acute testing and assurance challenges are concentrated within data-center fabrics and data-center interconnect, where AI workloads are most tightly coupled. These environments must support massive GPU clusters with highly synchronized communication patterns, meaning the performance of each individual link directly impacts the entire workload.

“In AI-scale deployments, even a single link failure or instability can stall a distributed training job, creating a level of sensitivity that is far greater than in traditional networks,” he said.

Operators are deploying millions of optical links while increasing per-link speeds to 800G and 1.6T. Jack said the combination of high density, strict latency requirements, and diverse interconnect technologies makes intra-data-center networks the most challenging segment to validate.

Testing is also shifting from a traditional, component-level approach to a holistic, system-level model that reflects the behavior of real AI workloads. AI networks require validation not just of individual links or devices, but of the entire system operating under realistic conditions.

Jack said validation must occur continuously across the lifecycle, from design through deployment and operations. AI environments are highly dynamic, with workloads that vary between training and inference, each placing different demands on the network.

This requires flexible, scalable testing approaches that can replicate diverse traffic patterns and ensure that performance, scalability, and reliability are maintained as conditions change. Jack also said the complexity of AI workloads requires validation of how network topology, protocols, and security mechanisms interact under real-world conditions.

AI workloads are changing requirements for network observability as well. Jack said operators need a step change from coarse, aggregated monitoring to fine-grained, real-time visibility into network behavior.

“Because distributed AI jobs depend on synchronized communication, operators must be able to detect and diagnose issues such as congestion, latency spikes, and link instability as they occur,” he said.

Jack said advanced telemetry and analytics are being adopted to correlate network conditions directly with application outcomes. Operators increasingly need to understand not just whether the network is performing within specification, but how it is impacting GPU utilization, job completion times, and overall system efficiency.

Traditional testing approaches are no longer sufficient for AI-scale environments, Jack said. Legacy methods focused on validating average performance and individual components, while AI workloads demand assurance of system-level behavior under extreme and highly variable conditions. “As a result, operators are adopting more advanced validation techniques, including AI workload emulation, digital twins, and continuous validation strategies,” he said.