
AWS has validated a solution using Amazon SageMaker HyperPod and Qumulo Cloud Data Fabric to train large AI models using datasets stored in different regions. The architecture uses predictive caching to maintain local-level throughput over high-latency network links, eliminating the need for petabyte-scale data copying between geographic locations.
Remote data access architecture
The validated configuration pairs SageMaker HyperPod managed infrastructure with Qumulo Cloud Native Qumulo and Cloud Data Fabric. This setup allows training nodes to mount a local spoke that connects to a central hub holding the primary dataset via VPC peering. By projecting the dataset to the compute region, developers can run jobs without modifying existing training code.
To mitigate network latency, the system utilizes a feature called NeuralCache. This component uses a model to predict the specific 4 KB data blocks required by the training job, pre-caching them at the remote site. This predictive approach aims to make remote datasets perform as if they were stored on local hardware.
Performance testing and regional availability
AWS tested this approach by running identical training jobs on two clusters. A hub cluster was placed in US East (Ohio) near the data, while a spoke cluster operated in US West (Oregon) with 60 ms of network latency. Following an initial warmup period, the remote cluster achieved throughput identical to the co-located hub cluster.
The solution is currently presented as a validated architecture using Amazon SageMaker HyperPod, Amazon EKS, and Qumulo software. While it maintained performance at 60 ms latency, Qumulo reports the fabric remains functional on links exceeding 100 ms. Users must manage their own Qumulo instances and VPC peering to implement the design.
Original source
This report summarises the source below. Analysis is labelled separately; product and research claims remain attributed to their source.
Read the original at AWS ML