Modal's multi-node GPU clusters go generally available, billed by the second
Distributed training and inference can now rent RDMA-connected GPU clusters by the second; all performance figures are the company's own.
Modal announced on October 1 that Modal Clusters, its multi-node GPU cluster product, is now generally available, requested through a single decorator.
Clusters connect over RDMA, which the company says reaches up to 6.4 Tbps per node and is auto-configured for PyTorch and NCCL; a gang scheduler allocates all nodes at once from Modal's shared capacity pool, and usage is billed by the second. The company says the product was battle-tested for a year and a half before this release.
For teams without their own machine room, this means distributed training and inference can rent multi-node capacity in short bursts instead of hourly reservations. The bandwidth and acquisition-speed figures come from Modal's own blog and have no third-party measurements yet.
Sources:https://modal.com/blog/modal-clusters-generally-available