
Running Low-Latency Analytical Workloads with GPU-Accelerated Presto on NVIDIA GB200 NVL72 | NVIDIA Technical Blog
7/8/2026
This post demonstrates the performance of GPU-accelerated Presto on NVIDIA DGX B200 and GB200 NVL72 systems for low-latency analytical workloads. It highlights performance improvements of up to 8x lower latency compared to CPU-based Presto on TPC-H benchmarks using single-node DGX B200. It further details scaling out to GB200 NVL72, showcasing the benefits of NVIDIA GPUDirect Storage (GDS) with IBM Storage Scale for high I/O throughput. The post details specific optimizations like increasing I/O task size, using more I/O threads, and query rewrites (e.g., Q11) that led to a 64% reduction in query runtimes. It also mentions the importance of NVLink for GPU-to-GPU communication and UcxExchange for high-performance communications.
