Serverless GPU Inference
Introducing: L40S GPUs on Modal

Introducing: L40S GPUs on Modal

12/19/2024

What this post added

Introduced support for NVIDIA L40S GPUs, offering 48GB of DDR6 RAM and significantly improved FP16/BF16 and FP8 Tensor Core arithmetic bandwidth compared to the A10 GPU. This enables running larger models and provides substantial speedups for both memory-bound and compute-bound inference workloads without requiring user tuning.

Read the original post ↗