Model Quantization for Inference Optimization
NVIDIA Blackwell Sets STAC-AI Record for LLM Inference in Finance | NVIDIA Technical Blog

NVIDIA Blackwell Sets STAC-AI Record for LLM Inference in Finance | NVIDIA Technical Blog

5/27/2026

What this post added

This post details the STAC-AI LANG6 benchmark results for LLM inference in finance, specifically highlighting the performance of NVIDIA Blackwell GPUs (HGX B200 and RTX PRO 6000) compared to Hopper (GH200). It showcases the use of NVFP4 quantization for Blackwell and FP8 for Hopper, both optimized with TensorRT LLM. The results demonstrate significant throughput and latency improvements, with HGX B200 achieving up to 2.8x performance uplift over GH200 in batch mode. It also touches upon interactive mode metrics like reaction time and words per second, and the importance of chat template application during inference.

Read the original post ↗