
5/27/2026
What this post added
This post details the STAC-AI LANG6 benchmark results for LLM inference in finance, specifically highlighting the performance of NVIDIA Blackwell GPUs (HGX B200 and RTX PRO 6000) compared to Hopper (GH200). It showcases the use of NVFP4 quantization for Blackwell and FP8 for Hopper, both optimized with TensorRT LLM. The results demonstrate significant throughput and latency improvements, with HGX B200 achieving up to 2.8x performance uplift over GH200 in batch mode. It also touches upon interactive mode metrics like reaction time and words per second, and the importance of chat template application during inference.