
1/28/2026 · Lu Wang, Chintan Parikh, Jingjiang Li, Terry Heo
What this post added
LiteRT has fully graduated into production, offering a unified on-device AI inference framework. Key advancements include 1.4x faster GPU performance than TFLite, new NPU acceleration, simplified GPU/NPU workflows across platforms (Android, iOS, macOS, Windows, Linux, Web) via ML Drift (OpenCL, OpenGL, Metal, WebGPU), asynchronous execution, and zero-copy buffer interoperability for reduced latency. NPU integration is streamlined with AOT/JIT compilation and partnerships with MediaTek and Qualcomm, achieving up to 100x faster speeds than CPU. LiteRT provides superior cross-platform GenAI support with LiteRT Torch Generative API, LiteRT-LM, and LiteRT Converter & Runtime, outperforming Llama.cpp on CPU/GPU and offering 3x NPU acceleration for Gemma models. It also supports a broad range of open-weight models and offers broad ML framework support.