
5/20/2025 · Mogan Shieh, Terry Heo, Jingjiang Li
What this post added
Introduced LiteRT, a new platform for on-device ML inference on mobile GPUs and NPUs. Key features include MLDrift for improved GPU acceleration with optimized data organization and workgroup optimization, NPU support co-developed with MediaTek and Qualcomm, simplified APIs for specifying target hardware accelerators (GPU, NPU), seamless buffer interoperability using TensorBuffer for zero-copy data transfers between hardware memory types, and asynchronous execution leveraging OS-level mechanisms for parallel processing across CPU, GPU, and NPUs.