
11/24/2025 · Lu Wang, Weiyi Wang, Andrew Zhang
What this post added
Introduced the LiteRT Qualcomm AI Engine Direct (QNN) Accelerator, replacing the TFLite QNN delegate. This provides a unified, simplified mobile deployment workflow for Android developers by abstracting vendor-specific SDKs and SoC fragmentation. It enables seamless deployment across supported devices with AOT or on-device compilation. The accelerator supports extensive LiteRT ops for full model delegation to the NPU, achieving state-of-the-art performance for LLMs and GenAI models. Detailed benchmarks show significant speedups over CPU and GPU, and a streamlined 3-step getting started guide is provided for AOT compilation and deployment via Google Play AI Packs.