
Qualcomm has introduced a new Hexagon NPU.
The architecture is aimed at agentic AI, which can understand context, work across apps, and perform multiple steps to complete a task.
The new NPU includes an Element Accelerator focused on transformer-based AI models. It also has 50% more shared memory, allowing more AI data to stay closer to the NPU and reducing the need to access the phone’s external memory.
Qualcomm says the NPU can run Mixture-of-Experts (MoE) models, which use only selected parts of a larger AI model for each task. For example, a 30-billion-parameter MoE model can activate around 3 billion parameters for each token generation step. This can reduce the amount of computing and memory needed.
The NPU supports INT2, INT4, INT8, FP8, and FP16 precision formats, giving developers different options for balancing AI performance, memory use, and model quality.
For INT4 models, Qualcomm claims up to 50% faster prefill performance, along with improvements in decoding and speculative decoding.
The Hexagon NPU will also work with the CPU to handle tasks such as routing and coordinating AI workloads. Qualcomm says this setup is intended to allow more AI processing to run directly on the device.

More details about Qualcomm’s next-generation mobile platform are expected at Snapdragon Summit 2026, which will take place September 22–24 in Maui, Hawaii.

ManilaShaker is a tech media producing insightful and helpful content for our local and growing international audience. Our goal is to create a premier Philippine digital consumer electronics resource that provides the most objective reviews and comparisons globally.