Hardware

  1. llama.cpp Hexagon NPU backend tested on a Snapdragon 8 Gen 3 phone

    A user report puts Gemma 3 4B at 12.5 tokens per second of generation on a OnePlus 12 using llama.cpp’s Hexagon NPU backend, at about CPU speed but without the heat.

  2. iPhone 18 Pro: A20 Pro adds a dual 16-core Neural Engine

    Apple says the A20 Pro carries 32 Neural Engine cores in total, double the AI processing power of A19 Pro, with 50 percent more memory bandwidth.

  3. Pixel Watch 5 adds offline Gemini commands and faster on-device smart replies

    Offline voice commands on the Pixel Watch 5 run on an on-device model, and a Gemini Nano upgrade makes smart replies 50 percent faster.

  4. Pixel 11 series: Tensor G6 adds 50 percent more TPU compute

    Google says Tensor G6 with the latest Gemini Nano model processes on-device AI tasks up to 3.5 times faster while using up to 3.5 times less energy.

  5. Qualcomm CEO says agents will become the new app, cites more than 40 device designs

    Cristiano Amon tells CNBC that Qualcomm has over 40 designs for AI wearables and that smart glasses shipments could reach hundreds of millions a year.

  6. iPhone 16 Pro settles 41.5 percent below peak over 20 back-to-back prompts

    Four authors ran Qwen 2.5 1.5B 20 times in a row on four edge platforms and report the iPhone 16 Pro settling 41.5 percent below its peak throughput.

  7. Intelligence per watt puts local model coverage at 88.7% of real queries

    Stanford and Together AI measured 20+ local models on 1M real queries and propose accuracy per watt as the metric for local inference.

  8. A19 Pro puts Neural Accelerators in every GPU core

    Apple says the iPhone 17 Pro chip pairs Neural Accelerators in each of six GPU cores with a 16-core Neural Engine to run large local language models.

  9. ShadowNPU scores attention on the Snapdragon NPU and reports 4.5 times faster inference

    Peking University and BUPT move attention token scoring onto a Snapdragon NPU and report up to 4.5 times faster inference using a single CPU core.

  10. ROMA keeps a 4-bit 3B model in on-chip ROM and reports 31,800 tok/s in synthesis

    A synthesised 7 nm accelerator holds a 4-bit 3B Llama in read-only memory and the LoRA adapter in SRAM, with no chip and no FPGA prototype built.

  11. Raspberry Pi 5 beats human reading speed only up to Phi 3.5 mini

    Helsinki and EURECOM researchers measured 11 models on a Raspberry Pi 5 and a Jetson Orin Nano, where memory and battery gave out before speed did.

  12. GenAI at the edge survey lists 12 accelerators, 8 of them only simulated

    A Johns Hopkins and Duke survey of generative AI on edge devices puts peak accelerator efficiency at 74.34 TOPS/W, with 8 of 12 designs only simulated.