Mobile AI news timeline

September 2026

  1. MediaTek launches Dimensity 9600 Pro, a 2nm chip for on-device models up to 30B parameters
  2. Arm recaps Arm Create China and shows Qwen3-TTS 0.6B running on a vivo X300 CPU
  3. Edge0 releases MoE expert-offloading framework, demos 35B model on iPhone
  4. llama.cpp Hexagon NPU backend tested on a Snapdragon 8 Gen 3 phone
  5. iPhone 18 Pro: A20 Pro adds a dual 16-core Neural Engine
  6. Arm unveils CSS for Mobile 2 with C2 CPU cluster, up to 1.7x faster on AI models
  7. Arm unveils Mali G2-Ultra NX GPU with neural accelerators in every shader core
  8. OpenBMB releases MiniCPM5-2B for local deployment

August 2026

  1. Google lists first Gemini Nano 4 phones, requires Nano 3 for Gemini Intelligence
  2. Artificial Analysis benchmarks 33 local models on an iPhone 17 Pro
  3. Ornith releases Ornith-1.5, a 9B model with a mobile build for iPhone and Android
  4. Online-SDFT reports 70.28% routing accuracy and updates a LoRA adapter on the phone
  5. Pixel Watch 5 adds offline Gemini commands and faster on-device smart replies
  6. Pixel 11 series: Tensor G6 adds 50 percent more TPU compute
  7. RikkaHub Agent test: Android phone agent compiles whisper.cpp on its own
  8. Liquid AI releases LFM2.5-2.6B for on-device agents

July 2026

  1. Gemini Nano 4 ships on Samsung foldables with ML Kit Prompt API access
  2. FBLayout fine-tunes transformers on phone GPUs 2.2 to 5.7 times faster

June 2026

  1. Qualcomm CEO says agents will become the new app, cites more than 40 device designs
  2. MLPerf Mobile v6.0 adds Llama tests in 1B, 3B and 8B sizes
  3. llada.cpp runs a diffusion LLM on a Snapdragon NPU up to 42 times faster
  4. CAPED redacts phone screenshots before a cloud GUI agent sees them

May 2026

  1. Apertus Mini distils an open-data 8B into 0.5B, 1.5B and 4B models
  2. Meta AI builds MobileMoE, 5.3B parameters with 0.9B active per token
  3. LiteRT-LM reports 52 decode tokens per second for Gemma 4 E2B on an Android GPU
  4. Quant.npu runs 4-bit LLMs up to 15.1 percent faster on a Qualcomm NPU

April 2026

  1. Tencent open-sources a 440 MB offline translation model for phones
  2. Airgap is a React Native kit for support chatbots that answer offline
  3. Tempo compresses hour-long video with a 2B vision model so a 4B LLM can answer
  4. Gemma 4 lands on the edge with Agent Skills in Google AI Edge Gallery

March 2026

  1. iPhone 16 Pro settles 41.5 percent below peak over 20 back-to-back prompts
  2. Meta's MobileLLM-Flash goes shallow and wide for 1.8 times faster prefill
  3. FlexServe runs on-device LLM inference inside TrustZone for 4.41 percent more latency
  4. Alibaba adds 0.8B and 2B sizes to Qwen3.5, with 262K context and a vision encoder

February 2026

  1. ClawMobile tries system commands before screen taps and finishes all six test tasks
  2. Apple ships Python bindings for the on-device Foundation Models framework
  3. Alibaba releases GUI-Owl-1.5 agent models from 2B to 32B under MIT
  4. Show HN: Off Grid runs text, image, vision and speech models offline on phones

November 2025

  1. Nemotron-Flash-1B decodes 1.9 times faster than Qwen3-0.6B on an H100
  2. Intelligence per watt puts local model coverage at 88.7% of real queries
  3. Meta's MobileLLM-Pro runs a 128k context from a 590 MB 4-bit build
  4. Confidant fine-tunes Phi2-2.7B across a phone and two laptops in 40.1 hours

October 2025

  1. ExecuTorch 1.0 reaches general availability for on-device PyTorch models
  2. lm-Meter times on-device inference and finds prefill, not decode, is the bottleneck

September 2025

  1. Meta trains 140M to 950M reasoning models on 4.2T tokens
  2. A19 Pro puts Neural Accelerators in every GPU core

August 2025

  1. ShadowNPU scores attention on the Snapdragon NPU and reports 4.5 times faster inference
  2. P/D-Device prefills in the cloud and decodes on the phone, cutting TTFT 60%

July 2025

  1. Benchmark of 68 small language models finds architecture outweighs size on device
  2. Apple puts the cost of 2-bit compression at 3.4 MMLU points
  3. Liquid AI releases LFM2, three CPU-first models from 350M to 1.2B
  4. Hugging Face's SmolLM3 is a 3B model with 128k context and two reasoning modes

June 2025

  1. Apple opens its on-device model to all apps with the Foundation Models framework

May 2025

  1. Vijay Janapa Reddi argues edge generative AI needs models under 1B parameters
  2. Gemma 3n runs 5B and 8B models in 2 GB and 3 GB of memory
  3. Google AI Edge Gallery runs Gemma 3 1B and Qwen2.5 offline on Android
  4. Alibaba gives Qwen3's 0.6B and 1.7B models a reasoning switch
  5. Google's ML Drift runs Llama 3.1 8B on a phone GPU at 12.7 tokens per second

April 2025

  1. D2MoE picks a bit-width per token, 1.39 times the throughput at up to 53 percent less memory
  2. Locally AI runs Llama, Gemma and Qwen offline on iPhone and iPad through MLX

March 2025

  1. HuggingSnap describes what the iPhone camera sees with a 500M model on the phone
  2. ROMA keeps a 4-bit 3B model in on-chip ROM and reports 31,800 tok/s in synthesis
  3. Gemma 3 adds a 1B size that fits in 0.5 GB as an int4 checkpoint
  4. Flower Intelligence runs models on device, with remote handoff off by default
  5. Raspberry Pi 5 beats human reading speed only up to Phi 3.5 mini

February 2025

  1. MobiLLM fine-tunes OPT-1.3B in 4.5 GB on device by moving backpropagation to a server
  2. GenAI at the edge survey lists 12 accelerators, 8 of them only simulated
  3. Meta's ParetoQ compares five bit widths and puts the accuracy cliff at 1 bit

January 2025

  1. Alibaba MNN runs 4-bit LLMs on phone CPUs and GPUs, with a multimodal Android app
  2. Amazon survey puts some small models at 10 to 100 times their parameter count

December 2024

  1. EXO Labs benchmarks put Llama 3.1 8B at 14 tok/s on an iPhone 15 Pro

November 2024

  1. BlueLM-V-3B runs a multimodal model on a Dimensity 9300 NPU in 2.2 GB
  2. PhoneLM searches for a fast architecture before training it and hits 58 tok/s

October 2024

  1. Hugging Face trains SmolLM2 at 135M, 360M and 1.7B on up to 11T tokens
  2. Apple team finds H100 last on tokens per dollar for models up to 2B
  3. Mistral puts Ministral 3B and 8B on devices with 128k context
  4. PalmBench finds iPhones running local LLMs about three times faster than Android phones

September 2024

  1. AMD trains its first small language model, AMD-Llama-135M, on MI250 accelerators
  2. Meta ships Llama Stack with Swift and Kotlin clients for on-device inference
  3. Meta releases Llama 3.2 1B and 3B for phones and edge devices
  4. CoMiGS splits on-device fine-tuning into shared generalists and private specialists
  5. ElastiLM resizes a shared phone LLM per request and switches in 0.31 seconds
  6. Ai2 releases OLMoE, 7B parameters with 1B active per token

July 2024

  1. Gemma 2 2B is distilled from a larger model and scores 1126 on Chatbot Arena
  2. torchchat runs Llama 3 8B on a Galaxy S23 and iPhone at more than 8 tok/s
  3. Spectra ships ternary 3.9B models that match 4-bit quantisation at half the bits
  4. Hugging Face releases SmolLM at 135M, 360M and 1.7B parameters
  5. Alibaba builds Qwen2's 0.5B and 1.5B sizes for phones, earphones and glasses
  6. llm.npu prefills 1,106 tok/s for Qwen1.5-1.8B on a Snapdragon 8 Gen 3

June 2024

  1. Apple's MobileCLIP-S0 encodes an image in 1.5 ms on an iPhone 12 Pro Max
  2. TensorOpera releases Fox-1, a 1.6B model trained on 3 trillion tokens
  3. BUPT measures 22 LLMs on four Android phones at about 200 ms per token
  4. Apple Intelligence pairs a 3-billion-parameter on-device model with a server model
  5. PowerInfer-2 runs a 47B model on a OnePlus 12 at 11.68 tokens per second

May 2024

  1. BUPT proposes one 9.2B model in the OS that all apps call through adapters

April 2024

  1. ExecuTorch alpha runs Llama 2 7B on iPhone 15 Pro and Galaxy phones
  2. Apple releases OpenELM at 270M to 3B, with parameters spread unevenly across layers
  3. Microsoft runs Phi-3-mini offline on an iPhone 14 at over 12 tokens per second
  4. Octopus v3 picks an action from an image and a query in under 1B parameters
  5. Octopus v2 is a 2B model that calls Android APIs with one token per function
  6. Octopus fine-tunes a 2B model to 93 percent on API function calls

March 2024

  1. One shared on-device LLM keeps a context per app and switches in 0.27 seconds
  2. Google ships an experimental MediaPipe LLM Inference API for web, Android and iOS

February 2024

  1. MeRino designs sub-100M language models that run 4.9 times faster than OPT-350M
  2. Stability AI trains Stable LM 2 1.6B on seven languages and 2 trillion tokens
  3. BitNet b1.58 gives every weight three values and runs 3B in 2.22 GB
  4. MobiLlama is a fully transparent 0.5B model that runs in 770 MB on a phone
  5. Qualcomm AI Hub opens 75 pre-optimised models and profiling on hosted Snapdragon phones
  6. Meta's MobileLLM trades width for depth and gains 2.7 and 4.3 points
  7. TinyLLaVA's 3.1B model outscores 7B LLaVA-1.5 on seven of nine benchmarks
  8. Gemma 2B and 7B open the Gemma line, built on Gemini research
  9. MobileVLM V2 runs 1.7B, 3B and 7B vision models on 144 image tokens

January 2024

  1. Galaxy S24 becomes the second phone line to run Gemini Nano
  2. Arm runs Llama 2 7B at 9.6 tokens per second on three mobile CPU cores
  3. TinyLlama pretrains a 1.1B model on 3 trillion tokens

December 2023

  1. Microsoft releases Phi-2, a 2.7B model it says matches models 25 times larger
  2. Apple researchers run models twice the size of available DRAM from flash
  3. Stability AI releases StableLM Zephyr 3B for edge devices
  4. Gemini Nano ships on the Pixel 8 Pro and Android gets AICore
  5. Apple publishes MLX, where CPU and GPU share arrays without copies
  6. LLM.swift wraps llama.cpp for on-device text generation in Swift apps

October 2023

  1. Snapdragon 8 Gen 3 targets 10-billion-parameter models on device

September 2023

  1. Microsoft carries its textbook data recipe from code to reasoning with the 1.3B phi-1.5

August 2023

  1. MIT HAN Lab publishes TinyChatEngine for 4-bit LLMs on laptops and Raspberry Pi
  2. Hugging Face publishes swift-transformers for Core ML models in Swift apps

June 2023

  1. Microsoft trains phi-1 to 50.6 percent on HumanEval with 1.3B parameters
  2. LLMFarm runs llama.cpp models offline on iOS and macOS

May 2023

  1. RWKV trains like a transformer and runs with constant memory per token
  2. Qualcomm argues for hybrid AI, citing 10 times the cost per generative AI search query
  3. MLC LLM brings local language models to iPhone, browsers and consumer GPUs

April 2023

  1. LaMini-LM distils models from 61M parameters up on 2.58M instructions

March 2023

  1. Sherpa runs LLaMA on an Android phone through a Flutter chat app
  2. Georgi Gerganov publishes llama.cpp, LLM inference in plain C and C++

February 2023

  1. Qualcomm runs Stable Diffusion on an Android phone for the first time