Models
Alibaba adds 0.8B and 2B sizes to Qwen3.5, with 262K context and a vision encoder
Alibaba released Qwen3.5-0.8B and Qwen3.5-2B, dense vision-language models with a 262,144-token context, Apache 2.0 weights and 4-bit MNN builds.
Alibaba releases GUI-Owl-1.5 agent models from 2B to 32B under MIT
Tongyi Lab open-sourced six GUI agent checkpoints from 2B to 32B and reports 71.6 on AndroidWorld, with every benchmark run server-side, not on a phone.
Nemotron-Flash-1B decodes 1.9 times faster than Qwen3-0.6B on an H100
Nvidia designed a hybrid 1B and 3B model family around measured decoding latency on an H100 rather than parameter count, and released three checkpoints.
Meta's MobileLLM-Pro runs a 128k context from a 590 MB 4-bit build
Meta Reality Labs released a 1.08B on-device model with a 128k context window, measured at 33.6 tok/s decode on a Galaxy S25 CPU.
Meta trains 140M to 950M reasoning models on 4.2T tokens
MobileLLM-R1 spans 140M to 950M parameters, trained on 4.2T tokens, and Meta scores the 950M model at 74.0 on MATH500.
Benchmark of 68 small language models finds architecture outweighs size on device
A study of 68 models from 100M to 5B puts Phi-3 near 70 percent accuracy and finds first-token time and memory tracking architecture, not parameter count.
Apple puts the cost of 2-bit compression at 3.4 MMLU points
Apple measures its on-device model at 67.8 MMLU in 16 bits and 64.4 after compression to 2 bits per weight, and details the distillation pipeline behind it.
Liquid AI releases LFM2, three CPU-first models from 350M to 1.2B
Liquid AI released open-weight models of 350M, 700M and 1.2B parameters and reports 2x faster decode and prefill than Qwen3 on CPU.
Hugging Face's SmolLM3 is a 3B model with 128k context and two reasoning modes
Hugging Face released a 3B model with a 128k context window, six languages and a switchable reasoning mode, along with the full training recipe.
Gemma 3n runs 5B and 8B models in 2 GB and 3 GB of memory
Google previewed a mobile-first model whose Per-Layer Embeddings cut RAM use, and said the same architecture powers the next Gemini Nano.
Alibaba gives Qwen3's 0.6B and 1.7B models a reasoning switch
Qwen3-0.6B and Qwen3-1.7B carry the family's switch between a reasoning mode and a fast mode, with 32K context and Apache 2.0 weights.
Gemma 3 adds a 1B size that fits in 0.5 GB as an int4 checkpoint
Google released Gemma 3 at 1B, 4B, 12B and 27B with quantisation-aware int4 checkpoints of 0.5 GB and 2.6 GB for the two smallest sizes.