iPhone
17 updates on iPhone.
Edge0 releases MoE expert-offloading framework, demos 35B model on iPhone
Edge0 streams mixture-of-experts weights from storage. Its 35B tier reports 2.9 GiB peak memory on a Mac mini M4 Pro; a launch post shows a 35B model on an iPhone.
iPhone 18 Pro: A20 Pro adds a dual 16-core Neural Engine
Apple says the A20 Pro carries 32 Neural Engine cores in total, double the AI processing power of A19 Pro, with 50 percent more memory bandwidth.
Artificial Analysis benchmarks 33 local models on an iPhone 17 Pro
A benchmark of quantised small models on an iPhone 17 Pro reports intelligence scores, generation times and peak memory between 0.4 and 6.9 GB.
Meta AI builds MobileMoE, 5.3B parameters with 0.9B active per token
Meta AI trained three on-device mixture-of-experts models that store 1.3B to 5.3B parameters and run 272M to 922M of them per token.
iPhone 16 Pro settles 41.5 percent below peak over 20 back-to-back prompts
Four authors ran Qwen 2.5 1.5B 20 times in a row on four edge platforms and report the iPhone 16 Pro settling 41.5 percent below its peak throughput.
A19 Pro puts Neural Accelerators in every GPU core
Apple says the iPhone 17 Pro chip pairs Neural Accelerators in each of six GPU cores with a 16-core Neural Engine to run large local language models.
Locally AI runs Llama, Gemma and Qwen offline on iPhone and iPad through MLX
Adrien Grondin shipped a free iPhone and iPad app that downloads open-weight models and runs them on device, built on Apple silicon through MLX.
HuggingSnap describes what the iPhone camera sees with a 500M model on the phone
Hugging Face released an iPhone app that runs SmolVLM2 at 500M parameters through MLX, describing camera scenes, photos and video with no cloud call.
EXO Labs benchmarks put Llama 3.1 8B at 14 tok/s on an iPhone 15 Pro
EXO Labs published automated inference benchmarks from real devices, covering iPhone 15 Pro, Galaxy S24 Ultra, Mac mini M4 Pro clusters and desktop GPUs.
PalmBench finds iPhones running local LLMs about three times faster than Android phones
A benchmark of quantised LLMs on eight phones and boards reports throughput, memory, power and heat, with iPhones ahead and 4-bit drawing more power than 3-bit.
Apple's MobileCLIP-S0 encodes an image in 1.5 ms on an iPhone 12 Pro Max
Apple timed its image-text models on an iPhone and released four variants, the weights and the reinforced DataCompDR dataset.
Apple Intelligence pairs a 3-billion-parameter on-device model with a server model
Apple reports 0.6 ms per prompt token and 30 tokens per second on an iPhone 15 Pro for a model compressed to an average of 3.7 bits per weight.