Apple
14 updates on Apple.
iPhone 18 Pro: A20 Pro adds a dual 16-core Neural Engine
Apple says the A20 Pro carries 32 Neural Engine cores in total, double the AI processing power of A19 Pro, with 50 percent more memory bandwidth.
Ornith releases Ornith-1.5, a 9B model with a mobile build for iPhone and Android
Ornith-1.5 comes in 9B, 35B and 397B sizes, and the 9B model scores 70.6 on SWE-bench Verified and has a mobile build for iPhone and Android.
Apple ships Python bindings for the on-device Foundation Models framework
The apple-fm-sdk package calls the on-device Apple Intelligence model from Python on macOS 26, for scripting and batch evaluation outside Swift.
A19 Pro puts Neural Accelerators in every GPU core
Apple says the iPhone 17 Pro chip pairs Neural Accelerators in each of six GPU cores with a 16-core Neural Engine to run large local language models.
Apple puts the cost of 2-bit compression at 3.4 MMLU points
Apple measures its on-device model at 67.8 MMLU in 16 bits and 64.4 after compression to 2 bits per weight, and details the distillation pipeline behind it.
Apple opens its on-device model to all apps with the Foundation Models framework
Any app can call the roughly 3-billion-parameter on-device model from Swift, offline and free of charge, with guided generation and tool calling.
Locally AI runs Llama, Gemma and Qwen offline on iPhone and iPad through MLX
Adrien Grondin shipped a free iPhone and iPad app that downloads open-weight models and runs them on device, built on Apple silicon through MLX.
Apple team finds H100 last on tokens per dollar for models up to 2B
Seven Apple authors measured tokens per dollar for LLaMA-style models from 100M to 2B and found H100s last at every size, behind cheaper A100s.
Apple's MobileCLIP-S0 encodes an image in 1.5 ms on an iPhone 12 Pro Max
Apple timed its image-text models on an iPhone and released four variants, the weights and the reinforced DataCompDR dataset.
Apple Intelligence pairs a 3-billion-parameter on-device model with a server model
Apple reports 0.6 ms per prompt token and 30 tokens per second on an iPhone 15 Pro for a model compressed to an average of 3.7 bits per weight.
ExecuTorch alpha runs Llama 2 7B on iPhone 15 Pro and Galaxy phones
PyTorch's edge runtime brought 4-bit Llama 2 7B to iPhone and Galaxy handsets, added early Llama 3 8B support and leaned on Apple, Arm and Qualcomm.
Apple releases OpenELM at 270M to 3B, with parameters spread unevenly across layers
Apple published four models from 270M to 3B parameters with the full training framework and code to run them through MLX on Apple silicon.