WebGPU

5 updates on WebGPU.

  1. Flower Intelligence runs models on device, with remote handoff off by default

    Flower Labs released a preview library that runs Llama 3.2 and SmolLM2 locally via WebLLM or MLX Swift, and calls its remote service only if an app enables it.

  2. Hugging Face trains SmolLM2 at 135M, 360M and 1.7B on up to 11T tokens

    Hugging Face released SmolLM2 in three sizes trained on up to 11 trillion tokens, with 4-bit builds from 118 MB for on-device runtimes.

  3. Hugging Face releases SmolLM at 135M, 360M and 1.7B parameters

    Three base models trained on the newly released SmolLM-Corpus, with published memory footprints from 109.78 MB to 3422.76 MB.

  4. MLC LLM brings local language models to iPhone, browsers and consumer GPUs

    A compiler stack built on Apache TVM deploys chat models natively to iOS, browsers and consumer GPUs, with an iPhone build handed out through TestFlight.

  5. George Hotz starts tinygrad, a framework that ports to an accelerator in about 25 ops

    The tiny corp framework caps its repository at 26,500 lines, ships Metal, Adreno and WebGPU backends, and runs openpilot on a Snapdragon 845 GPU.