<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Google · LLMobile.news</title><link>https://llmobile.kavents.com/tags/google/</link><description>A concise news ticker covering AI on mobile devices, local models, apps and hardware.</description><language>en-GB</language><atom:link href="https://llmobile.kavents.com/tags/google/index.xml" rel="self" type="application/rss+xml"/><item><title>Google lists first Gemini Nano 4 phones, requires Nano 3 for Gemini Intelligence</title><link>https://llmobile.kavents.com/ticker/gemini-nano-4-first-devices/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/gemini-nano-4-first-devices/</guid><pubDate>Thu, 27 Aug 2026 16:00:00 +0200</pubDate><description>Google&amp;amp;rsquo;s ML Kit GenAI documentation now lists the first devices running nano-v4, 9to5Google reports. The list covers the Pixel 11, Pixel 11 Pro, Pixel 11 Pro XL and Pixel 11 Pro Fold, plus Samsung&amp;amp;rsquo;s Galaxy Z Flip8, Galaxy Z Fold8 and Galaxy Z Fold8 Ultra.
The same documentation sets Nano v3 or greater as the requirement for Gemini Intelligence, Google&amp;amp;rsquo;s on-device feature set. According to the report, that requirement first appeared in May 2026, was removed, and has now been reinstated. Listed hardware requirements include 12 GB or more of RAM, a qualified flagship system-on-chip, five or more OS upgrades and six years of security support.
Gemini Intelligence features named in the report include Rambler and Proactive Assistance on Pixel 11, and task automation across more than 40 apps on Samsung&amp;amp;rsquo;s foldables.
Source: https://9to5google.com/2026/08/27/gemini-intelligence-nano-4/
Read the article: https://llmobile.kavents.com/ticker/gemini-nano-4-first-devices/</description><category>Google</category><category>Android</category><category>Pixel</category><category>Samsung</category><category>Gemini Nano</category></item><item><title>Pixel Watch 5 adds offline Gemini commands and faster on-device smart replies</title><link>https://llmobile.kavents.com/ticker/pixel-watch-5-offline-gemini/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/pixel-watch-5-offline-gemini/</guid><pubDate>Wed, 12 Aug 2026 19:30:00 +0200</pubDate><description>Gemini Intelligence is coming to the Pixel Watch 5, 9to5Google reports. According to the report, offline Gemini commands on the watch use a separate on-device model when the phone or an internet connection is unavailable. Those commands cover timers and alarms, brightness and modes, music control, opening apps and starting workouts.
On-device smart replies now offer three responses instead of one, which the report attributes to a Gemini Nano upgrade that makes them 50 percent faster.
Other parts of the feature set depend on a connection. Proactive Suggestions, formerly Magic Cue, are generated on a paired Pixel 11 and bridged to the watch, and Personal Intelligence draws on Gmail, Calendar and Keep. The update also brings a new At a Glance space on the watch face for timers, workouts, music, flight details and navigation.
Source: https://9to5google.com/2026/08/12/pixel-watch-5-gemini-intelligence/
Read the article: https://llmobile.kavents.com/ticker/pixel-watch-5-offline-gemini/</description><category>Google</category><category>Pixel</category><category>Wearables</category><category>Gemini Nano</category></item><item><title>Pixel 11 series: Tensor G6 adds 50 percent more TPU compute</title><link>https://llmobile.kavents.com/ticker/pixel-11-tensor-g6/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/pixel-11-tensor-g6/</guid><pubDate>Wed, 12 Aug 2026 19:00:00 +0200</pubDate><description>Google has announced the Pixel 11, Pixel 11 Pro and Pixel 11 Pro XL, built around the Tensor G6 chip. Google states that Tensor G6 packs 50 percent more TPU compute and, paired with the latest Gemini Nano model, processes on-device AI tasks up to 3.5 times faster while using up to 3.5 times less energy.
The company also cites an upgraded CPU with 25 percent faster web browsing and 15 percent quicker app launches, and says the chip powers the 30x Super Zoom on the 5x telephoto lens. Google does not publish RAM figures, model sizes or per-task latency in the announcement.
Pre-orders opened on 12 August, with retail availability from 20 August.
▶Meet Google Pixel 11 ProLoading connects your browser to www.youtube-nocookie.com, which may process your IP address and use cookies.
Load contentOpen externallyVideo: Made by Google.
Source: https://blog.google/products-and-platforms/devices/pixel/google-pixel-11-pro-xl/
Read the article: https://llmobile.kavents.com/ticker/pixel-11-tensor-g6/</description><category>Google</category><category>Pixel</category><category>Chips</category><category>NPU</category><category>Gemini Nano</category></item><item><title>Gemini Nano 4 ships on Samsung foldables with ML Kit Prompt API access</title><link>https://llmobile.kavents.com/ticker/gemini-nano-4-mlkit-prompt-api/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/gemini-nano-4-mlkit-prompt-api/</guid><pubDate>Wed, 22 Jul 2026 18:00:00 +0200</pubDate><description>Google&amp;amp;rsquo;s Android developer blog states that Samsung&amp;amp;rsquo;s new foldable devices come with Gemini Nano 4, which it calls its latest on-device model. The post credits Nano 4 with support for more than 140 languages and better multimodal understanding.
Apps reach the model through ML Kit&amp;amp;rsquo;s Prompt API, which sends natural language requests on-device to Gemini Nano. It takes text, or a combination of image and text, and returns text or structured output. Google names structured output and thinking mode as the features to use for on-device intelligence.
The ML Kit release notes dated 14 July 2026 record the structured output API, system instructions and thinking mode arriving in the Prompt API, along with multi-image support and an output token limit raised to 4,096 tokens. A note dated 21 July records a fix for Gemini Nano v4 compatibility in the Prompt API on non-Pixel devices. The Prompt API moved from alpha to beta in January 2026 and carries no service level agreement or deprecation policy.
The post also points developers to app functions, which share an app&amp;amp;rsquo;s capabilities with the Gemini Intelligence system. Its remaining sections cover adaptive layouts, fold-aware design, CameraX and Wear OS widgets.
Source: https://android-developers.googleblog.com/2026/07/optimize-galaxy-screen-sizes.html
Read the article: https://llmobile.kavents.com/ticker/gemini-nano-4-mlkit-prompt-api/</description><category>Google</category><category>Android</category><category>Samsung</category><category>Gemini Nano</category><category>Developer tools</category></item><item><title>LiteRT-LM reports 52 decode tokens per second for Gemma 4 E2B on an Android GPU</title><link>https://llmobile.kavents.com/ticker/litert-lm-on-device-genai/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/litert-lm-on-device-genai/</guid><pubDate>Tue, 19 May 2026 18:00:00 +0200</pubDate><description>Google has detailed LiteRT-LM, the runtime it uses to run Gemma 4 on-device in Chrome, ChromeOS, the Pixel Watch and the Google AI Edge Gallery app. It sits on top of LiteRT, formerly TensorFlow Lite, and adds GenAI libraries over the XNNPACK and ML Drift kernels.
Running Gemma 4 E2B without multi-token prediction, Google measures 52 decode tokens per second on the Android GPU through OpenCL, 56 on iOS through Metal, and up to 76 in the browser on WebGPU. The Android figures come from a Samsung S26 Ultra, iOS from an iPhone 17 Pro, and web from Chrome on a 2024 MacBook Pro with an M4 Max.
Chart: Google. The published comparison puts LiteRT-LM prefill at 3,808 tokens per second against 292 for llama.cpp on the Android GPU, 2,878 against 1,493 for MLX on iOS, and 4,182 against 1,127 for ONNX on the web. Decode figures in the same chart are 52 against 19, 56 against 32 and 76 against 29.
Multi-token prediction raises decode further. Google reports up to a 2.2x speedup, taking Gemma 4 E2B from 52 to 85 tokens per second and E4B from 21 to 47, measured on a Samsung S26 Ultra GPU. The drafter and the primary model run on the same hardware block so the shared KV cache and activations stay in local memory.
Chart: Google. On memory, LiteRT-LM keeps per-layer embeddings out of memory and loads image and audio encoders only when a task needs them. Google states that the 2.58 GB Gemma 4 E2B model runs with a physical footprint of 607 MB on Apple mobile CPUs using XNNPACK weight caching.
For agentic use the runtime supports Thinking Mode, which developers can stream to the interface or strip to save KV cache space:
By dedicating a scratchpad for step-by-step reasoning before the model commits to an action, LiteRT-LM can significantly improve the output quality.
Constrained decoding enforces JSON schemas on tool payloads, and native function calling pauses execution, returns the tool-call request to the app and resumes on the tool&amp;amp;rsquo;s output. Session save and restore serialise KV cache state so long conversations skip the prefill phase on return.
Beyond the existing Kotlin and C++ interfaces on Android, Google is adding an open-source Swift API for Apple platforms and a JavaScript API for the web through WebAssembly and WebGPU.
Source: https://developers.googleblog.com/blazing-fast-on-device-genai-with-litert-lm/
Read the article: https://llmobile.kavents.com/ticker/litert-lm-on-device-genai/</description><category>Google</category><category>LiteRT</category><category>Gemma</category><category>Benchmarks</category><category>Android</category></item><item><title>Gemma 4 lands on the edge with Agent Skills in Google AI Edge Gallery</title><link>https://llmobile.kavents.com/ticker/gemma-4-edge-agent-skills/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/gemma-4-edge-agent-skills/</guid><pubDate>Thu, 02 Apr 2026 18:00:00 +0200</pubDate><description>Google DeepMind has launched Gemma 4 under the Apache 2.0 licence, and the Google AI Edge team has described how the E2B and E4B variants run on edge hardware. The models support more than 140 languages and a 128K context window.
Gemma 4 enables multi-step planning, autonomous action, offline code generation, and even audio-visual processing, all without specialized fine-tuning.
Image: Google. The Google AI Edge Gallery app on iOS and Android gains Agent Skills, which Google calls one of the first applications to run multi-step, autonomous agentic workflows entirely on-device. The examples given are querying Wikipedia, turning speech input into charts and flashcards, pairing photos with generated music, and building a small app that plays animal calls through conversation alone.
For in-app deployment Google points to LiteRT-LM. It reports that Gemma 4 E2B runs in under 1.5 GB of memory on some devices using 2-bit and 4-bit weights with memory-mapped per-layer embeddings, and that GPU optimisations process 4,000 input tokens across two skills in under 3 seconds. Constrained decoding enforces structured output, and dynamic context lengths let a single model span CPUs and GPUs.
On smaller hardware, LiteRT-LM reaches 133 prefill and 7.6 decode tokens per second on a Raspberry Pi 5 CPU, and 3,700 prefill and 31 decode tokens per second on the NPU of a Qualcomm Dragonwing IQ8, the processor behind the Arduino VENTUNO Q.
Gemma 4 is available with CPU and GPU support on Android and iOS, system-wide on Android through the new AICore Developer Preview, on Windows, Linux and macOS via Metal, and in the browser through WebGPU. Google also released a Python package and a litert-lm CLI for Linux, macOS and Raspberry Pi, with tool calling included.
Source: https://developers.googleblog.com/bring-state-of-the-art-agentic-skills-to-the-edge-with-gemma-4/
Read the article: https://llmobile.kavents.com/ticker/gemma-4-edge-agent-skills/</description><category>Google</category><category>Gemma</category><category>Agents</category><category>Open weights</category><category>LiteRT</category></item><item><title>Gemma 3n runs 5B and 8B models in 2 GB and 3 GB of memory</title><link>https://llmobile.kavents.com/ticker/gemma-3n/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/gemma-3n/</guid><pubDate>Tue, 20 May 2025 19:00:00 +0200</pubDate><description>Google previewed Gemma 3n on May 20, 2025, a model built for phones, tablets and laptops. Its 5B and 8B versions run with dynamic memory footprints of 2 GB and 3 GB, which Google says matches what 2B and 4B models normally need.
The saving comes from Per-Layer Embeddings, a Google DeepMind technique that keeps embedding parameters off the accelerator. For E2B that cuts what has to be loaded into accelerator memory from 5.44B parameters to 1.91B, with the 2.55B embedding parameters cached to fast storage instead. A MatFormer setup nests a smaller but fully functional 2B model inside the 4B one, so developers can trade quality against footprint. Google reports the model starting to respond about 1.5 times faster on mobile than Gemma 3 4B.
Gemma 3n takes audio, text, images and video as input, including interleaved inputs across modalities, and handles speech recognition and translation. Google states that the same architecture powers the next generation of Gemini Nano, which was to reach Android and Chrome later in 2025. Qualcomm Technologies, MediaTek and Samsung System LSI worked on the optimisation.
Diagram: Google. ▶Announcing Gemma 3n Preview: Powerful, Efficient, Mobile-First AILoading connects your browser to www.youtube-nocookie.com, which may process your IP address and use cookies.
Load contentOpen externallyVideo: Google for Developers. Update, June 26, 2025. Google released Gemma 3n as E2B and E4B, naming them for 2B and 4B effective parameters out of 5B and 8B raw ones. Google reports E4B scoring 1303 on LMArena and calls it the first model under 10 billion parameters to pass 1300, ahead of Llama 4 Maverick 17B 128E at 1292 and GPT 4.1-nano at 1288. The release adds a MobileNet-V5-300M vision encoder, which Google says handles up to 60 frames per second on a Pixel, and KV cache sharing that doubles prefill performance. Text covers 140 languages, multimodal understanding 35. Weights are on Hugging Face and Kaggle, with support in Ollama, llama.cpp, MLX, LiteRT and the Google AI Edge Gallery.
Chart: Google.
Source: https://developers.googleblog.com/en/introducing-gemma-3n/
Read the article: https://llmobile.kavents.com/ticker/gemma-3n/</description><category>Google</category><category>Gemma</category><category>Gemini Nano</category><category>Android</category><category>Open weights</category></item><item><title>Google AI Edge Gallery runs Gemma 3 1B and Qwen2.5 offline on Android</title><link>https://llmobile.kavents.com/ticker/googleai-edge-gallery/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/googleai-edge-gallery/</guid><pubDate>Sun, 18 May 2025 00:21:00 +0200</pubDate><description>Google&amp;amp;rsquo;s AI Edge team has published Google AI Edge Gallery, an app that runs generative models entirely on the device. The home screen offers three tasks, AI Chat for multi-turn conversation, Ask Image for questions about a photo from the camera or the gallery, and Prompt Lab for single-turn work, which covers free-form prompts, rewriting text in a chosen tone, summarising and generating code. Once a model file is on the phone, the app needs no connection and Google states that all processing happens on the device.
Screenshot: Google. The third screen shows the per-model settings, including the GPU and CPU toggle. Models come from the LiteRT Community organisation on Hugging Face, which the app&amp;amp;rsquo;s welcome text links to directly. The first release listed three downloads, Gemma 3 1B IT at 4-bit in 554.7 MB, Hammer 2.1 1.5B at 8-bit and Qwen2.5 1.5B Instruct at 8-bit, the latter two at roughly 1.6 GB each. Gated repositories ask for a Hugging Face access token, which the settings dialog stores and can clear again.
Anything outside that list has to arrive as a LiteRT .task bundle picked from device storage, and the app tells the user that no other file type is supported. Each model then carries a settings dialog with an accelerator switch between GPU and CPU, next to max tokens, top-k, top-p and temperature. The shipped list sets a different default per model, GPU first for Gemma 3 1B and CPU only for Qwen2.5 1.5B Instruct.
After every reply the app prints four of its own measurements, time to first token in seconds, prefill speed and decode speed in tokens per second, and total latency in seconds. It derives them from the session itself, dividing the prompt&amp;amp;rsquo;s token count by the time the first token took and counting decoded tokens against the clock that starts with that first token, so the numbers describe whatever hardware the app happens to be running on.
Google calls this an &amp;amp;ldquo;experimental Alpha release&amp;amp;rdquo; and ships it under the Apache 2.0 licence. There is no store listing at this point, so installing means taking the APK from the project&amp;amp;rsquo;s releases page on a device running Android 8.0 or newer, with a project wiki covering the corporate-device case. The README lists Android as available now and iOS as coming soon. Underneath, the app builds on LiteRT and Google&amp;amp;rsquo;s MediaPipe LLM inference API for Android.
Source: https://github.com/google-ai-edge/gallery
Read the article: https://llmobile.kavents.com/ticker/googleai-edge-gallery/</description><category>Google</category><category>Android</category><category>LiteRT</category><category>Hugging Face</category><category>Open source</category></item><item><title>Google's ML Drift runs Llama 3.1 8B on a phone GPU at 12.7 tokens per second</title><link>https://llmobile.kavents.com/ticker/scaling-on-device-gpu-inference/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/scaling-on-device-gpu-inference/</guid><pubDate>Thu, 01 May 2025 02:44:00 +0200</pubDate><description>Eight researchers at Google published ML Drift on May 1, 2025, a GPU inference framework for running large generative models on the graphics chip of a phone, a laptop or a Mac. On the Adreno 750 GPU of a Samsung S24, the authors measure 1370 tokens per second of prefill and 37.1 tokens per second of decode for Gemma2 2B, and 412 and 12.7 tokens per second for Llama 3.1 8B. They report a 5 times to 11 times prefill speedup over llama.cpp, MLC LLM and the Qualcomm AI Hub benchmark on Qualcomm GPUs, and the work is due at the CVPR 2025 Workshop on Efficient and On-Device Generation (EDGE).
The team benchmarked five mobile GPUs, the Adreno 830 in a Xiaomi 15 Pro with 16 GB of RAM, the Adreno 750 in a Samsung S24 with 8 GB, the Adreno 740 in a Samsung S23 Ultra with 8 GB, the Arm Immortalis-G720 in a vivo X100 Pro with 16 GB and the Arm Mali-G715 in a Pixel 9 with 12 GB, each at a fixed 1024 prefill and 256 generated tokens. Llama 3.1 8B with every weight at 8-bit ran out of memory on the S24, the S23 Ultra and the Pixel 9 and finished only on the two 16 GB phones, at 7.70 and 4.72 tokens per second of decode. A mixed scheme that holds attention weights at 8-bit and drops embedding and feed-forward weights to 4-bit ran the 8B model on all five devices and lifted decode by up to 1.9 times. ML Drift also treats the two phases of generation separately, since filling the prompt is limited by arithmetic while producing each further token is limited by memory bandwidth.
Chart: Google. Solid bars are prefill on the left axis, cross-hatched bars decode on the right axis. Gaps mark runs the authors could not measure. A mobile GPU constrains inference in ways the paper works through one at a time. Memory comes first, and the authors report that Stable Diffusion 1.4 would need 4.31 GB for its intermediate 16-bit activations, which handing the same buffers to operations whose lifetimes do not overlap cuts to 387 MB. Weight layout comes second, and rearranging weight tensors into slices of four channels that match the 4-element units a GPU computes with buys up to a 20% speedup in matrix multiplication, which the authors call the primary driver of the framework&amp;amp;rsquo;s advantage. Shader code comes third, and ML Drift writes its kernels once at startup from hand-tuned templates, so the translation between a tensor&amp;amp;rsquo;s logical indices and the buffer or texture that actually holds it costs nothing while tokens are being produced.
The same engine renders Stable Diffusion 1.4 at 512 by 512 pixels over 20 iterations in 10.96 seconds on the S23 Ultra, under 9 seconds on the S24 and in 3.4 seconds on a laptop with an Intel Lunar Lake Ultra 7 258V. On an NVIDIA RTX 4090 the OpenCL path runs 5% to 25% slower than llama.cpp on CUDA, which the authors put down to NVIDIA&amp;amp;rsquo;s Tensor Cores being unreachable through OpenCL and WebGPU. The paper names sub-channel quantisation, sparsity and vendor extensions for matrix multiplication as the next steps, and points to no public release of the framework.
Source: https://arxiv.org/abs/2505.00232
Read the article: https://llmobile.kavents.com/ticker/scaling-on-device-gpu-inference/</description><category>Google</category><category>Qualcomm</category><category>Android</category><category>Benchmarks</category><category>Research</category></item><item><title>Gemma 3 adds a 1B size that fits in 0.5 GB as an int4 checkpoint</title><link>https://llmobile.kavents.com/ticker/gemma-3/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/gemma-3/</guid><pubDate>Wed, 12 Mar 2025 17:00:00 +0100</pubDate><description>Google released Gemma 3 on March 12, 2025 in four sizes, 1B, 4B, 12B and 27B, and describes them as designed to run fast directly on devices, from phones and laptops to workstations. The 1B is a new size for the line, which started at 2B with Gemma and stayed there through Gemma 2. The technical report puts the 1B at 2.0 GB of weights in bfloat16 and at 0.5 GB as a quantised int4 checkpoint, and the 4B at 8.0 GB and 2.6 GB.
Context length is not the same across the four. The report gives 128K tokens for the 4B, 12B and 27B and 32K for the 1B, with the three larger models pre-trained at 32K and stretched to 128K at the end of pre-training by rescaling their rotary position embeddings. To stop the key-value cache from swallowing memory at long context, Gemma 3 places five local attention layers between each global one and holds the local layers to a 1024-token window, where Gemma 2 alternated one for one. Google reports that in an ablation on a 2B model a global-attention-only configuration carries a 60 percent memory overhead from the cache at a 32K prefill, and that interleaving with a 1024-token window brings that under 15 percent.
Multimodality starts at 4B. The 4B, 12B and 27B share one 400M-parameter SigLIP vision encoder, frozen during training, which takes square images resized to 896 by 896 pixels and condenses each one into 256 tokens for the language model to read. An inference-time step Google calls Pan and Scan cuts non-square or high-resolution images into equal crops and resizes each to the encoder&amp;amp;rsquo;s resolution, which the report measures as raising the 4B from 72.8 to 81.0 on DocVQA and from 44.1 to 57.0 on InfoVQA in a 4-shot evaluation of a pre-trained checkpoint. The 1B has no vision encoder and takes text only.
Google published quantisation-aware trained checkpoints next to the raw weights, made by fine-tuning each model for about 5,000 steps against the probabilities of the unquantised model, so the weights are trained to tolerate the lower precision rather than rounded down afterwards. The three formats are per-channel int4, per-block int4 and switched fp8, which Google picked for what open source inference engines such as llama.cpp already support. The report&amp;amp;rsquo;s footprint table counts weights alone and weights plus a key-value cache for a 32,768-token sequence, and on the second measure the 1B needs 1.4 GB at per-channel int4 against 2.9 GB in bfloat16, while the 4B needs 7.3 GB against 12.7 GB.
Google&amp;amp;#39;s own footprint figures for the raw and quantised checkpoints, with and without a key-value cache at 32,768 tokens. Table: Google DeepMind, Gemma 3 technical report, CC BY 4.0. Google reports its own zero-shot figures for the instruction-tuned models and puts Gemma 3 4B at 43.6 on MMLU-Pro and 75.6 on MATH, against 15.6 and 27.2 for Gemma 2 2B and 56.9 and 55.6 for Gemma 2 27B, with the 1B at 14.7 and 48.0. The 4B also scores 48.8 on the MMMU validation set, a benchmark the text-only sizes do not run, and the report calls the instruction-tuned 4B competitive with the instruction-tuned Gemma 2 27B. Google put every size on Hugging Face, Kaggle and Ollama as a pre-trained and an instruction-tuned checkpoint under the Gemma Terms of Use, and names LiteRT, Gemma.cpp, llama.cpp and Google AI Edge among the runtimes that take them. The same release added ShieldGemma 2, a 4B image safety classifier built on Gemma 3 that labels images for dangerous content, sexually explicit material and violence.
▶What&amp;amp;#39;s new in Gemma 3?Loading connects your browser to www.youtube-nocookie.com, which may process your IP address and use cookies.
Load contentOpen externallyVideo: Google for Developers.
Source: https://blog.google/technology/developers/gemma-3/
Read the article: https://llmobile.kavents.com/ticker/gemma-3/</description><category>Google</category><category>Gemma</category><category>Quantisation</category><category>Open weights</category><category>Benchmarks</category></item><item><title>Gemma 2 2B is distilled from a larger model and scores 1126 on Chatbot Arena</title><link>https://llmobile.kavents.com/ticker/gemma-2/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/gemma-2/</guid><pubDate>Wed, 31 Jul 2024 17:00:00 +0200</pubDate><description>Google released Gemma 2 2B on July 31, 2024, an open model of 2.6B parameters that Google says produces outsized results by learning from larger models through distillation. Google reports it at 1126 Elo on the LMSYS Chatbot Arena leaderboard, captured on July 30, 2024, ahead of Mixtral 8x7B Instruct at 1114, GPT 3.5 Turbo 0314 at 1106 and Llama 2 70B chat at 1093, and states that the model surpasses all GPT-3.5 models there.
Chart: Google, from the announcement. Scores are Google&amp;amp;#39;s own reading of the leaderboard on July 30, 2024. The technical report describes training the 2B size on 2 trillion tokens with knowledge distillation, where a bigger teacher model supplies a probability for every token in context and the small student is trained to match those probabilities rather than to predict the next token. In an ablation the authors ran, a 2B model trained over 500B tokens distilled from a 7B teacher averaged 67.7 across three benchmarks, against 60.3 for the same model trained from scratch. The released model holds 2.02B non-embedding parameters plus 590M embedding parameters, which come from the 256k-entry Gemini vocabulary, and it interleaves local sliding-window attention with global attention over an 8192-token context. Google DeepMind trained it on 512 TPUv5e chips.
Google names edge devices and laptops as deployment targets next to cloud serving on Vertex AI and Google Kubernetes Engine, and says the model is small enough for the free tier of T4 GPUs in Google Colab. For running it locally Google points to Ollama and to Gemma.cpp, its own lightweight C++ inference engine, and says support in MediaPipe, the on-device inference stack for Android, iOS and the web, is coming. Google adds that NVIDIA optimised the model with the TensorRT-LLM library for GeForce RTX cards and for Jetson modules at the edge.
Google put the weights on Kaggle, Hugging Face and Vertex AI Model Garden under the Gemma terms, which it calls commercially friendly for research and commercial applications, and made the model available to try in Google AI Studio. The same announcement added ShieldGemma, a set of safety classifiers built on Gemma 2 that filter hate speech, harassment, sexually explicit and dangerous content, and Gemma Scope, over 400 sparse autoencoders for inspecting the inner workings of Gemma 2 2B and 9B.
Source: https://developers.googleblog.com/en/smaller-safer-more-transparent-advancing-responsible-ai-with-gemma/
Read the article: https://llmobile.kavents.com/ticker/gemma-2/</description><category>Google</category><category>Gemma</category><category>Distillation</category><category>Open weights</category><category>Benchmarks</category></item><item><title>Google ships an experimental MediaPipe LLM Inference API for web, Android and iOS</title><link>https://llmobile.kavents.com/ticker/mediapipe-llm-inference-api/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/mediapipe-llm-inference-api/</guid><pubDate>Thu, 07 Mar 2024 17:00:00 +0100</pubDate><description>Google released the MediaPipe LLM Inference API on March 7, 2024, an experimental way to run language models fully on device from web, Android and iOS apps. Four open models are supported at launch, Gemma 2B, Phi 2, Falcon 1B and Stable LM 3B.
The API quantises weights to int8, with Gemma 2B using mixed 4-bit and 8-bit weights. Google measured throughput on unnamed high-end devices with a 1024-token input prompt and a maximum of 1280 tokens. In its charts, Gemma 2B at int4 prefills at roughly 680 tokens per second on WebGPU and on an Android GPU, and decodes at about 57 tokens per second on WebGPU against 31 on an Android GPU and 27 on iOS.
Chart: Google. Chart: Google. Gemma 2B at int4 was the only model that ran on iOS. Google marks the Android version as intended for experimental and research use only, and points production apps to the Gemini API or to Gemini Nano through Android AICore instead. On iOS, Gemma 2B at int4 was the only model the team could run, which Google attributes to the memory available on the platform.
Source: https://developers.googleblog.com/en/large-language-models-on-device-with-mediapipe-and-tensorflow-lite/
Read the article: https://llmobile.kavents.com/ticker/mediapipe-llm-inference-api/</description><category>Google</category><category>MediaPipe</category><category>Developer tools</category><category>Android</category><category>iOS</category></item><item><title>Gemma 2B and 7B open the Gemma line, built on Gemini research</title><link>https://llmobile.kavents.com/ticker/gemma/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/gemma/</guid><pubDate>Wed, 21 Feb 2024 18:00:00 +0100</pubDate><description>Google released Gemma on February 21, 2024, the first two models in the line, at 2B and 7B parameters and in a pretrained and an instruction-tuned checkpoint each. The technical report presents the 7B as a model for deployment on GPU and TPU and the 2B as one for CPU and on-device applications. Google says Gemma was built from the same research and technology used to create Gemini and shares technical and infrastructure components with it.
Both sizes are decoder-only transformers trained on a context length of 8192 tokens. The 2B has 18 layers, a model dimension of 2048 and a single key-value head, since Google&amp;amp;rsquo;s ablations found multi-query attention works well at small scale, while the 7B has 28 layers and keeps standard multi-head attention across 16 heads. Both inherit Gemini&amp;amp;rsquo;s 256k-entry vocabulary, which puts 524M of the 2B&amp;amp;rsquo;s parameters into embeddings and leaves 1.98B elsewhere. Google reports training the 2B on 3T tokens and the 7B on 6T, mostly English web documents, mathematics and code.
Google reports Gemma 7B at 64.3 on MMLU 5-shot against 54.8 for Llama-2 13B, 46.4 on GSM8K against 28.7, and 32.3 on HumanEval against 18.3, for an average of 56.9 across 18 academic benchmarks. Google puts the 2B at 42.3 on MMLU and 45.0 on average, ahead of Llama-2 7B on the mathematics and coding tasks and behind it overall. Google states that Gemma outperforms similarly sized open models on 11 of the 18 text-based tasks, and notes it could not rerun the Llama-2 evaluations itself because of that model&amp;amp;rsquo;s licensing, so it cites Meta&amp;amp;rsquo;s published figures.
Google&amp;amp;#39;s own figures for Gemma 7B against Llama-2. Chart: Google. The weights went up on Kaggle and Hugging Face, with the models also runnable from Colab and Vertex AI, and Google provided toolchains for inference and supervised fine-tuning across JAX, PyTorch and TensorFlow through native Keras 3.0. Hugging Face added Gemma to Transformers 4.38 on announcement day. llama.cpp merged Gemma support the same day, within half an hour of the pull request opening, which is what brought the 2B into quantised local runs on consumer hardware. Google states that the models run across laptop, desktop, IoT, mobile and cloud.
The weights are open but the licence is not a standard open source one. Google publishes them under its own Gemma Terms of Use, which allow use, modification and redistribution provided that downstream recipients get the same terms and the separate Prohibited Use Policy, and which reserve Google&amp;amp;rsquo;s right to restrict uses it considers non-compliant. Google says the terms permit responsible commercial usage and distribution for all organisations regardless of size. Alongside the models Google shipped a Responsible Generative AI Toolkit with a safety classification method, a tool for debugging model behaviour and written guidance for model builders.
Source: https://blog.google/technology/developers/gemma-open-models/
Read the article: https://llmobile.kavents.com/ticker/gemma/</description><category>Google</category><category>Gemma</category><category>Open weights</category><category>llama.cpp</category><category>Benchmarks</category></item><item><title>Galaxy S24 becomes the second phone line to run Gemini Nano</title><link>https://llmobile.kavents.com/ticker/galaxy-s24-gemini-nano/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/galaxy-s24-gemini-nano/</guid><pubDate>Wed, 17 Jan 2024 20:00:00 +0100</pubDate><description>Google announced on January 17, 2024 that the Galaxy S24 series runs Gemini Nano on device, which made it the first phone line outside the Pixel 8 Pro to do so. Google names Magic Compose in Google Messages as the feature that runs locally, and states that the data does not leave the phone.
Image: Google. The rest of the announced features run in the cloud. Google lists Gemini Pro behind summarisation in Samsung Notes and Voice Recorder as well as keyboard features, with Generative Edit in the Gallery app built on Imagen 2. Gemini Ultra was still in testing at the time.
The launch also introduced Circle to Search, a gesture that searches whatever is circled or highlighted on screen without switching apps. Samsung published its own account of the launch on the Samsung Newsroom.
Source: https://blog.google/products/android/google-ai-samsung-galaxy-s24/
Read the article: https://llmobile.kavents.com/ticker/galaxy-s24-gemini-nano/</description><category>Samsung</category><category>Google</category><category>Gemini Nano</category><category>Android</category></item><item><title>Gemini Nano ships on the Pixel 8 Pro and Android gets AICore</title><link>https://llmobile.kavents.com/ticker/gemini-nano-pixel-8-pro/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/gemini-nano-pixel-8-pro/</guid><pubDate>Wed, 06 Dec 2023 18:00:00 +0100</pubDate><description>Google brought Gemini Nano to the Pixel 8 Pro in its December 2023 feature drop, where it powers Summarize in Recorder and Smart Reply in Gboard. Google calls the Pixel 8 Pro the first smartphone engineered for Gemini Nano and runs the model on the Tensor G3.
Your browser does not support this video. This video could not be loaded. Use the link below to open it directly.
Video: Google. Gemini Nano summarising a recording in the Recorder app. Open the Android Developers post According to Google, running the model locally helps prevent sensitive data from leaving the phone and lets the features work without a network connection. Summarize in Recorder launched in English. Smart Reply in Gboard launched globally on the United States English keyboard layout, starting with WhatsApp, Line and KakaoTalk.
On the same day Google introduced AICore, a system service in Android 14 that handles model management, runtimes and safety features for Gemini Nano. It supports Low Rank Adaptation, so developers can build small adapters trained on their own data, and it targets the Google Tensor TPU as well as NPUs from Qualcomm, Samsung and MediaTek. Google describes the service as isolated from the network by design and opened access through an early access programme.
Diagram: Google.
Source: https://blog.google/products/pixel/pixel-feature-drop-december-2023/
Read the article: https://llmobile.kavents.com/ticker/gemini-nano-pixel-8-pro/</description><category>Google</category><category>Gemini Nano</category><category>Pixel</category><category>Android</category><category>Developer tools</category></item><item><title>MobileBERT runs in 62 ms on a Pixel 4 with 25.3M parameters</title><link>https://llmobile.kavents.com/ticker/mobilebert/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/mobilebert/</guid><pubDate>Mon, 06 Apr 2020 22:20:00 +0200</pubDate><description>Researchers at Carnegie Mellon University and Google Brain published MobileBERT on April 6, 2020, a compressed version of the BERT language model built for phones. It has 25.3M parameters against 109M for BERT-base, and the authors measured 62 ms per inference on a Pixel 4, which they report as 4.3 times smaller and 5.5 times faster than BERT-base.
The saving comes from the shape of each layer. MobileBERT keeps the 24 layers of the much larger BERT-large but makes every block narrow, so that a block works internally at a width of 128 while the representation flowing between blocks stays 512 wide, with a small linear layer at each end to shrink the input and widen the output again, an arrangement the paper calls a bottleneck. Narrowing the block leaves the attention module holding too large a share of the parameters, so each block stacks 4 feed-forward networks behind its single attention module to restore the usual balance. The authors also traced a large part of the remaining latency to layer normalisation and the gelu activation and replaced both with cheaper operations, which cut inference from 192 ms to 62 ms without changing the number of arithmetic operations.
A network that deep and thin is hard to train directly, so the team first trained a teacher and then copied its behaviour layer by layer. The teacher is BERT-large fitted with inverted bottlenecks, which widen inside the block but narrow the representation passing between blocks to the same 512 the student uses, so the two models&amp;amp;rsquo; layer outputs line up and can be compared one to one during the transfer. That transfer happens only during pre-training, which keeps the result task-agnostic, so one distilled model is fine-tuned separately for each downstream task and no task-specific teacher is needed.
MobileBERT scores 77.7 on the GLUE language understanding benchmark against 78.3 for BERT-base, and on the SQuAD question answering sets v1.1 and v2.0 it reaches dev F1 scores of 90.0 and 79.2, which the authors put 1.5 and 2.1 above BERT-base. The latency figures are the authors&amp;amp;rsquo; own runs, with the models exported to TensorFlow Lite and timed on a 4-thread Pixel 4 at a fixed sequence length of 128, where BERT-base took 342 ms. A smaller variant with 15.1M parameters runs in 40 ms and scores 75.8 on GLUE, and 8-bit quantisation leaves the reported accuracies almost unchanged. Code and pre-trained weights are published in Google Research&amp;amp;rsquo;s repository.
Source: https://arxiv.org/abs/2004.02984
Read the article: https://llmobile.kavents.com/ticker/mobilebert/</description><category>Google</category><category>Pixel</category><category>Research</category><category>Benchmarks</category></item></channel></rss>