<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Pixel · LLMobile.news</title><link>https://llmobile.kavents.com/tags/pixel/</link><description>A concise news ticker covering AI on mobile devices, local models, apps and hardware.</description><language>en-GB</language><atom:link href="https://llmobile.kavents.com/tags/pixel/index.xml" rel="self" type="application/rss+xml"/><item><title>Google lists first Gemini Nano 4 phones, requires Nano 3 for Gemini Intelligence</title><link>https://llmobile.kavents.com/ticker/gemini-nano-4-first-devices/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/gemini-nano-4-first-devices/</guid><pubDate>Thu, 27 Aug 2026 16:00:00 +0200</pubDate><description>Google&amp;amp;rsquo;s ML Kit GenAI documentation now lists the first devices running nano-v4, 9to5Google reports. The list covers the Pixel 11, Pixel 11 Pro, Pixel 11 Pro XL and Pixel 11 Pro Fold, plus Samsung&amp;amp;rsquo;s Galaxy Z Flip8, Galaxy Z Fold8 and Galaxy Z Fold8 Ultra.
The same documentation sets Nano v3 or greater as the requirement for Gemini Intelligence, Google&amp;amp;rsquo;s on-device feature set. According to the report, that requirement first appeared in May 2026, was removed, and has now been reinstated. Listed hardware requirements include 12 GB or more of RAM, a qualified flagship system-on-chip, five or more OS upgrades and six years of security support.
Gemini Intelligence features named in the report include Rambler and Proactive Assistance on Pixel 11, and task automation across more than 40 apps on Samsung&amp;amp;rsquo;s foldables.
Source: https://9to5google.com/2026/08/27/gemini-intelligence-nano-4/
Read the article: https://llmobile.kavents.com/ticker/gemini-nano-4-first-devices/</description><category>Google</category><category>Android</category><category>Pixel</category><category>Samsung</category><category>Gemini Nano</category></item><item><title>Pixel Watch 5 adds offline Gemini commands and faster on-device smart replies</title><link>https://llmobile.kavents.com/ticker/pixel-watch-5-offline-gemini/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/pixel-watch-5-offline-gemini/</guid><pubDate>Wed, 12 Aug 2026 19:30:00 +0200</pubDate><description>Gemini Intelligence is coming to the Pixel Watch 5, 9to5Google reports. According to the report, offline Gemini commands on the watch use a separate on-device model when the phone or an internet connection is unavailable. Those commands cover timers and alarms, brightness and modes, music control, opening apps and starting workouts.
On-device smart replies now offer three responses instead of one, which the report attributes to a Gemini Nano upgrade that makes them 50 percent faster.
Other parts of the feature set depend on a connection. Proactive Suggestions, formerly Magic Cue, are generated on a paired Pixel 11 and bridged to the watch, and Personal Intelligence draws on Gmail, Calendar and Keep. The update also brings a new At a Glance space on the watch face for timers, workouts, music, flight details and navigation.
Source: https://9to5google.com/2026/08/12/pixel-watch-5-gemini-intelligence/
Read the article: https://llmobile.kavents.com/ticker/pixel-watch-5-offline-gemini/</description><category>Google</category><category>Pixel</category><category>Wearables</category><category>Gemini Nano</category></item><item><title>Pixel 11 series: Tensor G6 adds 50 percent more TPU compute</title><link>https://llmobile.kavents.com/ticker/pixel-11-tensor-g6/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/pixel-11-tensor-g6/</guid><pubDate>Wed, 12 Aug 2026 19:00:00 +0200</pubDate><description>Google has announced the Pixel 11, Pixel 11 Pro and Pixel 11 Pro XL, built around the Tensor G6 chip. Google states that Tensor G6 packs 50 percent more TPU compute and, paired with the latest Gemini Nano model, processes on-device AI tasks up to 3.5 times faster while using up to 3.5 times less energy.
The company also cites an upgraded CPU with 25 percent faster web browsing and 15 percent quicker app launches, and says the chip powers the 30x Super Zoom on the 5x telephoto lens. Google does not publish RAM figures, model sizes or per-task latency in the announcement.
Pre-orders opened on 12 August, with retail availability from 20 August.
▶Meet Google Pixel 11 ProLoading connects your browser to www.youtube-nocookie.com, which may process your IP address and use cookies.
Load contentOpen externallyVideo: Made by Google.
Source: https://blog.google/products-and-platforms/devices/pixel/google-pixel-11-pro-xl/
Read the article: https://llmobile.kavents.com/ticker/pixel-11-tensor-g6/</description><category>Google</category><category>Pixel</category><category>Chips</category><category>NPU</category><category>Gemini Nano</category></item><item><title>ClawMobile tries system commands before screen taps and finishes all six test tasks</title><link>https://llmobile.kavents.com/ticker/clawmobile/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/clawmobile/</guid><pubDate>Thu, 26 Feb 2026 13:34:00 +0100</pubDate><description>Seven researchers at MBZUAI, City University of Hong Kong and one working independently published ClawMobile on February 26, 2026, an agent runtime that runs on the Android phone itself with no tethering to a host machine. The model is not on the device. The authors state that ClawMobile runs locally while model inference is performed remotely, and all agents in their tests use GPT-5.2, so the split is an on-device runtime driving a cloud model. Across six real-life phone tasks the authors report ClawMobile completing all six at 100 percent, against 33 to 100 percent for the UI agent they compare against.
What separates it from screen-reading GUI agents is that the reasoning loop never touches the screen. The orchestrator issues tool calls to control backends instead, and the authors describe a deterministic-first policy in which the runtime first checks its memory for a structured interface, reaching for ADB system commands and the Termux hardware API before anything visual. A semantic UI agent is invoked only when a step genuinely depends on reading dynamic interface state, and raw taps on interface elements sit below that as a fallback. After every action the runtime re-queries device state to confirm the step landed, which the authors contrast with inferring execution success purely through model reasoning.
Architecture diagram: Du et al., CC BY 4.0. The tests ran on a Google Pixel 9 on Android 16, with DroidRun as the UI agent baseline and a stripped ClawMobile without DroidRun as a second baseline, all three on GPT-5.2 and scored by human annotators. In the authors&amp;amp;rsquo; table DroidRun reaches 100 percent on switching the system to dark theme and on playing a YouTube video, 85 percent on posting a YouTube comment, 73 percent on a Chrome search and on a cross-app task that searches football results and writes them into Notes, and 33 percent on installing an app from the Play Store, where ClawMobile scores 100 percent throughout. The authors put the cost of that at 57.5 seconds slower per task on average, with ClawMobile taking 21 seconds for the dark theme switch and 235 seconds for the YouTube comment. Their deterministic-only variant hits the 600-second timeout on both YouTube tasks.
The authors name the remote model as the main open problem, writing that fully remote inference introduces network latency and privacy concerns and calling on-device inference a promising direction they have not taken. They also flag that serialising full UI trees or screenshots into model context can dominate token use and latency on mobile, and present the results as preliminary, six tasks on one device. The code is open source on GitHub, built on the OpenClaw agent framework with a Telegram bot as the chat channel, and the paper was accepted at EuroMLSys 2026 under a Creative Commons Attribution 4.0 license.
Source: https://arxiv.org/abs/2602.22942
Read the article: https://llmobile.kavents.com/ticker/clawmobile/</description><category>Agents</category><category>Android</category><category>Pixel</category><category>Research</category><category>Open source</category></item><item><title>lm-Meter times on-device inference and finds prefill, not decode, is the bottleneck</title><link>https://llmobile.kavents.com/ticker/lm-meter/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/lm-meter/</guid><pubDate>Tue, 07 Oct 2025 19:05:00 +0200</pubDate><description>Researchers at Georgia State University and Toyota InfoTech Labs published lm-Meter, a latency profiler that runs inside the inference engine on the phone and splits each generation into embedding, prefill, decode, softmax and sampling. Measuring the Pythia models on a Google Pixel 8 Pro, they report that scaling from 70M to 1.4B parameters raises prefill latency from 0.012 s to 1.9 s per input token, a 158x slowdown, while decode latency per output token grows from 0.015 s to 0.15 s, a 10x slowdown. The authors write that this inverts the server picture, where decode is usually the limiting phase for single-request inference.
Below the phase level the profiler times individual GPU kernels through OpenCL event timestamps, which give queue, submit, start and end times without access to the closed-source driver. Running a 4-bit quantised Gemma-2-2B-it on a Pixel 8 Pro, the authors report that fused matrix-multiplication kernels dominate a decode step and that the GPU sits idle for more than 21% of it, the second-largest contributor to the step, which they attribute to host-side data preparation and I/O stalls. The paged attention kernel that scans the growing key-value cache is the only one whose cost rises with position in the sequence, climbing from roughly 0.2 ms to about 0.8 ms per token over 250 decode steps, and idle time drops from about 21% to 12% when the model generates 256 tokens instead of 16.
Whether those measurements mean anything depends on what the profiler itself costs. lm-Meter sits in the MLC LLM runtime and TVM in about 3,500 lines of code and needs no host machine attached, and under the Powersave CPU governor, the most constrained setting they tested, the authors measure a throughput loss of 2.58% in prefill and 0.99% in decode. They put the same figures for MELTing Point, the on-device profiler they compare against, at 22% for prefill and more than 93% for decode. Checked against traces from Android GPU Inspector, they report end-to-end phase accuracy of at least 99.99% and mean kernel-level accuracy of 96.82% on the Pixel 8 Pro and 96.61% on a Pixel 7.
The code is on GitHub under the MIT license, with the MLC LLM path released for Android GPUs through OpenCL and support for llama.cpp, vLLM, iOS Metal and Nvidia Jetson listed as unfinished. The work was accepted to the ACM/IEEE Symposium on Edge Computing 2025 and funded by Toyota Motor North America. The measurements come from three phones, the Pixel 8 Pro, Pixel 7 and Pixel 6, and the authors state that other edge platforms such as Jetson boards and Intel NPUs may show different bottlenecks.
Source: https://arxiv.org/abs/2510.06126
Read the article: https://llmobile.kavents.com/ticker/lm-meter/</description><category>Research</category><category>Benchmarks</category><category>Developer tools</category><category>MLC LLM</category><category>Pixel</category></item><item><title>BUPT proposes one 9.2B model in the OS that all apps call through adapters</title><link>https://llmobile.kavents.com/ticker/mobile-foundation-model-as-firmware/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/mobile-foundation-model-as-firmware/</guid><pubDate>Wed, 29 May 2024 15:00:00 +0200</pubDate><description>Researchers at Beijing University of Posts and Telecommunications proposed that a phone ship one shared multimodal model instead of letting every app bundle its own, in a paper published in the ACM MobiCom 2024 proceedings on May 29, 2024. The operating system and the hardware co-manage that model like firmware, unchangeable by apps or by the OS itself, exposed to applications as a system service, and each app reaches it through a small adapter fine-tuned offline for its own task. Their prototype, called M4, holds 9.2B parameters and needs 7.5 GB of peak memory, and the authors report it reaching accuracy comparable to purpose-built models on 85% of the 50 datasets in a benchmark they assembled from 38 mobile AI tasks across five input types.
What the shared model replaces is one small model per app per task. The paper&amp;amp;rsquo;s baselines are 50 task-specific models of 1M to 500M parameters each, one per dataset, against which M4&amp;amp;rsquo;s adapters run from 1,000 to 10 million parameters, so each added task costs under 10 MB. Measured on an Nvidia Jetson Orin NX, 4-bit M4 needs 6.1 GB of storage to serve all 50 tasks against 15.2 GB for the 50 separate models, with the crossover at about 15 tasks, and 7.5 GB of peak memory against roughly five times that. The authors state that on a device with 12 GB of memory the 4-bit model plus all 50 adapters fits, where only 20 of the 50 task-specific models would.
The prototype is slower than the models it replaces. On the Jetson Orin NX with 16 GB, the authors measured M4 averaging 18 times the inference latency of the task-specific models across the 50 tasks and 19 times the energy, 3.6 s against 0.2 s. On a Pixel 7 Pro CPU they measured an average of 6.8 s against 0.54 s, and their per-task breakdown puts image classification at 2.10 s and question answering at 6.34 s to the first token and 0.24 s per token after it. They state that M4 cannot currently run on a stock smartphone GPU or NPU at all, because those processors lack support for the operators it uses.
The NPU numbers in the paper are a projection rather than a measurement. The authors estimate that M4 on an NPU would average 0.48 s and 1.3 J, under the 0.54 s and 2.9 J they measured for task-specific models on the Pixel 7 Pro CPU, but they derive that by applying the CPU-to-NPU ratio they observed for task-specific models, not by running M4 on an NPU. Their case for a simpler accelerator rests on a separate Pixel 7 Pro measurement, where they converted 110 downloaded models to TensorFlow Lite and only 8% ran entirely on the NPU, those gaining a median speedup above 20 times over the CPU. M4 itself uses 39 operator types against the 156 that the 50 task-specific models need between them.
The authors name their own limits. They write that the accuracy results come from an A100 and the Jetson board rather than from phones, that M4 underperforms task-specific models on some tasks including translation, and that a prototype assembled from off-the-shelf pre-trained models is &amp;amp;ldquo;still highly inefficient in terms of accuracy and model parameter size&amp;amp;rdquo;. Its backbone is Meta&amp;amp;rsquo;s LLaMA-7B at 8-bit, with encoders taken from ImageBind and Whisper, and they note that adapters trained against one backbone stop working when the backbone is upgraded, so the design still needs a stable interface between the two. Code and benchmark are published at github.com/UbiquitousLearning/MobileFM, and the paper carries ACM copyright rather than an open license.
Source: https://dl.acm.org/doi/10.1145/3636534.3649361
Read the article: https://llmobile.kavents.com/ticker/mobile-foundation-model-as-firmware/</description><category>Research</category><category>Benchmarks</category><category>NPU</category><category>Pixel</category><category>Llama</category></item><item><title>Gemini Nano ships on the Pixel 8 Pro and Android gets AICore</title><link>https://llmobile.kavents.com/ticker/gemini-nano-pixel-8-pro/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/gemini-nano-pixel-8-pro/</guid><pubDate>Wed, 06 Dec 2023 18:00:00 +0100</pubDate><description>Google brought Gemini Nano to the Pixel 8 Pro in its December 2023 feature drop, where it powers Summarize in Recorder and Smart Reply in Gboard. Google calls the Pixel 8 Pro the first smartphone engineered for Gemini Nano and runs the model on the Tensor G3.
Your browser does not support this video. This video could not be loaded. Use the link below to open it directly.
Video: Google. Gemini Nano summarising a recording in the Recorder app. Open the Android Developers post According to Google, running the model locally helps prevent sensitive data from leaving the phone and lets the features work without a network connection. Summarize in Recorder launched in English. Smart Reply in Gboard launched globally on the United States English keyboard layout, starting with WhatsApp, Line and KakaoTalk.
On the same day Google introduced AICore, a system service in Android 14 that handles model management, runtimes and safety features for Gemini Nano. It supports Low Rank Adaptation, so developers can build small adapters trained on their own data, and it targets the Google Tensor TPU as well as NPUs from Qualcomm, Samsung and MediaTek. Google describes the service as isolated from the network by design and opened access through an early access programme.
Diagram: Google.
Source: https://blog.google/products/pixel/pixel-feature-drop-december-2023/
Read the article: https://llmobile.kavents.com/ticker/gemini-nano-pixel-8-pro/</description><category>Google</category><category>Gemini Nano</category><category>Pixel</category><category>Android</category><category>Developer tools</category></item><item><title>MobileBERT runs in 62 ms on a Pixel 4 with 25.3M parameters</title><link>https://llmobile.kavents.com/ticker/mobilebert/</link><guid isPermaLink="true">https://llmobile.kavents.com/ticker/mobilebert/</guid><pubDate>Mon, 06 Apr 2020 22:20:00 +0200</pubDate><description>Researchers at Carnegie Mellon University and Google Brain published MobileBERT on April 6, 2020, a compressed version of the BERT language model built for phones. It has 25.3M parameters against 109M for BERT-base, and the authors measured 62 ms per inference on a Pixel 4, which they report as 4.3 times smaller and 5.5 times faster than BERT-base.
The saving comes from the shape of each layer. MobileBERT keeps the 24 layers of the much larger BERT-large but makes every block narrow, so that a block works internally at a width of 128 while the representation flowing between blocks stays 512 wide, with a small linear layer at each end to shrink the input and widen the output again, an arrangement the paper calls a bottleneck. Narrowing the block leaves the attention module holding too large a share of the parameters, so each block stacks 4 feed-forward networks behind its single attention module to restore the usual balance. The authors also traced a large part of the remaining latency to layer normalisation and the gelu activation and replaced both with cheaper operations, which cut inference from 192 ms to 62 ms without changing the number of arithmetic operations.
A network that deep and thin is hard to train directly, so the team first trained a teacher and then copied its behaviour layer by layer. The teacher is BERT-large fitted with inverted bottlenecks, which widen inside the block but narrow the representation passing between blocks to the same 512 the student uses, so the two models&amp;amp;rsquo; layer outputs line up and can be compared one to one during the transfer. That transfer happens only during pre-training, which keeps the result task-agnostic, so one distilled model is fine-tuned separately for each downstream task and no task-specific teacher is needed.
MobileBERT scores 77.7 on the GLUE language understanding benchmark against 78.3 for BERT-base, and on the SQuAD question answering sets v1.1 and v2.0 it reaches dev F1 scores of 90.0 and 79.2, which the authors put 1.5 and 2.1 above BERT-base. The latency figures are the authors&amp;amp;rsquo; own runs, with the models exported to TensorFlow Lite and timed on a 4-thread Pixel 4 at a fixed sequence length of 128, where BERT-base took 342 ms. A smaller variant with 15.1M parameters runs in 40 ms and scores 75.8 on GLUE, and 8-bit quantisation leaves the reported accuracies almost unchanged. Code and pre-trained weights are published in Google Research&amp;amp;rsquo;s repository.
Source: https://arxiv.org/abs/2004.02984
Read the article: https://llmobile.kavents.com/ticker/mobilebert/</description><category>Google</category><category>Pixel</category><category>Research</category><category>Benchmarks</category></item></channel></rss>