Nvidia

8 updates on Nvidia.

  1. Nemotron-Flash-1B decodes 1.9 times faster than Qwen3-0.6B on an H100

    Nvidia designed a hybrid 1B and 3B model family around measured decoding latency on an H100 rather than parameter count, and released three checkpoints.

  2. Intelligence per watt puts local model coverage at 88.7% of real queries

    Stanford and Together AI measured 20+ local models on 1M real queries and propose accuracy per watt as the metric for local inference.

  3. D2MoE picks a bit-width per token, 1.39 times the throughput at up to 53 percent less memory

    A MobiCom 2025 paper routes every token to an expert and a bit-width, reporting up to 1.39 times the throughput of EdgeMoE at up to 53 percent less memory.

  4. Raspberry Pi 5 beats human reading speed only up to Phi 3.5 mini

    Helsinki and EURECOM researchers measured 11 models on a Raspberry Pi 5 and a Jetson Orin Nano, where memory and battery gave out before speed did.

  5. MobiLLM fine-tunes OPT-1.3B in 4.5 GB on device by moving backpropagation to a server

    Researchers fine-tuned OPT-1.3B on a Jetson Xavier NX in 4.5 GB by keeping the frozen model on the device and the trainable side network on a server.

  6. Apple team finds H100 last on tokens per dollar for models up to 2B

    Seven Apple authors measured tokens per dollar for LLaMA-style models from 100M to 2B and found H100s last at every size, behind cheaper A100s.

  7. MeRino designs sub-100M language models that run 4.9 times faster than OPT-350M

    Researchers design 52M to 64M parameter transformers by maximising entropy under a compute budget, matching OPT-350M accuracy on an NVIDIA Jetson Nano.

  8. MobileVLM V2 runs 1.7B, 3B and 7B vision models on 144 image tokens

    Meituan and Zhejiang University report a 1.7B vision language model at 64.2 on six benchmarks and 51.63 tok/s on an NVIDIA Jetson Orin.