Benchmarks
52 updates on Benchmarks.
BUPT proposes one 9.2B model in the OS that all apps call through adapters
A BUPT team proposes that the phone OS ship one 9.2B multimodal model all apps call through small adapters, and matched app models on 85% of 50 datasets.
Apple releases OpenELM at 270M to 3B, with parameters spread unevenly across layers
Apple published four models from 270M to 3B parameters with the full training framework and code to run them through MLX on Apple silicon.
Octopus fine-tunes a 2B model to 93 percent on API function calls
Stanford and Harvard authors fine-tuned open 2B to 7B models on 20,000 RapidAPI functions and report up to 97 percent function call accuracy.
MeRino designs sub-100M language models that run 4.9 times faster than OPT-350M
Researchers design 52M to 64M parameter transformers by maximising entropy under a compute budget, matching OPT-350M accuracy on an NVIDIA Jetson Nano.
Stability AI trains Stable LM 2 1.6B on seven languages and 2 trillion tokens
The technical report details a 1.6B model pre-trained on seven languages and measures 127 tok/s for a 4-bit build on an M2 Mac mini.
Meta's MobileLLM trades width for depth and gains 2.7 and 4.3 points
Meta Reality Labs built 125M and 350M models around deep and thin layers and profiled them on an iPhone 13 through ExecuTorch.
TinyLLaVA's 3.1B model outscores 7B LLaVA-1.5 on seven of nine benchmarks
Beihang and Tsinghua researchers report a 3.1B vision-language model that beats the 7B LLaVA-1.5 on seven of nine image benchmarks.
Gemma 2B and 7B open the Gemma line, built on Gemini research
Google released Gemma 2B and 7B with an 8192-token context, weights on Kaggle and Hugging Face under a custom Gemma licence, not an open source one.
MobileVLM V2 runs 1.7B, 3B and 7B vision models on 144 image tokens
Meituan and Zhejiang University report a 1.7B vision language model at 64.2 on six benchmarks and 51.63 tok/s on an NVIDIA Jetson Orin.
Microsoft releases Phi-2, a 2.7B model it says matches models 25 times larger
The 2.7B base model was trained on 1.4 trillion tokens in 14 days on 96 A100 GPUs, and Microsoft says it matches models up to 25 times larger.
Stability AI releases StableLM Zephyr 3B for edge devices
Stability AI tuned a 3B chat model with direct preference optimisation for edge devices and reports an MT-Bench score of 6.64.
Microsoft carries its textbook data recipe from code to reasoning with the 1.3B phi-1.5
The 1.3-billion-parameter model trains on 30B tokens of mostly synthetic data and posts reasoning scores above Llama2-7B in Microsoft evaluations.