GPU
2 updates on GPU.
S2-MoE speeds up MoE decoding by up to 5.3x in llama.cpp
S2-MoE is a self-speculative decoding method for mixture-of-experts models that reaches up to 5.3x speedup and about 2.0x on average.
MobiBench paper benchmarks small LLMs in llama.cpp on two laptops and an RTX 3060
MobiBench is a llama.cpp benchmark suite for speed, memory and accuracy of small LLMs, but the reported tests ran on two laptops and not on phones.