GPU

2 updates on GPU.

  1. S2-MoE speeds up MoE decoding by up to 5.3x in llama.cpp

    S2-MoE is a self-speculative decoding method for mixture-of-experts models that reaches up to 5.3x speedup and about 2.0x on average.

  2. MobiBench paper benchmarks small LLMs in llama.cpp on two laptops and an RTX 3060

    MobiBench is a llama.cpp benchmark suite for speed, memory and accuracy of small LLMs, but the reported tests ran on two laptops and not on phones.