Arm

2 updates on Arm.

  1. FlexServe runs on-device LLM inference inside TrustZone for 4.41 percent more latency

    Shanghai Jiao Tong University researchers kept model weights and prompts inside ARM TrustZone at 4.41 percent more time to first token.

  2. Arm runs Llama 2 7B at 9.6 tokens per second on three mobile CPU cores

    Arm demonstrated Llama 2 7B generating 9.6 tokens per second on three Cortex-A700 CPU cores, using int4 kernels it says beat llama.cpp by 20 percent.