Apple

14 updates on Apple.

  1. Apple researchers run models twice the size of available DRAM from flash

    The LLM in a flash paper loads parameters from flash on demand and reports 4 to 5 times faster CPU and 20 to 25 times faster GPU inference.

  2. Apple publishes MLX, where CPU and GPU share arrays without copies

    Apple machine learning research released an array framework for Apple silicon with a unified memory model, lazy evaluation and Swift bindings for iOS.