WASTE Inference Engine Enables Kimi K3 2.78T Model to Run Locally on 64GB Laptop
Source:
github.com
The open-source inference engine WASTE enables running the full Kimi K3 model on a MacBook Pro equipped with 64GB memory. This solution reduces minimum memory usage to approximately 29.05 GiB by streaming expert weights on demand, enabling local inference without quantization.