FareedKhan-dev/kimi-k3-in-c
Fareed Khan has built a portable C99 implementation of the Kimi K3 model, which can run inference on a single CPU with 8.24 GB of RAM. This implementation does not require any frameworks or GPU acceleration.
This project demonstrates the feasibility of running large language models on resource-constrained hardware, making AI more accessible to a wider range of users.