Self-Hosting & Open SourceShimmy: Pure-Rust WebGPU Local LLM Inference Engine Without Python or C++ Bloat
Shimmy is a standalone, lightweight LLM inference server written in pure Rust that leverages WebGPU compute shaders to run GGUF models across Vulkan, Metal, and DirectX 12 hardware with zero CUDA or Python dependencies.


