§ models · storyline
Smaller, faster, safer: running Kimi and GLM at scale
Running Kimi and GLM at scale uses quantized KV caches, compressed weights, and integrity checks to reduce costs and latency.
Serving frontier models like Kimi and GLM means fighting for GPU memory. Here's how we quantize KV caches, compress model weights, and add integrity checks to serve them faster, cheaper, and safely.
§ sources1 publication · timeline below
- blog.cloudflare.comSmaller, faster, safer: running Kimi and GLM at scaleprimary