Posts Tagged "kimi"

Kimi K3 from day zero: the largest model we operate, without a partner deal

Kimi K3 is 2.8 trillion parameters, the largest and costliest model we have ever served. It was live on umans.ai the day the weights dropped — no seat in Moonshot's partner program, no revenue share, just a public download and a serving stack that was already warm. Then we spent a week running five topologies and two engines against each other. The results: B300 wins everywhere, disaggregation earns its complexity, and draft models are not worth it yet.

GLM-5.2 NVFP4: fast, cheap, and not worth serving

We put an NVFP4 build of GLM-5.2 in front of real users for four days, at 200+ tokens per second. Users called it intoxicating. We retired it anyway: if the tokens are not useful, they are not worth serving, no matter how efficiently we can produce them.

GLM-5 vs Kimi-K2.5: Long-context serving at scale

Why GLM-5 is a real step forward compared to GLM-4.7 for serving long-context coding agents, and why we still keep Kimi-K2.5 as the default for the best experience.