Posts Tagged "inference"
Kimi K3 from day zero: the largest model we operate, without a partner deal
Kimi K3 is 2.8 trillion parameters, the largest and costliest model we have ever served. It was live on umans.ai the day the weights dropped — no seat in Moonshot's partner program, no revenue share, just a public download and a serving stack that was already warm. Then we spent a week running five topologies and two engines against each other. The results: B300 wins everywhere, disaggregation earns its complexity, and draft models are not worth it yet.
Read Post
DeepSeek V4 Pro DSpark: the model isn't ready, the architecture is
We served DeepSeek V4 Pro for two days. What we wanted was hands on the architecture: attention that scales almost linearly instead of quadratically, and a speculative decoder that adapts to load. Both delivered. The preview checkpoint did not.
Read Post
GLM-5.2 NVFP4: fast, cheap, and not worth serving
We put an NVFP4 build of GLM-5.2 in front of real users for four days, at 200+ tokens per second. Users called it intoxicating. We retired it anyway: if the tokens are not useful, they are not worth serving, no matter how efficiently we can produce them.
Read Post
GLM-5 vs Kimi-K2.5: Long-context serving at scale
Why GLM-5 is a real step forward compared to GLM-4.7 for serving long-context coding agents, and why we still keep Kimi-K2.5 as the default for the best experience.
Read Post