umans/status/umans-kimi-k3
Live · refreshes every 30s
← all models
Umans Kimi K3 Experimental
umans-kimi-k3 · Kimi K3 · Moonshot
In testing
60.3tok/s
throughput · p50 · last 5 min
2.24s
TTFT · p50 · last 5 min
98.99%
uptime · 24h

Kimi K3 in prerelease: Moonshot's largest open-weight release, a 2.8T-parameter mixture-of-experts with a 1M-token context window and native vision, built for repository-scale code understanding and multi-step agentic work. It thinks by default at maximum reasoning effort; select none, low, high, or max to trade depth for speed. While we scale capacity, access is seat-gated through the Labs page, and availability is limited: expect occasional errors during the ramp. For production work today we recommend umans-coder or umans-glm-5.2.

90 days agoin production since Jul 31, 2026today
Context
1049K
Max output
131K
Recommended
131K
Vision
Yes
Tools
Yes
Reasoning
Toggle · none/low/high/max
Trends

Speed over the last 90 days

daily medians · dashed line = target
throughput p50 · output tokens per second, higher is better
now 48.1 tok/s
90 days agopre-release before Jul 31, 2026today
TTFT p50 · time to first token, lower is better
best 1.46s · Jul 28now 3.17s
90 days agopre-release before Jul 31, 2026today
Changelog

Events for Umans Kimi K3

incl. gateway-wide announcements
Jul 312026
Released to production: Umans Kimi K3 Released
umans-kimi-k3 graduated from its Labs prerelease to the production lineup: Moonshot's largest open model, a 1M context window, native vision, and max reasoning effort by default, billed per token at $3.00 / $15.00 / $0.30 per 1M (input / output / cache read).