Kimi K3-256k

472 points · 142 comments on HN · read original →

Points and comments are a snapshot, not live.

Kimi launches a 256k-context model version that consumes half the quota of the 1M flagship.

Kimi Code now offers the K3-256k model, a 256k-context variant of its flagship K3 (1M tokens, 2.8T parameters). The smaller version delivers identical results within its window while consuming roughly half the quota of the full 1M model. It supports image input but not video. Users on Moderato or above can switch via CLI (/model) or VS Code dropdown. The article advises starting a new session when switching models to avoid cache invalidation and extra token consumption. For third-party tools, set the Model ID and manually configure the context window to 1048576 for full 1M access.

What commenters are saying

Commenters mostly agree this is an API-level configuration change, not a retrained model-vLLM and other inference engines already allow capping context windows for cheaper serving. Several note the waitlist for subscriptions, with speculation that demand surged after K3's open-weight launch and hardware export constraints. Some debate whether US export restrictions could affect Chinese open-weight models in Europe, with skepticism about enforcement. A few users compare Anthropic outages to this launch, joking about the timing.