Kimi API guide · sources captured 2026-08-18

Kimi API pricing

Kimi API pricing starts with the exact model ID, then separates input, output, and cache-hit usage. Thinking behavior and retained conversation context can change the token shape behind those categories. This guide explains the provider mechanics without copying a unit-price snapshot into a production forecast.

Kimi API model inference pricing explanation · verified 2026-08-18 · Kimi K2.7 Code pricing documentation · verified 2026-08-18

How is Kimi API pricing calculated?

Kimi’s chat-completion documentation says both input and output are billed by usage. If a document is extracted and its contents are passed into the model, that extracted text is input usage too. The billing subject is therefore the full request sent to inference, not only the user’s latest message.

Forecast fieldProvider factWhat to measure
Model IDKimi publishes model-specific pricing records.Keep the exact ID on every sample and quote.
InputChat-completion input is billed by usage.Count system prompts, tools, documents, history, and user content.
Cache statusK2.7 Code separates cache-hit and cache-miss input.Retain hit and miss tokens instead of assuming a hit rate.
OutputOutput is a separate usage category.Measure reasoning and final-answer behavior on representative tasks.

Monthly total tokens alone cannot reproduce this calculation. Preserve the model ID and usage categories beside the measurement date, then recheck the provider record before a commercial decision.

Source: Kimi API model inference pricing explanation · verified 2026-08-18 · Source: Kimi K2.7 Code pricing documentation · verified 2026-08-18

Which model ID belongs in a Kimi API pricing comparison?

Use the exact ID that will receive traffic. Kimi’s current model list distinguishes kimi-k3, kimi-k2.7-code, its high-speed variant, andkimi-k2.6. The K2.7 Code pricing record gives the standard and high-speed IDs separate rows, while describing the high-speed option as the same model with a different serving profile.

This is also why one “Kimi K2 API pricing” average is not a durable comparison. The model list says the older kimi-k2 preview series was discontinued on 2026-05-25. Pin a supported ID, record the provider’s lifecycle state, and reopen the evaluation before changing either the model or serving variant.

Source: Kimi API model list · verified 2026-08-18 · Source: Kimi K2.7 Code pricing documentation · verified 2026-08-18

How does automatic context caching affect Kimi cost?

Kimi says context caching is automatically enabled for all model requests. There is no cache ID to create and no TTL to manage. The system looks for repeated initial context such as system prompts, knowledge documents, or tool definitions, so stable prefixes are the relevant unit of reuse.

A cache hit is conditional, not guaranteed. The documentation says a new request can hit the prefix cache only when the previous request’s prompt exceeds 256 tokens; shorter prompts are discarded rather than cached. Keep an uncached baseline, report observed hit and miss tokens separately, and do not apply a headline cache assumption to every request.

Source: Kimi API context caching documentation · verified 2026-08-18

Why do thinking tokens belong in a Kimi forecast?

Kimi documents reasoning in the response’s reasoning_content field and says thinking increases token usage. For kimi-k2.7-code, thinking and Preserved Thinking are always on. Historical assistant messages must retain their reasoning_content in multi-turn conversations.

The output ceiling also covers more than the visible answer: Kimi says the combined tokens in reasoning_content and content must not exceed max_tokens. Measure real task traces, including tool-call loops and retained history, instead of multiplying final-answer length by request count.

Source: Kimi API thinking-model documentation · verified 2026-08-18

How should a team forecast Kimi API usage?

Start with representative request traces. Kimi provides an estimate-token-count endpoint whose input closely matches chat completion; a successful response exposes the estimate in data.total_tokens. Use it before calls, then retain actual response usage for the production distribution.

Bucket the result by exact model ID, cache hit or miss, input size, reasoning behavior, final output, and traffic class. Keep percentiles and peak concurrency beside totals. This turns a provider pricing page into an auditable workload record without freezing a public rate that may change after publication.

Source: Kimi API token estimation documentation · verified 2026-08-18 · Source: Kimi API model inference pricing explanation · verified 2026-08-18

What should a production Kimi API quote include?

Send the exact model and serving variant, representative input and output distributions, cache-hit observations, reasoning and tool-call behavior, expected concurrency and bursts, version policy, and required service terms. Reopen the quote when any of those inputs changes.

Continue with the sourced kimi-k2.7-code model record, the availability ledger, and SteadyGateway’s production pricing inputs. These pages keep model facts, provider availability, and commercial scope as separate decisions.

Source: Kimi K2.7 Code pricing documentation · verified 2026-08-18 · Source: Kimi API thinking-model documentation · verified 2026-08-18

What do buyers ask about Kimi API pricing?

The recurring questions concern token categories, exact model IDs, automatic caching, reasoning usage, estimation, and production quote scope. Each answer uses provider records captured on 2026-08-18.

How is Kimi API pricing calculated?

Kimi bills chat-completion input and output by usage. The applicable model record can also separate cache-hit input from cache-miss input, so a forecast needs the exact model ID, token mix, and observed cache status.

Source: Kimi API model inference pricing explanation · verified 2026-08-18

How does Kimi K2 API pricing differ by model?

Kimi publishes separate records for current model IDs. For Kimi K2.7 Code, the standard and high-speed IDs have distinct pricing rows even though the provider describes them as the same underlying model.

Source: Kimi K2.7 Code pricing documentation · verified 2026-08-18

Does Kimi API cache prompts automatically?

Yes. Kimi says context caching is automatically enabled for all model requests and requires no cache ID or TTL management. A new request can hit the prefix cache only when the previous prompt exceeds 256 tokens.

Source: Kimi API context caching documentation · verified 2026-08-18

Do Kimi reasoning tokens affect usage?

Yes. Kimi says thinking adds token usage, and the reasoning_content and final content tokens together must remain within max_tokens. Preserve this class in workload measurements instead of forecasting only visible answer length.

Source: Kimi API thinking-model documentation · verified 2026-08-18

How can I estimate Kimi input tokens before a call?

Kimi provides an estimate-token-count endpoint whose request shape closely matches chat completion. A successful response returns the estimated total in data.total_tokens.

Source: Kimi API token estimation documentation · verified 2026-08-18

Does SteadyGateway publish Kimi unit prices?

No. SteadyGateway scopes the exact model, token distribution, cache behavior, capacity, version policy, and service terms before returning current production terms in writing.

Source: Kimi K2.7 Code pricing documentation · verified 2026-08-18

Request a workload-scoped Kimi quote.

Send the exact model ID, request distribution, cache observations, reasoning behavior, and capacity plan. We will return production access and service terms in writing.

Request production accesshello@steadygateway.com

99.9% availability commitment with tiered service credits · Reply within one business day · NDA available on request