/v1/chat/completions for OpenAI-compatible text generation.
Streaming
Setstream: true to receive server-sent events. Each event is an OpenAI
chat.completion.chunk, and the stream ends with data: [DONE].
Parameters
OpenAI-standard parameters are accepted, includingtemperature, top_p,
max_tokens, max_completion_tokens, stop, n, frequency_penalty, presence_penalty, seed,
response_format, tools, tool_choice, logprobs, top_logprobs,
logit_bias, user, and stream.
Reasoning effort
Set top-levelreasoning_effort for models that support configurable reasoning. The Responses-style reasoning.effort is also accepted; if both are provided, they must match or the request returns 400.
GET /v1/models under capability_metadata.reasoning, so clients can configure a model without hardcoding a table.
Output limits such as max_completion_tokens, max_output_tokens, and the native equivalents include reasoning and final-answer tokens. Effort is not an output-token limit. A request can still exhaust its budget (finish_reason: "length" in Chat Completions) and return an incomplete or empty final answer.
Kimi K3 compatibility
kimi-k3 uses the same OpenAI-compatible endpoint, but its native contract has
model-specific rules:
- K3 always reasons and cannot be turned off. It accepts
low,highandmax; the default ismax, and a lower effort trades reasoning depth for latency. The K2.xthinkingobject is accepted only as an enabled compatibility alias. - Prefer
max_completion_tokens. APIAny accepts deprecatedmax_tokensand usesmax_completion_tokensfor Kimi K3. If no limit is provided, the model’s default is used. The explicit maximum is 1,048,576 tokens. - K3 sampling settings are fixed. APIAny removes
temperature,top_p,frequency_penalty, andpresence_penaltyonly for K3;nmust be1. reasoning_contentis the reasoning trace andcontentis the final answer. Streaming returns them as separate deltas. Do not merge the reasoning trace into the final answer.- Replay the complete assistant message, including
reasoning_contentandtool_calls, in multi-turn and tool-calling requests. - The last assistant message may use
partial: true. A content-less system message may carrytoolsto load K3 tools dynamically. - K3 image/video content blocks must use a Base64
data:URL or anms://file ID. Public HTTP media URLs are rejected before the model request is sent.
finish_reason is length, the output budget was exhausted. The response
is still OpenAI-compatible, but the final answer may be missing or incomplete;
increase max_completion_tokens or reduce the input before retrying.
Notes
- The
modelis a public APIAny.AI model ID. - Token usage is normalized into the OpenAI-style
usageobject when usage data is available.
GPT-6 output and stream completion
Forgpt-6-astra, recognizable leading thinking wrappers are separated from the final answer. Read content / output_text for the answer and the independent reasoning fields or summary events for thinking. Explicit tag examples and ambiguous markup are preserved. Reasoning effort and output-token limits are not changed by this formatting.
An HTTP 200 stream is not proof of a successful generation. Handle error events, keep any partial answer clearly marked as incomplete, and wait for the protocol’s completion event. In Chat, read the final usage chunk before [DONE]; in Responses, distinguish response.completed, response.incomplete, and response.failed / error.
For GPT-6 Chat streams, [DONE] ends the stream; it does not override an earlier error or guarantee that the answer is complete. Always inspect finish_reason: length means the output limit was reached, and content_filter means content was filtered. Neither is rewritten to stop. If the connection closes before completion or an error event is received, do not treat the partial text as a complete answer.