Skip to main content
Use /v1/chat/completions for OpenAI-compatible text generation.

Streaming

Set stream: true to receive server-sent events. Each event is an OpenAI chat.completion.chunk, and the stream ends with data: [DONE].

Parameters

OpenAI-standard parameters are accepted, including temperature, top_p, max_tokens, max_completion_tokens, stop, n, frequency_penalty, presence_penalty, seed, response_format, tools, tool_choice, logprobs, top_logprobs, logit_bias, user, and stream.

Reasoning effort

Set top-level reasoning_effort for models that support configurable reasoning. The Responses-style reasoning.effort is also accepted; if both are provided, they must match or the request returns 400.
Supported values vary by model; omission preserves the existing default. See model-specific reasoning controls. Each model’s accepted efforts, its default and whether reasoning can be disabled are published in GET /v1/models under capability_metadata.reasoning, so clients can configure a model without hardcoding a table. Output limits such as max_completion_tokens, max_output_tokens, and the native equivalents include reasoning and final-answer tokens. Effort is not an output-token limit. A request can still exhaust its budget (finish_reason: "length" in Chat Completions) and return an incomplete or empty final answer.

Kimi K3 compatibility

kimi-k3 uses the same OpenAI-compatible endpoint, but its native contract has model-specific rules:
  • K3 always reasons and cannot be turned off. It accepts low, high and max; the default is max, and a lower effort trades reasoning depth for latency. The K2.x thinking object is accepted only as an enabled compatibility alias.
  • Prefer max_completion_tokens. APIAny accepts deprecated max_tokens and uses max_completion_tokens for Kimi K3. If no limit is provided, the model’s default is used. The explicit maximum is 1,048,576 tokens.
  • K3 sampling settings are fixed. APIAny removes temperature, top_p, frequency_penalty, and presence_penalty only for K3; n must be 1.
  • reasoning_content is the reasoning trace and content is the final answer. Streaming returns them as separate deltas. Do not merge the reasoning trace into the final answer.
  • Replay the complete assistant message, including reasoning_content and tool_calls, in multi-turn and tool-calling requests.
  • The last assistant message may use partial: true. A content-less system message may carry tools to load K3 tools dynamically.
  • K3 image/video content blocks must use a Base64 data: URL or an ms:// file ID. Public HTTP media URLs are rejected before the model request is sent.
If finish_reason is length, the output budget was exhausted. The response is still OpenAI-compatible, but the final answer may be missing or incomplete; increase max_completion_tokens or reduce the input before retrying.

Notes

  • The model is a public APIAny.AI model ID.
  • Token usage is normalized into the OpenAI-style usage object when usage data is available.

GPT-6 output and stream completion

For gpt-6-astra, recognizable leading thinking wrappers are separated from the final answer. Read content / output_text for the answer and the independent reasoning fields or summary events for thinking. Explicit tag examples and ambiguous markup are preserved. Reasoning effort and output-token limits are not changed by this formatting. An HTTP 200 stream is not proof of a successful generation. Handle error events, keep any partial answer clearly marked as incomplete, and wait for the protocol’s completion event. In Chat, read the final usage chunk before [DONE]; in Responses, distinguish response.completed, response.incomplete, and response.failed / error. For GPT-6 Chat streams, [DONE] ends the stream; it does not override an earlier error or guarantee that the answer is complete. Always inspect finish_reason: length means the output limit was reached, and content_filter means content was filtered. Neither is rewritten to stop. If the connection closes before completion or an error event is received, do not treat the partial text as a complete answer.