/v1/responses when your client already speaks the OpenAI Responses API
format, or when you need Responses-native features such as reasoning controls,
web search tools, and official Responses server-sent events.
Reasoning controls
gpt-5.5 and gpt-5.4 accept Responses reasoning controls. Use
reasoning.effort to set the reasoning strength:
gpt-5.5 and gpt-5.4, supported efforts are none, low, medium, high, and xhigh—not minimal or max. Top-level reasoning_effort is accepted as an alias; if both forms are present, their values must match or the request returns 400.
Supported values vary by model; omission preserves the existing default. See model-specific reasoning controls.
Each model’s accepted efforts, its default and whether reasoning can be disabled are published in GET /v1/models under capability_metadata.reasoning, so clients can configure a model without hardcoding a table.
Output limits such as max_completion_tokens, max_output_tokens, and the native equivalents include reasoning and final-answer tokens. Effort is not an output-token limit. A request can still exhaust its budget (finish_reason: "length" in Chat Completions) and return an incomplete or empty final answer.
Web search
Pass OpenAI Responses tools directly. For web-connected answers, include a web search tool intools:
Streaming
Setstream: true to receive the Responses SSE event stream. Events follow
the OpenAI Responses format, such as response.created,
response.output_text.delta, tool events, and response.completed.
Notes
- The
modelis a public APIAny.AI model ID. max_output_tokensis the Responses-style output token limit.- Responses-native fields such as
reasoning,text,tools,tool_choice,include,previous_response_id, andstoreare forwarded when supported by the selected model.
GPT-6 output and stream completion
Forgpt-6-astra, recognizable leading thinking wrappers are separated from the final answer. Read content / output_text for the answer and the independent reasoning fields or summary events for thinking. Explicit tag examples and ambiguous markup are preserved. Reasoning effort and output-token limits are not changed by this formatting.
An HTTP 200 stream is not proof of a successful generation. Handle error events, keep any partial answer clearly marked as incomplete, and wait for the protocol’s completion event. In Chat, read the final usage chunk before [DONE]; in Responses, distinguish response.completed, response.incomplete, and response.failed / error.