Skip to main content
Use /v1/responses when your client already speaks the OpenAI Responses API format, or when you need Responses-native features such as reasoning controls, web search tools, and official Responses server-sent events.

Reasoning controls

gpt-5.5 and gpt-5.4 accept Responses reasoning controls. Use reasoning.effort to set the reasoning strength:
For gpt-5.5 and gpt-5.4, supported efforts are none, low, medium, high, and xhigh—not minimal or max. Top-level reasoning_effort is accepted as an alias; if both forms are present, their values must match or the request returns 400. Supported values vary by model; omission preserves the existing default. See model-specific reasoning controls. Each model’s accepted efforts, its default and whether reasoning can be disabled are published in GET /v1/models under capability_metadata.reasoning, so clients can configure a model without hardcoding a table. Output limits such as max_completion_tokens, max_output_tokens, and the native equivalents include reasoning and final-answer tokens. Effort is not an output-token limit. A request can still exhaust its budget (finish_reason: "length" in Chat Completions) and return an incomplete or empty final answer. Pass OpenAI Responses tools directly. For web-connected answers, include a web search tool in tools:
APIAny.AI forwards Responses-native tools to compatible models that support them. Function tools are also supported.

Streaming

Set stream: true to receive the Responses SSE event stream. Events follow the OpenAI Responses format, such as response.created, response.output_text.delta, tool events, and response.completed.

Notes

  • The model is a public APIAny.AI model ID.
  • max_output_tokens is the Responses-style output token limit.
  • Responses-native fields such as reasoning, text, tools, tool_choice, include, previous_response_id, and store are forwarded when supported by the selected model.

GPT-6 output and stream completion

For gpt-6-astra, recognizable leading thinking wrappers are separated from the final answer. Read content / output_text for the answer and the independent reasoning fields or summary events for thinking. Explicit tag examples and ambiguous markup are preserved. Reasoning effort and output-token limits are not changed by this formatting. An HTTP 200 stream is not proof of a successful generation. Handle error events, keep any partial answer clearly marked as incomplete, and wait for the protocol’s completion event. In Chat, read the final usage chunk before [DONE]; in Responses, distinguish response.completed, response.incomplete, and response.failed / error.