Skip to main content
Chat models can be exposed through:
  • /v1/chat/completions
  • /v1/responses
  • /v1/messages
  • /v1beta/models/{model}:generateContent
Pricing can include:
  • Input tokens.
  • Output tokens.
  • Cached input tokens.
  • Minimum request credits.
Context window and maximum output limits are model-level capability metadata.

Available models

Grok 4.7

Use the independent model ID grok-4.7 with /v1/chat/completions. Grok 4.6 remains a separate model. APIAny input and cached-input pricing is 1.2 credits per 1,000 tokens; output pricing is 2.4 credits per 1,000 tokens, matching Grok 4.6 at launch. Credit-pack discounts continue to apply.
This integration initially advertises text input and text output only. Vision, document input, context-window limits and reasoning controls are not verified. Check /v1/models for current availability and pricing. Launch testing observed refusals for some harmless exact-phrase echo requests, despite successful ordinary text and code responses. Do not rely on exact-string echo behavior as an availability probe. This is an observed model limitation, not a guarantee that every instruction will be followed.

Reasoning effort

DeepSeek Flash and GLM 5.3 Flash

deepseek-v4-flash and deepseek-v4.1-flash are separate models, not aliases. V4 Flash accepts text; V4.1 Flash accepts text and images. Both support up to 393,216 output tokens within a 1,048,576-token context window. GLM 5.3 Flash supports the same context window and up to 131,072 output tokens. Use /v1/models for current availability and prices. glm-5.2 is no longer available for new requests; previous usage records remain available. DeepSeek Flash accepts none, low, high, and max. Compatibility values minimal map to low; medium and xhigh map to high. In Chat Completions, thinking: { "type": "disabled" } or effort none disables thinking. Do not combine contradictory settings. For deepseek-v4.1-flash Chat Completions, thinking is enabled by default. Forced tools (tool_choice: "required" or a named function) require reasoning_effort: "none" or thinking: { "type": "disabled" }. Otherwise the request returns HTTP 400 with invalid_tool_choice. To keep thinking enabled, use tool_choice: "auto"; APIAny.AI does not silently disable thinking for you. GLM 5.3 Flash always thinks. Its effort values are low, high, and max (model default); disabling thinking returns 400. For a quick smoke test, use low. Allocate enough tokens for thinking plus the final answer, for example max_tokens: 32768. Explicit limits are honored, not silently increased. max_completion_tokens is accepted as an alias; different simultaneous limits return 400. When continuing a tool conversation with these models, preserve the assistant’s reasoning_content, tool_calls, and matching tool result messages. Cache-hit input tokens are included in total input tokens, not added again. Submit images with Chat image_url or Responses input_image. glm-5.3-flash additionally accepts documents (file / input_file) and video; documents, audio, and video are not accepted for deepseek-v4.1-flash or deepseek-v4-flash.

Official model IDs

deepseek-flash and deepseek-v4-pro are accepted as equivalent IDs for deepseek-v4.1-flash and deepseek-v4, so a client written against those official names works unchanged. Usage and billing are recorded once under the APIAny.AI model ID. GET /v1/models lists both IDs and marks the equivalent entry with alias_of.
For GLM, use "model": "glm-5.3-flash" with the same example. In Responses, use input, max_output_tokens, and reasoning: { "effort": "low" }. For Responses tool follow-ups, append the complete returned output items (including reasoning and function calls) and the matching function_call_output to your next input. Do not discard reasoning items. DeepSeek Responses is stateless: resend the conversation history instead of relying on previous_response_id. Use reasoning_effort in Chat Completions and reasoning: { "effort": "high" } in Responses. Each endpoint also accepts the other form as a compatibility alias. If both are present, their values must match; conflicting values return 400. Omitting effort preserves the model default. The syntax accepts none, minimal, low, medium, high, xhigh, max, and ultra. This is not a promise that every model supports every value. Invalid values and values outside a documented model’s supported set return 400. Reasoning is always on for kimi-k3, glm-5.3-flash, gpt-6-astra, the Grok 4.5 family and the Gemini 3 family; a value that disables reasoning returns 400. gpt-5.4 and its mini and nano variants reason only when you ask for low or higher. minimax-m3 and gpt-4o-mini take no effort parameter at all; minimax-m3 exposes the thinking object instead. Some models also accept the compatibility spellings that older integrations used: grok-4.5, grok-4.5-latest, grok-4.6 and grok-4.7 accept max (executed as xhigh) and minimal (executed as low), and gemini-3.1-pro accepts minimal (executed as low). They are normalized before the request runs, so the model always receives an official level. capability_metadata.reasoning.compatEfforts publishes these mappings; new integrations should use the values in efforts. For Anthropic Messages, use output_config.effort. Native thinking remains independent: effort does not enable thinking or select budget_tokens. For Gemini 3, native requests use generationConfig.thinkingConfig.thinkingLevel; the accepted levels vary by model. Gemini 2.5 uses thinkingBudget. In the OpenAI-compatible format, Gemini 2.5 maps minimal/low to 1024, medium to 8192, high to 24576, and none to 0 (except Pro). Do not combine explicit effort with native thinkingLevel or thinkingBudget: this returns 400 even if the settings are equivalent. includeThoughts alone can be combined with effort. For other OpenAI-compatible models, syntactically valid efforts are accepted for model execution without a declared support guarantee. Consult the model’s capabilities before relying on an effort value. Output limits such as max_completion_tokens, max_output_tokens, and the native equivalents include reasoning and final-answer tokens. Effort is not an output-token limit. A request can still exhaust its budget (finish_reason: "length" in Chat Completions) and return an incomplete or empty final answer.

Model capability metadata

GET /v1/models returns a capability_metadata object for every model that declares its capabilities. Read it to configure a model without guessing:
  • input lists the accepted input types: text, image, video, audio, file.
  • output lists the produced types.
  • features reports stream, tools, vision, jsonMode, and whether the model reasons at all (thinking).
  • reasoning describes the reasoning controls: mode (always-on, optional, or none), the accepted efforts, the defaultEffort that applies when you omit effort, and disableWith for models whose reasoning can be turned off.
  • limits reports contextWindow and maxOutputTokens.
reasoning.mode is the field to check before you send an effort value: always-on means reasoning cannot be disabled, optional means disableWith applies, and none means the model has no reasoning mode.
API format references: OpenAI, Anthropic Messages, Gemini, Gemini OpenAI.