/v1/chat/completions/v1/responses/v1/messages/v1beta/models/{model}:generateContent
- Input tokens.
- Output tokens.
- Cached input tokens.
- Minimum request credits.
Available models
Grok 4.7
Use the independent model IDgrok-4.7 with /v1/chat/completions. Grok 4.6 remains a separate model. APIAny input and cached-input pricing is 1.2 credits per 1,000 tokens; output pricing is 2.4 credits per 1,000 tokens, matching Grok 4.6 at launch. Credit-pack discounts continue to apply.
/v1/models for current availability and pricing.
Launch testing observed refusals for some harmless exact-phrase echo requests, despite successful ordinary text and code responses. Do not rely on exact-string echo behavior as an availability probe. This is an observed model limitation, not a guarantee that every instruction will be followed.
Reasoning effort
DeepSeek Flash and GLM 5.3 Flash
deepseek-v4-flash and deepseek-v4.1-flash are separate models, not aliases. V4 Flash accepts text; V4.1 Flash accepts text and images. Both support up to 393,216 output tokens within a 1,048,576-token context window. GLM 5.3 Flash supports the same context window and up to 131,072 output tokens. Use /v1/models for current availability and prices. glm-5.2 is no longer available for new requests; previous usage records remain available.
DeepSeek Flash accepts none, low, high, and max. Compatibility values minimal map to low; medium and xhigh map to high. In Chat Completions, thinking: { "type": "disabled" } or effort none disables thinking. Do not combine contradictory settings.
For deepseek-v4.1-flash Chat Completions, thinking is enabled by default. Forced tools (tool_choice: "required" or a named function) require reasoning_effort: "none" or thinking: { "type": "disabled" }. Otherwise the request returns HTTP 400 with invalid_tool_choice. To keep thinking enabled, use tool_choice: "auto"; APIAny.AI does not silently disable thinking for you.
GLM 5.3 Flash always thinks. Its effort values are low, high, and max (model default); disabling thinking returns 400. For a quick smoke test, use low. Allocate enough tokens for thinking plus the final answer, for example max_tokens: 32768. Explicit limits are honored, not silently increased. max_completion_tokens is accepted as an alias; different simultaneous limits return 400.
When continuing a tool conversation with these models, preserve the assistant’s reasoning_content, tool_calls, and matching tool result messages. Cache-hit input tokens are included in total input tokens, not added again. Submit images with Chat image_url or Responses input_image. glm-5.3-flash additionally accepts documents (file / input_file) and video; documents, audio, and video are not accepted for deepseek-v4.1-flash or deepseek-v4-flash.
Official model IDs
deepseek-flash and deepseek-v4-pro are accepted as equivalent IDs for deepseek-v4.1-flash and deepseek-v4, so a client written against those official names works unchanged. Usage and billing are recorded once under the APIAny.AI model ID. GET /v1/models lists both IDs and marks the equivalent entry with alias_of.
"model": "glm-5.3-flash" with the same example. In Responses, use input, max_output_tokens, and reasoning: { "effort": "low" }.
For Responses tool follow-ups, append the complete returned output items (including reasoning and function calls) and the matching function_call_output to your next input. Do not discard reasoning items. DeepSeek Responses is stateless: resend the conversation history instead of relying on previous_response_id.
Use reasoning_effort in Chat Completions and reasoning: { "effort": "high" } in Responses. Each endpoint also accepts the other form as a compatibility alias. If both are present, their values must match; conflicting values return 400. Omitting effort preserves the model default.
The syntax accepts none, minimal, low, medium, high, xhigh, max, and ultra. This is not a promise that every model supports every value. Invalid values and values outside a documented model’s supported set return 400.
Reasoning is always on for
kimi-k3, glm-5.3-flash, gpt-6-astra, the Grok 4.5 family and the Gemini 3 family; a value that disables reasoning returns 400. gpt-5.4 and its mini and nano variants reason only when you ask for low or higher. minimax-m3 and gpt-4o-mini take no effort parameter at all; minimax-m3 exposes the thinking object instead.
Some models also accept the compatibility spellings that older integrations used: grok-4.5, grok-4.5-latest, grok-4.6 and grok-4.7 accept max (executed as xhigh) and minimal (executed as low), and gemini-3.1-pro accepts minimal (executed as low). They are normalized before the request runs, so the model always receives an official level. capability_metadata.reasoning.compatEfforts publishes these mappings; new integrations should use the values in efforts.
For Anthropic Messages, use output_config.effort. Native thinking remains independent: effort does not enable thinking or select budget_tokens.
For Gemini 3, native requests use generationConfig.thinkingConfig.thinkingLevel; the accepted levels vary by model. Gemini 2.5 uses thinkingBudget. In the OpenAI-compatible format, Gemini 2.5 maps minimal/low to 1024, medium to 8192, high to 24576, and none to 0 (except Pro). Do not combine explicit effort with native thinkingLevel or thinkingBudget: this returns 400 even if the settings are equivalent. includeThoughts alone can be combined with effort.
For other OpenAI-compatible models, syntactically valid efforts are accepted for model execution without a declared support guarantee. Consult the model’s capabilities before relying on an effort value.
Output limits such as max_completion_tokens, max_output_tokens, and the native equivalents include reasoning and final-answer tokens. Effort is not an output-token limit. A request can still exhaust its budget (finish_reason: "length" in Chat Completions) and return an incomplete or empty final answer.
Model capability metadata
GET /v1/models returns a capability_metadata object for every model that declares its capabilities. Read it to configure a model without guessing:
inputlists the accepted input types:text,image,video,audio,file.outputlists the produced types.featuresreportsstream,tools,vision,jsonMode, and whether the model reasons at all (thinking).reasoningdescribes the reasoning controls:mode(always-on,optional, ornone), the acceptedefforts, thedefaultEffortthat applies when you omit effort, anddisableWithfor models whose reasoning can be turned off.limitsreportscontextWindowandmaxOutputTokens.
reasoning.mode is the field to check before you send an effort value: always-on means reasoning cannot be disabled, optional means disableWith applies, and none means the model has no reasoning mode.