/v1beta/models/{model}:generateContent for Gemini-compatible integrations.
Streaming
Call:streamGenerateContent instead of :generateContent to receive a
Gemini server-sent event stream of incremental candidates chunks.
Structured output
UsegenerationConfig.responseMimeType: "application/json" with
responseSchema for JSON output:
Parameters
SupportedgenerationConfig: temperature, topP, topK, maxOutputTokens,
stopSequences, responseMimeType, responseSchema, responseModalities
(include IMAGE for image output). Also top-level tools (functionDeclarations)
and systemInstruction.
Reasoning effort
Gemini 3 usesgenerationConfig.thinkingConfig.thinkingLevel; Gemini 2.5 uses generationConfig.thinkingConfig.thinkingBudget. Supported levels and budgets vary by model. Gemini 3 does not accept none, and Gemini 2.5 Pro cannot disable thinking.
minimal/low to a 1024-token thinking budget, medium to 8192, high to 24576, and none to 0 where supported. Do not combine explicit effort with native thinkingLevel or thinkingBudget: this returns 400 even if the settings are equivalent. includeThoughts alone can be combined with effort. See model-specific reasoning controls.
Output limits such as max_completion_tokens, max_output_tokens, and the native equivalents include reasoning and final-answer tokens. Effort is not an output-token limit. A request can still exhaust its budget (finish_reason: "length" in Chat Completions) and return an incomplete or empty final answer.