Google Gemini
Run Google's Gemini model with vision capabilities.
v3
Ask a question to Google's Gemini model with vision capabilities.
You can specify arbitrary text prompts or predefined ones, the block supports the following types of prompt:
Open Prompt (
unconstrained) - Use any prompt to generate a raw responseText Recognition (OCR) (
ocr) - Model recognizes text in the imageVisual Question Answering (
visual-question-answering) - Model answers the question you submit in the promptCaptioning (short) (
caption) - Model provides a short description of the imageCaptioning (
detailed-caption) - Model provides a long description of the imageSingle-Label Classification (
classification) - Model classifies the image content as one of the provided classesMulti-Label Classification (
multi-label-classification) - Model classifies the image content as one or more of the provided classesUnprompted Object Detection (
object-detection) - Model detects and returns the bounding boxes for prominent objects in the imageStructured Output Generation (
structured-answering) - Model returns a JSON response with the specified fields
API Key Options
This block supports two API key modes:
Roboflow Managed API Key (Default) - Use
rf_key:accountto proxy requests through Roboflow's API:Simplified setup - no Google AI API key required
Secure - your workflow API key is used for authentication
Usage-based billing - charged per token based on the model used
Custom Google AI API Key - Provide your own Google AI API key:
Full control over API usage
You pay Google directly
WARNING!
This block makes use of /v1beta API of Google Gemini model - the implementation may change in the future, without guarantee of backward compatibility.
Type identifier
Use the following identifier in step "type" field: roboflow_core/google_gemini@v3 to add the block as a step in your workflow.
Properties
Name
Type
Description
Refs
name
str
Enter a unique identifier for this step..
❌
task_type
str
Task type to be performed by model. Value determines required parameters and output response..
❌
prompt
str
Text prompt to the Gemini model.
✅
output_structure
Dict[str, str]
Dictionary with structure of expected JSON response.
❌
classes
List[str]
List of classes to be used.
✅
api_key
str
Your Google AI API key or 'rf_key:account' to use Roboflow's managed API key.
✅
model_version
str
Model to be used.
✅
thinking_level
str
Controls the depth of internal reasoning for Gemini 3+ models. 'low' minimizes latency and cost (best for simple tasks), 'high' maximizes reasoning depth (default). Only supported by Gemini 3 and newer models..
✅
temperature
float
Temperature to sample from the model - value in range 0.0-2.0, the higher - the more random / "creative" the generations are..
✅
max_tokens
int
Maximum number of tokens the model can generate in it's response. If not specified, the model will use its default limit..
❌
google_code_execution
bool
Enable native code execution for the Gemini model..
✅
max_concurrent_requests
int
Number of concurrent requests that can be executed by block when batch of input images provided. If not given - block defaults to value configured globally in Workflows Execution Engine. Please restrict if you hit Google Gemini API limits..
❌
The Refs column marks possibility to parametrise the property with dynamic values available in workflow runtime. See Bindings for more info.
Runtime compatibility
requires_internet - air-gapped / offline deployments : This block depends on a service that is not reachable from fully offline / air-gapped deployments.
Input and Output Bindings
The available connections depend on its binding kinds. Check what binding kinds Google Gemini in version v3 has.
Input and output bindings
input
images(image): The image to infer on..prompt(string): Text prompt to the Gemini model.classes(list_of_values): List of classes to be used.api_key(Union[ROBOFLOW_MANAGED_KEY,secret,string]): Your Google AI API key or 'rf_key:account' to use Roboflow's managed API key.model_version(string): Model to be used.thinking_level(string): Controls the depth of internal reasoning for Gemini 3+ models. 'low' minimizes latency and cost (best for simple tasks), 'high' maximizes reasoning depth (default). Only supported by Gemini 3 and newer models..temperature(float): Temperature to sample from the model - value in range 0.0-2.0, the higher - the more random / "creative" the generations are..google_code_execution(boolean): Enable native code execution for the Gemini model..
output
output(Union[string,language_model_output]): String value ifstringor LLM / VLM output iflanguage_model_output.classes(list_of_values): List of values of any type.
v2
Ask a question to Google's Gemini model with vision capabilities.
You can specify arbitrary text prompts or predefined ones, the block supports the following types of prompt:
Open Prompt (
unconstrained) - Use any prompt to generate a raw responseText Recognition (OCR) (
ocr) - Model recognizes text in the imageVisual Question Answering (
visual-question-answering) - Model answers the question you submit in the promptCaptioning (short) (
caption) - Model provides a short description of the imageCaptioning (
detailed-caption) - Model provides a long description of the imageSingle-Label Classification (
classification) - Model classifies the image content as one of the provided classesMulti-Label Classification (
multi-label-classification) - Model classifies the image content as one or more of the provided classesUnprompted Object Detection (
object-detection) - Model detects and returns the bounding boxes for prominent objects in the imageStructured Output Generation (
structured-answering) - Model returns a JSON response with the specified fields
You need to provide your Google AI API key to use the Gemini model.
WARNING!
This block makes use of /v1beta API of Google Gemini model - the implementation may change in the future, without guarantee of backward compatibility.
Type identifier
Use the following identifier in step "type" field: roboflow_core/google_gemini@v2 to add the block as a step in your workflow.
Properties
Name
Type
Description
Refs
name
str
Enter a unique identifier for this step..
❌
task_type
str
Task type to be performed by model. Value determines required parameters and output response..
❌
prompt
str
Text prompt to the Gemini model.
✅
output_structure
Dict[str, str]
Dictionary with structure of expected JSON response.
❌
classes
List[str]
List of classes to be used.
✅
api_key
str
Your Google AI API key.
✅
model_version
str
Model to be used.
✅
thinking_level
str
Controls the depth of internal reasoning for Gemini 3+ models. 'low' minimizes latency and cost (best for simple tasks), 'high' maximizes reasoning depth (default). Only supported by Gemini 3 and newer models..
✅
temperature
float
Temperature to sample from the model - value in range 0.0-2.0, the higher - the more random / "creative" the generations are..
✅
max_tokens
int
Maximum number of tokens the model can generate in it's response. If not specified, the model will use its default limit..
❌
max_concurrent_requests
int
Number of concurrent requests that can be executed by block when batch of input images provided. If not given - block defaults to value configured globally in Workflows Execution Engine. Please restrict if you hit Google Gemini API limits..
❌
The Refs column marks possibility to parametrise the property with dynamic values available in workflow runtime. See Bindings for more info.
Runtime compatibility
requires_internet - air-gapped / offline deployments : This block depends on a service that is not reachable from fully offline / air-gapped deployments.
Input and Output Bindings
The available connections depend on its binding kinds. Check what binding kinds Google Gemini in version v2 has.
Input and output bindings
input
images(image): The image to infer on..prompt(string): Text prompt to the Gemini model.classes(list_of_values): List of classes to be used.model_version(string): Model to be used.thinking_level(string): Controls the depth of internal reasoning for Gemini 3+ models. 'low' minimizes latency and cost (best for simple tasks), 'high' maximizes reasoning depth (default). Only supported by Gemini 3 and newer models..temperature(float): Temperature to sample from the model - value in range 0.0-2.0, the higher - the more random / "creative" the generations are..
output
output(Union[string,language_model_output]): String value ifstringor LLM / VLM output iflanguage_model_output.classes(list_of_values): List of values of any type.
v1
Ask a question to Google's Gemini model with vision capabilities.
You can specify arbitrary text prompts or predefined ones, the block supports the following types of prompt:
Open Prompt (
unconstrained) - Use any prompt to generate a raw responseText Recognition (OCR) (
ocr) - Model recognizes text in the imageVisual Question Answering (
visual-question-answering) - Model answers the question you submit in the promptCaptioning (short) (
caption) - Model provides a short description of the imageCaptioning (
detailed-caption) - Model provides a long description of the imageSingle-Label Classification (
classification) - Model classifies the image content as one of the provided classesMulti-Label Classification (
multi-label-classification) - Model classifies the image content as one or more of the provided classesUnprompted Object Detection (
object-detection) - Model detects and returns the bounding boxes for prominent objects in the imageStructured Output Generation (
structured-answering) - Model returns a JSON response with the specified fields
You need to provide your Google AI API key to use the Gemini model.
WARNING!
This block makes use of /v1beta API of Google Gemini model - the implementation may change in the future, without guarantee of backward compatibility.
Type identifier
Use the following identifier in step "type" field: roboflow_core/google_gemini@v1 to add the block as a step in your workflow.
Properties
Name
Type
Description
Refs
name
str
Enter a unique identifier for this step..
❌
task_type
str
Task type to be performed by model. Value determines required parameters and output response..
❌
prompt
str
Text prompt to the Gemini model.
✅
output_structure
Dict[str, str]
Dictionary with structure of expected JSON response.
❌
classes
List[str]
List of classes to be used.
✅
api_key
str
Your Google AI API key.
✅
model_version
str
Model to be used.
✅
max_tokens
int
Maximum number of tokens the model can generate in it's response..
❌
temperature
float
Temperature to sample from the model - value in range 0.0-2.0, the higher - the more random / "creative" the generations are..
✅
max_concurrent_requests
int
Number of concurrent requests that can be executed by block when batch of input images provided. If not given - block defaults to value configured globally in Workflows Execution Engine. Please restrict if you hit Google Gemini API limits..
❌
The Refs column marks possibility to parametrise the property with dynamic values available in workflow runtime. See Bindings for more info.
Runtime compatibility
requires_internet - air-gapped / offline deployments : This block depends on a service that is not reachable from fully offline / air-gapped deployments.
Input and Output Bindings
The available connections depend on its binding kinds. Check what binding kinds Google Gemini in version v1 has.
Input and output bindings
input
images(image): The image to infer on..prompt(string): Text prompt to the Gemini model.classes(list_of_values): List of classes to be used.model_version(string): Model to be used.temperature(float): Temperature to sample from the model - value in range 0.0-2.0, the higher - the more random / "creative" the generations are..
output
output(Union[string,language_model_output]): String value ifstringor LLM / VLM output iflanguage_model_output.classes(list_of_values): List of values of any type.
Last updated
Was this helpful?