OpenAI
Run OpenAI's GPT models with vision capabilities.
v5
Ask a question to OpenAI's GPT models with vision capabilities (including GPT-5 and GPT-4o).
You can specify arbitrary text prompts or predefined ones, the block supports the following types of prompt:
Open Prompt (
unconstrained) - Use any prompt to generate a raw responseText Recognition (OCR) (
ocr) - Model recognizes text in the imageVisual Question Answering (
visual-question-answering) - Model answers the question you submit in the promptCaptioning (short) (
caption) - Model provides a short description of the imageCaptioning (
detailed-caption) - Model provides a long description of the imageSingle-Label Classification (
classification) - Model classifies the image content as one of the provided classesMulti-Label Classification (
multi-label-classification) - Model classifies the image content as one or more of the provided classesUnprompted Object Detection (
object-detection) - Model detects and returns the bounding boxes for prominent objects in the imageStructured Output Generation (
structured-answering) - Model returns a JSON response with the specified fields
The object-detection task uses a per-model prompt contract selected from a large-scale benchmark - use roboflow_core/vlm_as_detector@v2 to convert any of the outputs into predictions:
Most models (GPT-5.2 and newer, plus unknown/future models) return
{"detections": [...]}withbox_2dboxes in[x_min, y_min, x_max, y_max]format (absolute pixel coordinates of the uploaded image) and alabelkey, enforced via structured outputs.GPT-5.1/GPT-5 generation models return the legacy normalized dict (
x_min/y_min/x_max/y_maxin 0.0-1.0 withclass_nameandconfidence).GPT-4.x/GPT-4o and nano-tier models return a plain JSON list of
box_2dentries in absolute pixel coordinates.
The absolute-coordinate formats do not provide confidence scores. roboflow_core/vlm_as_detector@v2 assigns these detections a confidence of 1.0, so downstream confidence filtering is not meaningful for these formats.
Images are downscaled so that their longest edge does not exceed 2048px and are sent as lossless PNG for this task.
Provide your OpenAI API key or set the value to rf_key:account (or rf_key:user:<id>) to proxy requests through Roboflow's API.
Type identifier
Use the following identifier in step "type" field: roboflow_core/open_ai@v5 to add the block as a step in your workflow.
Properties
Name
Type
Description
Refs
name
str
Enter a unique identifier for this step..
❌
task_type
str
Task type to be performed by model. Value determines required parameters and output response..
❌
prompt
str
Text prompt to the OpenAI model.
✅
output_structure
Dict[str, str]
Dictionary with structure of expected JSON response.
❌
classes
List[str]
List of classes to be used.
✅
api_key
str
Your OpenAI API key.
✅
model_version
str
Model to be used.
✅
reasoning_effort
str
Controls reasoning. Reducing can result in faster responses and fewer tokens. GPT-5.1 and higher models default to 'none' (no reasoning) and support 'none', 'low', 'medium', 'high'. GPT-5.2 also supports 'xhigh'. GPT-5 models default to 'medium' and support 'minimal', 'low', 'medium', 'high'..
✅
image_detail
str
Indicates the image's quality, with 'high' suggesting it is of high resolution and should be processed or displayed with high fidelity. Not applied to the object-detection task, which always sends the image without the detail hint..
✅
max_tokens
int
Maximum number of tokens the model can generate in its response. If not specified, the model will use its default limit. Minimum value is 16..
❌
temperature
float
Temperature to sample from the model - value in range 0.0-2.0, the higher - the more random / "creative" the generations are..
✅
max_concurrent_requests
int
Number of concurrent requests that can be executed by block when batch of input images provided. If not given - block defaults to value configured globally in Workflows Execution Engine. Please restrict if you hit OpenAI limits..
❌
The Refs column marks possibility to parametrise the property with dynamic values available in workflow runtime. See Bindings for more info.
Runtime compatibility
requires_internet - air-gapped / offline deployments : This block depends on a service that is not reachable from fully offline / air-gapped deployments.
Input and Output Bindings
The available connections depend on its binding kinds. Check what binding kinds OpenAI in version v5 has.
Input and output bindings
input
images(image): The image to infer on..prompt(string): Text prompt to the OpenAI model.classes(list_of_values): List of classes to be used.api_key(Union[ROBOFLOW_MANAGED_KEY,secret,string]): Your OpenAI API key.model_version(string): Model to be used.reasoning_effort(string): Controls reasoning. Reducing can result in faster responses and fewer tokens. GPT-5.1 and higher models default to 'none' (no reasoning) and support 'none', 'low', 'medium', 'high'. GPT-5.2 also supports 'xhigh'. GPT-5 models default to 'medium' and support 'minimal', 'low', 'medium', 'high'..image_detail(string): Indicates the image's quality, with 'high' suggesting it is of high resolution and should be processed or displayed with high fidelity. Not applied to theobject-detectiontask, which always sends the image without the detail hint..temperature(float): Temperature to sample from the model - value in range 0.0-2.0, the higher - the more random / "creative" the generations are..
output
output(Union[string,language_model_output]): String value ifstringor LLM / VLM output iflanguage_model_output.classes(list_of_values): List of values of any type.
v4
Ask a question to OpenAI's GPT models with vision capabilities (including GPT-5 and GPT-4o).
You can specify arbitrary text prompts or predefined ones, the block supports the following types of prompt:
Open Prompt (
unconstrained) - Use any prompt to generate a raw responseText Recognition (OCR) (
ocr) - Model recognizes text in the imageVisual Question Answering (
visual-question-answering) - Model answers the question you submit in the promptCaptioning (short) (
caption) - Model provides a short description of the imageCaptioning (
detailed-caption) - Model provides a long description of the imageSingle-Label Classification (
classification) - Model classifies the image content as one of the provided classesMulti-Label Classification (
multi-label-classification) - Model classifies the image content as one or more of the provided classesUnprompted Object Detection (
object-detection) - Model detects and returns the bounding boxes for prominent objects in the imageStructured Output Generation (
structured-answering) - Model returns a JSON response with the specified fields
Provide your OpenAI API key or set the value to rf_key:account (or rf_key:user:<id>) to proxy requests through Roboflow's API.
Type identifier
Use the following identifier in step "type" field: roboflow_core/open_ai@v4 to add the block as a step in your workflow.
Properties
Name
Type
Description
Refs
name
str
Enter a unique identifier for this step..
❌
task_type
str
Task type to be performed by model. Value determines required parameters and output response..
❌
prompt
str
Text prompt to the OpenAI model.
✅
output_structure
Dict[str, str]
Dictionary with structure of expected JSON response.
❌
classes
List[str]
List of classes to be used.
✅
api_key
str
Your OpenAI API key.
✅
model_version
str
Model to be used.
✅
reasoning_effort
str
Controls reasoning. Reducing can result in faster responses and fewer tokens. GPT-5.1 and higher models default to 'none' (no reasoning) and support 'none', 'low', 'medium', 'high'. GPT-5.2 also supports 'xhigh'. GPT-5 models default to 'medium' and support 'minimal', 'low', 'medium', 'high'..
✅
image_detail
str
Indicates the image's quality, with 'high' suggesting it is of high resolution and should be processed or displayed with high fidelity..
✅
max_tokens
int
Maximum number of tokens the model can generate in its response. If not specified, the model will use its default limit. Minimum value is 16..
❌
temperature
float
Temperature to sample from the model - value in range 0.0-2.0, the higher - the more random / "creative" the generations are..
✅
max_concurrent_requests
int
Number of concurrent requests that can be executed by block when batch of input images provided. If not given - block defaults to value configured globally in Workflows Execution Engine. Please restrict if you hit OpenAI limits..
❌
The Refs column marks possibility to parametrise the property with dynamic values available in workflow runtime. See Bindings for more info.
Runtime compatibility
requires_internet - air-gapped / offline deployments : This block depends on a service that is not reachable from fully offline / air-gapped deployments.
Input and Output Bindings
The available connections depend on its binding kinds. Check what binding kinds OpenAI in version v4 has.
Input and output bindings
input
images(image): The image to infer on..prompt(string): Text prompt to the OpenAI model.classes(list_of_values): List of classes to be used.api_key(Union[ROBOFLOW_MANAGED_KEY,secret,string]): Your OpenAI API key.model_version(string): Model to be used.reasoning_effort(string): Controls reasoning. Reducing can result in faster responses and fewer tokens. GPT-5.1 and higher models default to 'none' (no reasoning) and support 'none', 'low', 'medium', 'high'. GPT-5.2 also supports 'xhigh'. GPT-5 models default to 'medium' and support 'minimal', 'low', 'medium', 'high'..image_detail(string): Indicates the image's quality, with 'high' suggesting it is of high resolution and should be processed or displayed with high fidelity..temperature(float): Temperature to sample from the model - value in range 0.0-2.0, the higher - the more random / "creative" the generations are..
output
output(Union[string,language_model_output]): String value ifstringor LLM / VLM output iflanguage_model_output.classes(list_of_values): List of values of any type.
v3
Ask a question to OpenAI's GPT models with vision capabilities (including GPT-5 and GPT-4o).
You can specify arbitrary text prompts or predefined ones, the block supports the following types of prompt:
Open Prompt (
unconstrained) - Use any prompt to generate a raw responseText Recognition (OCR) (
ocr) - Model recognizes text in the imageVisual Question Answering (
visual-question-answering) - Model answers the question you submit in the promptCaptioning (short) (
caption) - Model provides a short description of the imageCaptioning (
detailed-caption) - Model provides a long description of the imageSingle-Label Classification (
classification) - Model classifies the image content as one of the provided classesMulti-Label Classification (
multi-label-classification) - Model classifies the image content as one or more of the provided classesStructured Output Generation (
structured-answering) - Model returns a JSON response with the specified fields
Provide your OpenAI API key or set the value to rf_key:account (or rf_key:user:<id>) to proxy requests through Roboflow's API.
Type identifier
Use the following identifier in step "type" field: roboflow_core/open_ai@v3 to add the block as a step in your workflow.
Properties
Name
Type
Description
Refs
name
str
Enter a unique identifier for this step..
❌
task_type
str
Task type to be performed by model. Value determines required parameters and output response..
❌
prompt
str
Text prompt to the OpenAI model.
✅
output_structure
Dict[str, str]
Dictionary with structure of expected JSON response.
❌
classes
List[str]
List of classes to be used.
✅
api_key
str
Your OpenAI API key.
✅
model_version
str
Model to be used.
✅
image_detail
str
Indicates the image's quality, with 'high' suggesting it is of high resolution and should be processed or displayed with high fidelity..
✅
max_tokens
int
Maximum number of tokens the model can generate in it's response..
❌
temperature
float
Temperature to sample from the model - value in range 0.0-2.0, the higher - the more random / "creative" the generations are..
✅
max_concurrent_requests
int
Number of concurrent requests that can be executed by block when batch of input images provided. If not given - block defaults to value configured globally in Workflows Execution Engine. Please restrict if you hit OpenAI limits..
❌
The Refs column marks possibility to parametrise the property with dynamic values available in workflow runtime. See Bindings for more info.
Runtime compatibility
requires_internet - air-gapped / offline deployments : This block depends on a service that is not reachable from fully offline / air-gapped deployments.
Input and Output Bindings
The available connections depend on its binding kinds. Check what binding kinds OpenAI in version v3 has.
Input and output bindings
input
images(image): The image to infer on..prompt(string): Text prompt to the OpenAI model.classes(list_of_values): List of classes to be used.api_key(Union[ROBOFLOW_MANAGED_KEY,secret,string]): Your OpenAI API key.model_version(string): Model to be used.image_detail(string): Indicates the image's quality, with 'high' suggesting it is of high resolution and should be processed or displayed with high fidelity..temperature(float): Temperature to sample from the model - value in range 0.0-2.0, the higher - the more random / "creative" the generations are..
output
output(Union[string,language_model_output]): String value ifstringor LLM / VLM output iflanguage_model_output.classes(list_of_values): List of values of any type.
v2
Ask a question to OpenAI's GPT models with vision capabilities (including GPT-4o and GPT-5).
You can specify arbitrary text prompts or predefined ones, the block supports the following types of prompt:
Open Prompt (
unconstrained) - Use any prompt to generate a raw responseText Recognition (OCR) (
ocr) - Model recognizes text in the imageVisual Question Answering (
visual-question-answering) - Model answers the question you submit in the promptCaptioning (short) (
caption) - Model provides a short description of the imageCaptioning (
detailed-caption) - Model provides a long description of the imageSingle-Label Classification (
classification) - Model classifies the image content as one of the provided classesMulti-Label Classification (
multi-label-classification) - Model classifies the image content as one or more of the provided classesStructured Output Generation (
structured-answering) - Model returns a JSON response with the specified fields
You need to provide your OpenAI API key to use the GPT-4 with Vision model.
Type identifier
Use the following identifier in step "type" field: roboflow_core/open_ai@v2 to add the block as a step in your workflow.
Properties
Name
Type
Description
Refs
name
str
Enter a unique identifier for this step..
❌
task_type
str
Task type to be performed by model. Value determines required parameters and output response..
❌
prompt
str
Text prompt to the OpenAI model.
✅
output_structure
Dict[str, str]
Dictionary with structure of expected JSON response.
❌
classes
List[str]
List of classes to be used.
✅
api_key
str
Your OpenAI API key.
✅
model_version
str
Model to be used.
✅
image_detail
str
Indicates the image's quality, with 'high' suggesting it is of high resolution and should be processed or displayed with high fidelity..
✅
max_tokens
int
Maximum number of tokens the model can generate in it's response..
❌
temperature
float
Temperature to sample from the model - value in range 0.0-2.0, the higher - the more random / "creative" the generations are..
✅
max_concurrent_requests
int
Number of concurrent requests that can be executed by block when batch of input images provided. If not given - block defaults to value configured globally in Workflows Execution Engine. Please restrict if you hit OpenAI limits..
❌
The Refs column marks possibility to parametrise the property with dynamic values available in workflow runtime. See Bindings for more info.
Runtime compatibility
requires_internet - air-gapped / offline deployments : This block depends on a service that is not reachable from fully offline / air-gapped deployments.
Input and Output Bindings
The available connections depend on its binding kinds. Check what binding kinds OpenAI in version v2 has.
Input and output bindings
input
images(image): The image to infer on..prompt(string): Text prompt to the OpenAI model.classes(list_of_values): List of classes to be used.model_version(string): Model to be used.image_detail(string): Indicates the image's quality, with 'high' suggesting it is of high resolution and should be processed or displayed with high fidelity..temperature(float): Temperature to sample from the model - value in range 0.0-2.0, the higher - the more random / "creative" the generations are..
output
output(Union[string,language_model_output]): String value ifstringor LLM / VLM output iflanguage_model_output.classes(list_of_values): List of values of any type.
v1
Ask a question to OpenAI's GPT-4 with Vision model.
You can specify arbitrary text prompts to the OpenAIBlock.
You need to provide your OpenAI API key to use the GPT-4 with Vision model.
This model was previously part of the LMM block.
Type identifier
Use the following identifier in step "type" field: roboflow_core/open_ai@v1 to add the block as a step in your workflow.
Properties
Name
Type
Description
Refs
name
str
Enter a unique identifier for this step..
❌
prompt
str
Text prompt to the OpenAI model.
✅
openai_api_key
str
Your OpenAI API key.
✅
openai_model
str
Model to be used.
✅
json_output_format
Dict[str, str]
Holds dictionary that maps name of requested output field into its description.
❌
image_detail
str
Indicates the image's quality, with 'high' suggesting it is of high resolution and should be processed or displayed with high fidelity..
✅
max_tokens
int
Maximum number of tokens the model can generate in it's response..
❌
The Refs column marks possibility to parametrise the property with dynamic values available in workflow runtime. See Bindings for more info.
Runtime compatibility
requires_internet - air-gapped / offline deployments : This block depends on a service that is not reachable from fully offline / air-gapped deployments.
Input and Output Bindings
The available connections depend on its binding kinds. Check what binding kinds OpenAI in version v1 has.
Input and output bindings
input
output
parent_id(parent_id): Identifier of parent for step output.root_parent_id(parent_id): Identifier of parent for step output.image(image_metadata): Dictionary with image metadata required by supervision.structured_output(dictionary): Dictionary.raw_output(string): String value.*(*): Equivalent of any element.
Last updated
Was this helpful?