SmolVLM2
Run SmolVLM2 on an image.
Last updated
Was this helpful?
Run SmolVLM2 on an image.
This workflow block runs SmolVLM2, a multimodal vision-language model. You can ask questions about images -- including documents and photos -- and get answers in natural language.
Use the following identifier in step "type" field: roboflow_core/smolvlm2@v1 to add the block as a step in your workflow.
Name
Type
Description
Refs
name
str
Enter a unique identifier for this step..
❌
prompt
str
Optional text prompt to provide additional context to SmolVLM2. Otherwise it will just be None.
❌
model_version
str
The SmolVLM2 model to be used for inference..
✅
The Refs column marks possibility to parametrise the property with dynamic values available in workflow runtime. See Bindings for more info.
hard - runtime self_hosted_cpu; execution local : Requires a GPU; run_locally() loads a model that needs CUDA.
The available connections depend on its binding kinds. Check what binding kinds SmolVLM2 in version v1 has.
input
images (image): The image to infer on..
model_version (roboflow_model_id): The SmolVLM2 model to be used for inference..
output
parsed_output (dictionary): Dictionary.
Last updated
Was this helpful?
Was this helpful?
{
"name": "<your_step_name_here>",
"type": "roboflow_core/smolvlm2@v1",
"images": "$inputs.image",
"prompt": "What is in this image?",
"model_version": "smolvlm2/smolvlm-2.2b-instruct"
}