> For the complete documentation index, see [llms.txt](https://docs.roboflow.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.roboflow.com/workflows/blocks/blocks/run-a-model.md).

# Run a Model

Run fine-tuned models trained on Roboflow, foundation models, and third-party VLMs and LLMs directly in a Workflow.

| Block                                                                                                            | Description                                                                                 |
| ---------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- |
| [Anthropic Claude](/workflows/blocks/blocks/run-a-model/anthropic-claude.md)                                     | Run Anthropic Claude model with vision capabilities                                         |
| [Barcode Detection](/workflows/blocks/blocks/run-a-model/barcode-detection.md)                                   | Detect and read barcodes in an image                                                        |
| [Clip Comparison](/workflows/blocks/blocks/run-a-model/clip-comparison.md)                                       | Compare CLIP image and text embeddings                                                      |
| [CLIP Embedding Model](/workflows/blocks/blocks/run-a-model/clip-embedding-model.md)                             | Generate an embedding of an image or string                                                 |
| [CogVLM](/workflows/blocks/blocks/run-a-model/cog-vlm.md)                                                        | DEPRECATED! Run a self-hosted vision language model                                         |
| [Cosmos 3](/workflows/blocks/blocks/run-a-model/cosmos3.md)                                                      | Reason about physical scenes with NVIDIA Cosmos 3                                           |
| [Depth Estimation](/workflows/blocks/blocks/run-a-model/depth-estimation.md)                                     | Estimate relative scene depth in an image                                                   |
| [EasyOCR](/workflows/blocks/blocks/run-a-model/easy-ocr.md)                                                      | Extract text from an image using EasyOCR optical character recognition                      |
| [Florence-2 Model](/workflows/blocks/blocks/run-a-model/florence2-model.md)                                      | Run Florence-2 on an image                                                                  |
| [GLM-OCR](/workflows/blocks/blocks/run-a-model/glmocr.md)                                                        | Run GLM-OCR on an image to recognize text                                                   |
| [Google Gemini](/workflows/blocks/blocks/run-a-model/google-gemini.md)                                           | Run Google's Gemini model with vision capabilities                                          |
| [Google Gemma](/workflows/blocks/blocks/run-a-model/google-gemma.md)                                             | Run Google's Gemma model with vision capabilities via OpenRouter                            |
| [Google Vision OCR](/workflows/blocks/blocks/run-a-model/google-vision-ocr.md)                                   | Detect text in images using Google Vision API                                               |
| [Instance Segmentation Model](/workflows/blocks/blocks/run-a-model/instance-segmentation-model.md)               | Predict the shape, size, and location of objects                                            |
| [Keypoint Detection Model](/workflows/blocks/blocks/run-a-model/keypoint-detection-model.md)                     | Predict skeletons on objects                                                                |
| [Llama 3.2 Vision](/workflows/blocks/blocks/run-a-model/llama3-2-vision.md)                                      | Run Llama 3.2 Vision via OpenRouter                                                         |
| [Moondream2](/workflows/blocks/blocks/run-a-model/moondream2.md)                                                 | Run Moondream2 on an image                                                                  |
| [MoonshotAI Kimi](/workflows/blocks/blocks/run-a-model/moonshot-ai-kimi.md)                                      | Run Moonshot AI Kimi vision-language models via OpenRouter                                  |
| [Multi-Label Classification Model](/workflows/blocks/blocks/run-a-model/multi-label-classification-model.md)     | Apply multiple tags to an image                                                             |
| [Object Detection Model](/workflows/blocks/blocks/run-a-model/object-detection-model.md)                         | Predict the location of objects with bounding boxes                                         |
| [OCR Model](/workflows/blocks/blocks/run-a-model/ocr-model.md)                                                   | Extract text from an image using DocTR optical character recognition                        |
| [OpenAI](/workflows/blocks/blocks/run-a-model/open-ai.md)                                                        | Run OpenAI's GPT models with vision capabilities                                            |
| [OpenAI-Compatible LLM](/workflows/blocks/blocks/run-a-model/open-ai-compatible-llm.md)                          | Send prompts to any OpenAI-compatible API endpoint                                          |
| [OpenRouter](/workflows/blocks/blocks/run-a-model/open-router.md)                                                | Run any OpenRouter model by pasting its model slug                                          |
| [Perception Encoder Embedding Model](/workflows/blocks/blocks/run-a-model/perception-encoder-embedding-model.md) | Generate an embedding of an image or string                                                 |
| [PP-OCR](/workflows/blocks/blocks/run-a-model/ppocr.md)                                                          | Extract text from an image using PP-OCR (PaddleOCR) optical character recognition           |
| [QR Code Detection](/workflows/blocks/blocks/run-a-model/qr-code-detection.md)                                   | Detect and read QR codes in an image                                                        |
| [Qwen-VL](/workflows/blocks/blocks/run-a-model/qwen-vl.md)                                                       | Run any Qwen vision model - natively or via OpenRouter                                      |
| [Qwen3.5](/workflows/blocks/blocks/run-a-model/qwen3-5.md)                                                       | Run Qwen3.5 on an image                                                                     |
| [Roboflow Visual Search Classifier](/workflows/blocks/blocks/run-a-model/roboflow-visual-search-classifier.md)   | Classify an image by finding the most visually similar annotated image                      |
| [SAM 3](/workflows/blocks/blocks/run-a-model/sam3.md)                                                            | Run SAM3 with text prompts for zero-shot segmentation                                       |
| [SAM 3 Interactive](/workflows/blocks/blocks/run-a-model/sam3-interactive.md)                                    | Segment a specific object with SAM3 using point and/or bounding box prompts                 |
| [Seg Preview](/workflows/blocks/blocks/run-a-model/seg-preview.md)                                               | Seg Preview                                                                                 |
| [Segment Anything 2 Model](/workflows/blocks/blocks/run-a-model/segment-anything2-model.md)                      | Convert bounding boxes to polygons, or run SAM2 on an entire image to generate a mask       |
| [Semantic Segmentation Model](/workflows/blocks/blocks/run-a-model/semantic-segmentation-model.md)               | Assign a class label to every pixel in the image                                            |
| [Single-Label Classification Model](/workflows/blocks/blocks/run-a-model/single-label-classification-model.md)   | Apply a single tag to an image                                                              |
| [SmolVLM2](/workflows/blocks/blocks/run-a-model/smol-vlm2.md)                                                    | Run SmolVLM2 on an image                                                                    |
| [Stability AI Image Generation](/workflows/blocks/blocks/run-a-model/stability-ai-image-generation.md)           | generate new images from text, or create variations of existing images                      |
| [Stability AI Inpainting](/workflows/blocks/blocks/run-a-model/stability-ai-inpainting.md)                       | Use segmentation masks to inpaint objects within an image                                   |
| [Stability AI Outpainting](/workflows/blocks/blocks/run-a-model/stability-ai-outpainting.md)                     | Use object detection bounding box to crop the image and to outpaint within given directions |
| [YOLO-World Model](/workflows/blocks/blocks/run-a-model/yolo-world-model.md)                                     | Run a zero-shot object detection model                                                      |
| [Gaze Detection (Deprecated)](/workflows/blocks/blocks/run-a-model/gaze-detection.md)                            | Detect faces and estimate gaze direction (deprecated)                                       |
| [Google Gemma API (Deprecated)](/workflows/blocks/blocks/run-a-model/google-gemma-api.md)                        | Run Google's Gemma model with vision capabilities via OpenRouter                            |
| [LMM (Deprecated)](/workflows/blocks/blocks/run-a-model/lmm.md)                                                  | Run a large multimodal model such as ChatGPT-4v                                             |
| [LMM For Classification (Deprecated)](/workflows/blocks/blocks/run-a-model/lmm-for-classification.md)            | Run a large multimodal model such as ChatGPT-4v for classification                          |
| [Qwen 3.5 API (Deprecated)](/workflows/blocks/blocks/run-a-model/qwen3-5-api.md)                                 | Run Qwen 3.5 vision-language models via OpenRouter                                          |
| [Qwen 3.6 API (Deprecated)](/workflows/blocks/blocks/run-a-model/qwen3-6-api.md)                                 | Run Qwen 3.6 vision-language models via OpenRouter                                          |
| [Qwen2.5-VL (Deprecated)](/workflows/blocks/blocks/run-a-model/qwen2-5-vl.md)                                    | Run Qwen2.5-VL on an image                                                                  |
| [Qwen3-VL (Deprecated)](/workflows/blocks/blocks/run-a-model/qwen3-vl.md)                                        | Run Qwen3-VL on an image                                                                    |
| [Qwen3.5-VL (Deprecated)](/workflows/blocks/blocks/run-a-model/qwen3-5-vl.md)                                    | Run Qwen3.5-VL on an image                                                                  |
