> For the complete documentation index, see [llms.txt](https://docs.roboflow.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.roboflow.com/deployment/roboflow-cloud/serverless-api/model-pricing.md).

# Model pricing

Credit prices for the models served by the Serverless Cloud API, per 1,000 images.

Serverless Cloud API image inference is billed in credits per image. Every price on this page is credits per 1,000 images at the standard rate, rounded to four decimal places. Each row is a model architecture priced from its smallest size, and each section expands to show every size.

Volume discounts apply automatically once monthly usage crosses a threshold. [Pricing](/deployment/roboflow-cloud/serverless-api/pricing.md) covers how billing works, including Workflow runs and which failed requests are billable. Video is priced separately in [Serverless Video Streaming](/deployment/roboflow-cloud/serverless-api/serverless-video-streaming-api.md).

## Object detection

Sorted by credits per 1,000 images divided by COCO mAP 50:95, using the smallest size of each architecture. Accuracy figures come from the [Roboflow model leaderboard](https://leaderboard.roboflow.com/).

These are the five most efficient architectures for this task.

<table data-search="false"><thead><tr><th>Model</th><th>Credits per 1,000 images</th><th>Sizes</th><th>Best COCO mAP 50:95</th></tr></thead><tbody><tr><td>RF-DETR</td><td>0.125</td><td>7</td><td>59.9%</td></tr><tr><td>YOLO26</td><td>0.125</td><td>5</td><td>56.3%</td></tr><tr><td>YOLOv12</td><td>0.125</td><td>5</td><td>54.0%</td></tr><tr><td>YOLOv11</td><td>0.125</td><td>5</td><td>53.6%</td></tr><tr><td>YOLOv10</td><td>0.125</td><td>6</td><td>53.6%</td></tr></tbody></table>

<details>

<summary>Every object detection model and size</summary>

* RF-DETR
  * Nano: 0.125
  * Small, Base, Medium, Large: 0.1875
  * XL: 0.25
  * 2XL: 0.3125
* YOLO26
  * Nano: 0.125
  * Small, Medium, Large: 0.1875
  * XL: 0.25
* YOLOv12
  * Nano: 0.125
  * Small, Medium, Large: 0.1875
  * XL: 0.25
* YOLOv11
  * Nano: 0.125
  * Small, Medium, Large: 0.1875
  * XL: 0.25
* YOLOv10
  * Nano: 0.125
  * Small, Medium, Balanced, Large: 0.1875
  * XL: 0.25
* YOLOv8
  * Nano: 0.125
  * Small, Medium, Large: 0.1875
  * XL: 0.25
* YOLOv9
  * Tiny: 0.125
  * Small, Medium, Compact: 0.1875
  * Extended: 0.25
* YOLOv7
  * Tiny: 0.125
  * Base: 0.1875
  * X: 0.25
* YOLOv5
  * Nano: 0.125
  * Small, Medium, Large: 0.1875
  * XL: 0.25
* YOLOLite
  * Nano: 0.125
  * Small, Medium, Large: 0.1875
  * XL: 0.25
* YOLO-NAS
  * Small, Medium: 0.1875

</details>

## Instance segmentation

Sorted by price, lowest first.

These are the five lowest-priced architectures for this task.

<table data-search="false"><thead><tr><th>Model</th><th>Credits per 1,000 images</th><th>Sizes</th></tr></thead><tbody><tr><td>RF-DETR Seg</td><td>0.1875</td><td>7</td></tr><tr><td>YOLO26 Seg</td><td>0.1875</td><td>5</td></tr><tr><td>YOLOv11 Seg</td><td>0.1875</td><td>5</td></tr><tr><td>YOLOv8 Seg</td><td>0.1875</td><td>5</td></tr><tr><td>YOLOv5 Seg</td><td>0.1875</td><td>5</td></tr></tbody></table>

<details>

<summary>Every instance segmentation model and size</summary>

* RF-DETR Seg
  * Nano: 0.1875
  * Small, Medium, Large, Preview: 0.25
  * XL, 2XL: 0.3125
* YOLO26 Seg
  * Nano: 0.1875
  * Small, Medium, Large: 0.25
  * XL: 0.3125
* YOLOv11 Seg
  * Nano: 0.1875
  * Small, Medium, Large: 0.25
  * XL: 0.3125
* YOLOv8 Seg
  * Nano: 0.1875
  * Small, Medium, Large: 0.25
  * XL: 0.3125
* YOLOv5 Seg
  * Nano: 0.1875
  * Small, Medium, Large: 0.25
  * XL: 0.3125
* YOLACT
  * ResNet-101: 0.25

</details>

## Semantic segmentation

<table data-search="false"><thead><tr><th>Model</th><th>Credits per 1,000 images</th><th>Sizes</th></tr></thead><tbody><tr><td>YOLO26 Semantic Seg</td><td>0.25</td><td>5</td></tr><tr><td>DeepLabV3+</td><td>0.25</td><td>1</td></tr></tbody></table>

<details>

<summary>Every semantic segmentation model and size</summary>

* YOLO26 Semantic Seg
  * Nano, Small, Medium, Large, XL: 0.25
* DeepLabV3+
  * MobileNetV2: 0.25

</details>

## Keypoint detection

<table data-search="false"><thead><tr><th>Model</th><th>Credits per 1,000 images</th><th>Sizes</th></tr></thead><tbody><tr><td>YOLO26 Pose</td><td>0.125</td><td>5</td></tr><tr><td>YOLOv11 Pose</td><td>0.125</td><td>5</td></tr><tr><td>YOLOv8 Pose</td><td>0.125</td><td>5</td></tr><tr><td>RF-DETR Keypoint</td><td>0.25</td><td>1</td></tr></tbody></table>

<details>

<summary>Every keypoint detection model and size</summary>

* YOLO26 Pose
  * Nano: 0.125
  * Small, Medium, Large: 0.1875
  * XL: 0.25
* YOLOv11 Pose
  * Nano: 0.125
  * Small, Medium, Large: 0.1875
  * XL: 0.25
* YOLOv8 Pose
  * Nano: 0.125
  * Small, Medium, Large: 0.1875
  * XL: 0.25
* RF-DETR Keypoint
  * Preview: 0.25

</details>

## Classification

<table data-search="false"><thead><tr><th>Model</th><th>Credits per 1,000 images</th><th>Sizes</th></tr></thead><tbody><tr><td>ViT</td><td>0.0625</td><td>1</td></tr><tr><td>ResNet</td><td>0.0625</td><td>4</td></tr><tr><td>DINOv3 Probe</td><td>0.0625</td><td>2</td></tr></tbody></table>

<details>

<summary>Every classification model and size</summary>

* ViT
  * B/16: 0.0625
* ResNet
  * 18, 34, 50, 101: 0.0625
* DINOv3 Probe
  * Small: 0.0625
  * Base: 0.125

</details>

## Open-vocabulary detection

These models detect objects from a text prompt, with no training required.

<table data-search="false"><thead><tr><th>Model</th><th>Credits per 1,000 images</th><th>Sizes</th></tr></thead><tbody><tr><td>YOLO-World</td><td>0.1875</td><td>8</td></tr><tr><td>Roboflow Instant</td><td>0.75</td><td>1</td></tr><tr><td>Grounding DINO</td><td>0.75</td><td>1</td></tr><tr><td>OWLv2</td><td>0.75</td><td>2</td></tr></tbody></table>

<details>

<summary>Every open-vocabulary detection model and size</summary>

* YOLO-World
  * S, M, L, X, v2-S, v2-M, v2-L, v2-X: 0.1875
* Roboflow Instant
  * Single size: 0.75
* Grounding DINO
  * Swin-T: 0.75
* OWLv2
  * Large, Large (fine-tuned): 0.75

</details>

## Segment Anything

Promptable segmentation foundation models.

<table data-search="false"><thead><tr><th>Model</th><th>Credits per 1,000 images</th><th>Sizes</th></tr></thead><tbody><tr><td>SAM 2</td><td>0.3125</td><td>4</td></tr><tr><td>SAM 3</td><td>0.5</td><td>2</td></tr><tr><td>SAM</td><td>0.5</td><td>1</td></tr></tbody></table>

<details>

<summary>Every model and size in Segment Anything</summary>

* SAM 2
  * Hiera Tiny, Hiera Small: 0.3125
  * Hiera Base+, Hiera Large: 0.5
* SAM 3
  * Base, Large: 0.5
* SAM
  * ViT-H: 0.5

</details>

## OCR

PP-OCR is billed per run, not per stage. Calls with text detection turned off bill at the recognition-only rate, and every combination with detection on bills at the OCR rate.

<table data-search="false"><thead><tr><th>Model</th><th>Credits per 1,000 images</th><th>Sizes</th></tr></thead><tbody><tr><td>PP-OCR (recognition only)</td><td>0.125</td><td>3</td></tr><tr><td>TrOCR</td><td>0.3125</td><td>1</td></tr><tr><td>PP-OCR (detection + recognition)</td><td>0.375</td><td>3</td></tr><tr><td>DocTR</td><td>0.375</td><td>1</td></tr><tr><td>EasyOCR</td><td>0.375</td><td>1</td></tr></tbody></table>

<details>

<summary>Every model and size in OCR</summary>

* PP-OCR (recognition only)
  * Tiny, Small, Medium: 0.125
* TrOCR
  * Single size: 0.3125
* PP-OCR (detection + recognition)
  * Tiny, Small, Medium: 0.375
* DocTR
  * Single size: 0.375
* EasyOCR
  * Single size: 0.375

</details>

## Embeddings

<table data-search="false"><thead><tr><th>Model</th><th>Credits per 1,000 images</th><th>Sizes</th></tr></thead><tbody><tr><td>CLIP</td><td>0.125</td><td>3</td></tr><tr><td>Perception Encoder</td><td>0.3125</td><td>1</td></tr></tbody></table>

<details>

<summary>Every embeddings model and size</summary>

* CLIP
  * ViT-B/16, ViT-B/32: 0.125
  * ViT-L/14: 0.1875
* Perception Encoder
  * Single size: 0.3125

</details>

## Depth estimation

<table data-search="false"><thead><tr><th>Model</th><th>Credits per 1,000 images</th><th>Sizes</th></tr></thead><tbody><tr><td>Depth Anything V3</td><td>0.3125</td><td>2</td></tr><tr><td>Depth Anything V2</td><td>0.3125</td><td>1</td></tr></tbody></table>

<details>

<summary>Every depth estimation model and size</summary>

* Depth Anything V3
  * Small: 0.3125
  * Base: 0.75
* Depth Anything V2
  * Small: 0.3125

</details>

## Execution time pricing

Custom Python blocks and vision language models are billed for the time they run, not per image. One credit buys 500 execution seconds.

<table data-search="false"><thead><tr><th>Workload</th><th>Execution seconds per credit</th><th>Credits per second</th><th>Credits per hour</th></tr></thead><tbody><tr><td>Custom Python blocks</td><td>500</td><td>0.002</td><td>7.2</td></tr><tr><td>Vision language models</td><td>500</td><td>0.002</td><td>7.2</td></tr></tbody></table>

### Vision language models

Every vision language model is billed for its execution time at the rate above, whatever the prompt or the number of output tokens.

<table data-search="false"><thead><tr><th>Model</th><th>Execution seconds per credit</th></tr></thead><tbody><tr><td>PaliGemma 2</td><td>500</td></tr><tr><td>Florence 2</td><td>500</td></tr><tr><td>Qwen2.5 VL</td><td>500</td></tr><tr><td>Qwen3 VL</td><td>500</td></tr><tr><td>Qwen3.5</td><td>500</td></tr><tr><td>SmolVLM</td><td>500</td></tr><tr><td>Cosmos 3 (Edge) VLM</td><td>500</td></tr><tr><td>Moondream</td><td>500</td></tr><tr><td>Gemma 4</td><td>500</td></tr><tr><td>Mage-VL</td><td>500</td></tr><tr><td>GLM-OCR</td><td>500</td></tr></tbody></table>

A step that runs a vision language model for 2.5 seconds per image, across 10,000 images, uses 25,000 execution seconds, or 50 credits.

## Volume discounts

Volume discounts are graduated: each rate applies only to the images that fall inside its band, and images below a threshold keep the rate of the band they fall in. Bands count a workspace total across all models for the calendar month. Select a band to see the rates that apply to images in it.

{% tabs %}
{% tab title="First 250K" %}
Standard rates, charged on the first 250,000 images each month.

<table data-search="false"><thead><tr><th>Standard rate</th><th>Models on this rate</th><th>Rate in this band</th></tr></thead><tbody><tr><td>0.0625</td><td>Classification models</td><td>0.0625</td></tr><tr><td>0.125</td><td>Nano detection, CLIP, PP-OCR recognition only</td><td>0.125</td></tr><tr><td>0.1875</td><td>Small–Large detection, Nano segmentation</td><td>0.1875</td></tr><tr><td>0.25</td><td>XL detection, Small–Large segmentation</td><td>0.25</td></tr><tr><td>0.3125</td><td>XL segmentation, SAM 2 Tiny/Small</td><td>0.3125</td></tr><tr><td>0.375</td><td>OCR pipelines</td><td>0.375</td></tr><tr><td>0.5</td><td>Segment Anything</td><td>0.5</td></tr><tr><td>0.75</td><td>Open-vocabulary detection, Depth Anything V3 Base</td><td>0.75</td></tr></tbody></table>
{% endtab %}

{% tab title="250K–500K" %}
25% off.

<table data-search="false"><thead><tr><th>Standard rate</th><th>Models on this rate</th><th>Rate in this band</th></tr></thead><tbody><tr><td>0.0625</td><td>Classification models</td><td>0.0469</td></tr><tr><td>0.125</td><td>Nano detection, CLIP, PP-OCR recognition only</td><td>0.0938</td></tr><tr><td>0.1875</td><td>Small–Large detection, Nano segmentation</td><td>0.1406</td></tr><tr><td>0.25</td><td>XL detection, Small–Large segmentation</td><td>0.1875</td></tr><tr><td>0.3125</td><td>XL segmentation, SAM 2 Tiny/Small</td><td>0.2344</td></tr><tr><td>0.375</td><td>OCR pipelines</td><td>0.2813</td></tr><tr><td>0.5</td><td>Segment Anything</td><td>0.375</td></tr><tr><td>0.75</td><td>Open-vocabulary detection, Depth Anything V3 Base</td><td>0.5625</td></tr></tbody></table>
{% endtab %}

{% tab title="500K–1M" %}
50% off.

<table data-search="false"><thead><tr><th>Standard rate</th><th>Models on this rate</th><th>Rate in this band</th></tr></thead><tbody><tr><td>0.0625</td><td>Classification models</td><td>0.0313</td></tr><tr><td>0.125</td><td>Nano detection, CLIP, PP-OCR recognition only</td><td>0.0625</td></tr><tr><td>0.1875</td><td>Small–Large detection, Nano segmentation</td><td>0.0938</td></tr><tr><td>0.25</td><td>XL detection, Small–Large segmentation</td><td>0.125</td></tr><tr><td>0.3125</td><td>XL segmentation, SAM 2 Tiny/Small</td><td>0.1563</td></tr><tr><td>0.375</td><td>OCR pipelines</td><td>0.1875</td></tr><tr><td>0.5</td><td>Segment Anything</td><td>0.25</td></tr><tr><td>0.75</td><td>Open-vocabulary detection, Depth Anything V3 Base</td><td>0.375</td></tr></tbody></table>
{% endtab %}

{% tab title="1M–2M" %}
60% off.

<table data-search="false"><thead><tr><th>Standard rate</th><th>Models on this rate</th><th>Rate in this band</th></tr></thead><tbody><tr><td>0.0625</td><td>Classification models</td><td>0.025</td></tr><tr><td>0.125</td><td>Nano detection, CLIP, PP-OCR recognition only</td><td>0.05</td></tr><tr><td>0.1875</td><td>Small–Large detection, Nano segmentation</td><td>0.075</td></tr><tr><td>0.25</td><td>XL detection, Small–Large segmentation</td><td>0.1</td></tr><tr><td>0.3125</td><td>XL segmentation, SAM 2 Tiny/Small</td><td>0.125</td></tr><tr><td>0.375</td><td>OCR pipelines</td><td>0.15</td></tr><tr><td>0.5</td><td>Segment Anything</td><td>0.2</td></tr><tr><td>0.75</td><td>Open-vocabulary detection, Depth Anything V3 Base</td><td>0.3</td></tr></tbody></table>
{% endtab %}

{% tab title="2M–5M" %}
70% off.

<table data-search="false"><thead><tr><th>Standard rate</th><th>Models on this rate</th><th>Rate in this band</th></tr></thead><tbody><tr><td>0.0625</td><td>Classification models</td><td>0.0188</td></tr><tr><td>0.125</td><td>Nano detection, CLIP, PP-OCR recognition only</td><td>0.0375</td></tr><tr><td>0.1875</td><td>Small–Large detection, Nano segmentation</td><td>0.0563</td></tr><tr><td>0.25</td><td>XL detection, Small–Large segmentation</td><td>0.075</td></tr><tr><td>0.3125</td><td>XL segmentation, SAM 2 Tiny/Small</td><td>0.0938</td></tr><tr><td>0.375</td><td>OCR pipelines</td><td>0.1125</td></tr><tr><td>0.5</td><td>Segment Anything</td><td>0.15</td></tr><tr><td>0.75</td><td>Open-vocabulary detection, Depth Anything V3 Base</td><td>0.225</td></tr></tbody></table>
{% endtab %}

{% tab title="Over 5M" %}
80% off.

<table data-search="false"><thead><tr><th>Standard rate</th><th>Models on this rate</th><th>Rate in this band</th></tr></thead><tbody><tr><td>0.0625</td><td>Classification models</td><td>0.0125</td></tr><tr><td>0.125</td><td>Nano detection, CLIP, PP-OCR recognition only</td><td>0.025</td></tr><tr><td>0.1875</td><td>Small–Large detection, Nano segmentation</td><td>0.0375</td></tr><tr><td>0.25</td><td>XL detection, Small–Large segmentation</td><td>0.05</td></tr><tr><td>0.3125</td><td>XL segmentation, SAM 2 Tiny/Small</td><td>0.0625</td></tr><tr><td>0.375</td><td>OCR pipelines</td><td>0.075</td></tr><tr><td>0.5</td><td>Segment Anything</td><td>0.1</td></tr><tr><td>0.75</td><td>Open-vocabulary detection, Depth Anything V3 Base</td><td>0.15</td></tr></tbody></table>
{% endtab %}
{% endtabs %}

For example, 600,000 images through RF-DETR Small in one month bill as 250,000 at 0.1875, then 250,000 at 0.1406, then 100,000 at 0.0938, for a total of 91.41 credits.

## Models priced elsewhere

The video trackers SAM 2 Video and SAM 3 Video run on Serverless Video and are priced in [Serverless Video Streaming](/deployment/roboflow-cloud/serverless-api/serverless-video-streaming-api.md).
