For the complete documentation index, see llms.txt. This page is also available as Markdown.

Perception Encoder

Dedicated Deployment या self-hosted Inference पर image और text embeddings compute करने के लिए Meta के Perception Encoder का उपयोग करें

Perception Encoder Meta का vision-language embedding model है। यह images और text को similarity search, zero-shot classification, और retrieval के लिए एक shared embedding space में map करता है।

Perception Encoder Serverless Hosted API पर उपलब्ध नहीं है। इसे एक पर चलाएँ Dedicated Deployment या self-hosted Inference.

हम तीन Perception Encoder endpoints का समर्थन करते हैं:

  • /perception_encoder/embed_image — एक image embed करें

  • /perception_encoder/embed_text — एक string embed करें

  • /perception_encoder/compare — एक image और text prompts की सूची के बीच similarity compute करें

कोड नमूना

1

अपनी API Key प्राप्त करें

एक Roboflow खाता बनाएं, अपनी key यहाँ पर ढूँढें Roboflow API settings page और इसे अपने shell में उपलब्ध कराएँ:

export ROBOFLOW_API_KEY="your-key-here"
2

निर्भरताएँ इंस्टॉल करें

ये packages image को fetch करते हैं और API को call करते हैं:

pip install requests opencv-python
3

मॉडल चलाएँ

नीचे का sample एक image भेजता है /perception_encoder/embed_image और embedding shape प्रिंट करता है। सेट करें URL को अपने Dedicated Deployment URL या local Inference server पर।

import base64
import os
import cv2
import numpy as np
import requests

URL = "https://your-deployment.roboflow.cloud"

content = requests.get("https://media.roboflow.com/notebooks/examples/dog.jpeg").content
image = cv2.imdecode(np.frombuffer(content, np.uint8), cv2.IMREAD_COLOR)

_, buffer = cv2.imencode(".jpg", image)
image_base64 = base64.b64encode(buffer).decode("utf-8")

response = requests.post(
    f"{URL}/perception_encoder/embed_image",
    json={
        "api_key": os.environ["ROBOFLOW_API_KEY"],
        "image": {"type": "base64", "value": image_base64},
    },
)
result = response.json()
embedding = result["embeddings"][0]
print(f"Embedding length: {len(embedding)}")
print(f"First values: {embedding[:5]}")

ऊपर का code terminal में embedding shape प्रिंट करता है:

Embedding length: 1024
First values: [0.0545, -0.0338, -0.0355, -0.0062, 0.0154]

Inference speed

Latency मापी गई Roboflow Inference 1x NVIDIA L4 पर, batch size 1, warmup के बाद का औसत।

मॉडल
विलंबता (ms)

perception-encoder

25.2

के साथ मापा गया embed_image पर PE-Core-L14-336 checkpoint (केवल image embedding).

सेट करें URL को अपने deployment target से मिलाएँ:

अंतिम अपडेट

क्या यह उपयोगी था?