For the complete documentation index, see llms.txt. This page is also available as Markdown.

Moondream2

Dedicated Deployment 또는 self-hosted Inference에서 Moondream2를 사용해 open-vocabulary detection을 수행합니다.

Moondream2는 컴팩트한 vision-language model입니다. Roboflow Inference에서는 open-vocabulary object detector로 노출됩니다: class name을 prompt로 전달하면 일치하는 영역의 bounding box를 받을 수 있습니다.

Moondream2는 Serverless Hosted API에서 사용할 수 없습니다. 다음에서 실행하세요 Dedicated Deployment 또는 자체 호스팅 Inference.

코드 샘플

1

API Key를 받으세요

Roboflow 계정을 만들고, 다음에서 키를 찾으세요: Roboflow API 설정 페이지 그리고 이를 셸에서 사용할 수 있도록 설정하세요:

export ROBOFLOW_API_KEY="your-key-here"
2

종속성을 설치하세요

다음을 설치하세요 Inference SDKsupervision:

pip install inference-sdk supervision opencv-python
3

모델을 실행하세요

설정 api_url 를 Dedicated Deployment URL 또는 로컬 Inference server에.

import os
import cv2
import numpy as np
import requests
import supervision as sv
from inference_sdk import InferenceHTTPClient

content = requests.get("https://media.roboflow.com/notebooks/examples/dog.jpeg").content
image = cv2.imdecode(np.frombuffer(content, np.uint8), cv2.IMREAD_COLOR)
client = InferenceHTTPClient(
    api_url="https://your-deployment.roboflow.cloud",
    api_key=os.environ["ROBOFLOW_API_KEY"],
)
result = client.infer_lmm(
    image,
    model_id="moondream2",
    prompt="dog",
)

preds = result["predictions"]
xyxys = [
    [p["x"] - p["width"] / 2, p["y"] - p["height"] / 2,
     p["x"] + p["width"] / 2, p["y"] + p["height"] / 2]
    for p in preds
]
detections = sv.Detections(
    xyxy=np.array(xyxys, dtype=float),
    class_id=np.array([p.get("class_id", 0) for p in preds]),
    confidence=np.array([p.get("confidence", 1.0) for p in preds], dtype=float),
    data={"class_name": np.array([p["class"] for p in preds])},
)
labels = [f"{p['class']} {p.get('confidence', 1.0):.2f}" for p in preds]
annotated = sv.BoxAnnotator().annotate(image.copy(), detections)
annotated = sv.LabelAnnotator().annotate(annotated, detections, labels=labels)
cv2.imwrite("dog_annotated.png", annotated)

추론 속도

다음 기준으로 측정한 지연 시간 Roboflow Inference 1x NVIDIA L4에서 batch size 1로 이미지 한 장의 캡션을 생성할 때. Moondream2는 출력 길이를 고정할 수 없으므로, 지연 시간은 응답에 따라 달라집니다.

별칭
지연 시간(ms)

moondream2

1669

설정 api_url 을 배포 대상에 맞게 설정하세요:

마지막 업데이트

도움이 되었나요?