For the complete documentation index, see llms.txt. This page is also available as Markdown.

Grounding DINO

Dedicated Deployment 또는 self-hosted Inference에서 Grounding DINO를 사용해 텍스트 프롬프트 기반 객체 탐지를 수행합니다.

Grounding DINO는 개방형 어휘 객체 탐지기입니다. 이미지와 텍스트 클래스 목록을 전달하면, 모델은 추가 학습 없이 일치하는 영역의 바운딩 박스를 반환합니다.

Grounding DINO는 Serverless Hosted API에서 사용할 수 없습니다. 다음에서 실행하세요 Dedicated Deployment 또는 자체 호스팅 Inference.

코드 샘플

1

API Key를 받으세요

Roboflow 계정을 만들고, 다음에서 키를 찾으세요: Roboflow API 설정 페이지 그리고 이를 셸에서 사용할 수 있도록 설정하세요:

export ROBOFLOW_API_KEY="your-key-here"
2

종속성을 설치하세요

이 패키지들은 API를 호출하고 결과를 그립니다:

pip install requests supervision opencv-python
3

모델을 실행하세요

설정 URL 를 Dedicated Deployment URL 또는 로컬 Inference server에.

import base64
import os
import cv2
import numpy as np
import requests
import supervision as sv

URL = "https://your-deployment.roboflow.cloud"
content = requests.get("https://media.roboflow.com/notebooks/examples/dog.jpeg").content
image = cv2.imdecode(np.frombuffer(content, np.uint8), cv2.IMREAD_COLOR)
_, buffer = cv2.imencode(".jpg", image)
image_base64 = base64.b64encode(buffer).decode("utf-8")

response = requests.post(
    f"{URL}/grounding_dino/infer",
    json={
        "api_key": os.environ["ROBOFLOW_API_KEY"],
        "image": {"type": "base64", "value": image_base64},
        "text": ["dog", "person", "backpack"],
    },
)
preds = response.json()["predictions"]

xyxys = [
    [p["x"] - p["width"] / 2, p["y"] - p["height"] / 2,
     p["x"] + p["width"] / 2, p["y"] + p["height"] / 2]
    for p in preds
]
detections = sv.Detections(
    xyxy=np.array(xyxys, dtype=float),
    class_id=np.array([p.get("class_id", 0) for p in preds]),
    confidence=np.array([p["confidence"] for p in preds], dtype=float),
    data={"class_name": np.array([p["class"] for p in preds])},
)
labels = [f"{p['class']} {p['confidence']:.2f}" for p in preds]
annotated = sv.BoxAnnotator().annotate(image.copy(), detections)
annotated = sv.LabelAnnotator().annotate(annotated, detections, labels=labels)
cv2.imwrite("dog_annotated.png", annotated)

추론 속도

다음 기준으로 측정한 지연 시간 Roboflow Inference 1x NVIDIA L4에서, 배치 크기 1로, 워밍업 이후 평균값입니다.

모델
지연 시간(ms)

grounding-dino

165.4

두 개의 텍스트 프롬프트로 측정됨.

설정 URL 을 배포 대상에 맞게 설정하세요:

마지막 업데이트

도움이 되었나요?