For the complete documentation index, see llms.txt. This page is also available as Markdown.

LMM For Classification

Run a large multimodal model such as ChatGPT-4v for classification.

Classify an image into one or more categories using a Large Multimodal Model (LMM).

You can specify arbitrary classes to an LMMBlock.

The LLMBlock supports two LMMs:

  • OpenAI's GPT-4 with Vision.

You need to provide your OpenAI API key to use the GPT-4 with Vision model.

Type identifier

Use the following identifier in step "type" field: roboflow_core/lmm_for_classification@v1 to add the block as a step in your workflow.

Properties

Name

Type

Description

Refs

name

str

Enter a unique identifier for this step..

lmm_type

str

Type of LMM to be used.

classes

List[str]

List of classes that LMM shall classify against.

lmm_config

LMMConfig

Configuration of LMM.

remote_api_key

str

Holds API key required to call LMM model - in current state of development, we require OpenAI key when lmm_type=gpt_4v..

The Refs column marks possibility to parametrise the property with dynamic values available in workflow runtime. See Bindings for more info.

Runtime compatibility

requires_internet - air-gapped / offline deployments : This block depends on a service that is not reachable from fully offline / air-gapped deployments.

hard - runtime hosted_serverless; execution remote : LMM_ENABLED=False on Roboflow Hosted Serverless: the /llm_v1 endpoint is not registered, so run_remotely() returns 404.

Input and Output Bindings

The available connections depend on its binding kinds. Check what binding kinds LMM For Classification in version v1 has.

Input and output bindings
  • input

    • images (image): The image to infer on..

    • lmm_type (string): Type of LMM to be used.

    • classes (list_of_values): List of classes that LMM shall classify against.

    • remote_api_key (Union[secret, string]): Holds API key required to call LMM model - in current state of development, we require OpenAI key when lmm_type=gpt_4v..

  • output

    • raw_output (string): String value.

    • top (top_class): String value representing top class predicted by classification model.

    • parent_id (parent_id): Identifier of parent for step output.

    • root_parent_id (parent_id): Identifier of parent for step output.

    • image (image_metadata): Dictionary with image metadata required by supervision.

    • prediction_type (prediction_type): String value with type of prediction.

Example JSON definition

Last updated

Was this helpful?