For the complete documentation index, see llms.txt. This page is also available as Markdown.

Install on Mac

Install the Roboflow Inference Server on macOS with the native Apple Silicon app, with Docker, or outside Docker with MPS acceleration.

macOS native app (Apple Silicon)

You can run the Roboflow Inference Server on your Apple Silicon Mac with the native desktop app. Download the latest DMG disk image from the latest GitHub release: View the latest release and download installers on GitHub.

  1. Mount the disk image by double-clicking it.

  2. Drag the Roboflow Inference app to your Applications folder.

  3. Open your Applications folder and double-click the Roboflow Inference app to start the server.

Using Docker

First, install Docker Desktop. Then use the CLI to start the container:

pip install inference-cli
inference server start

If you want more control over the container settings, start it manually:

sudo docker run -d \
    --name inference-server \
    --read-only \
    -p 9001:9001 \
    --volume ~/.inference/cache:/tmp:rw \
    --security-opt="no-new-privileges" \
    --cap-drop="ALL" \
    --cap-add="NET_BIND_SERVICE" \
    roboflow/roboflow-inference-server-cpu:latest

Apple does not yet support passing the Metal Performance Shaders (MPS) device to Docker, so hardware acceleration is not possible inside a container on Mac. To use MPS you must run the server outside Docker.

We recommend pyenv and pyenv-virtualenv to manage your Python environments on Mac, especially because homebrew defaults to Python 3.13, which is not yet compatible with several of the machine learning dependencies Inference uses.

Once you have installed and set up pyenv and pyenv-virtualenv (follow the full instructions for setting up your shell), create and activate an inference virtual environment with Python 3.12:

pyenv install 3.12
pyenv virtualenv 3.12 inference
pyenv activate inference

To install and run the server outside Docker, clone the repo, install the dependencies, copy cpu_http.py into the top level of the repo, and start the server with uvicorn:

git clone https://github.com/roboflow/inference.git
cd inference
pip install .
cp docker/config/cpu_http.py .
uvicorn cpu_http:app --port 9001 --host 0.0.0.0

Your server is now running at http://localhost:9001 with MPS acceleration.

Next steps

Last updated

Was this helpful?