diff --git a/docs/data/nav.yml b/docs/data/nav.yml index 88e362a4d..1ad2b9067 100644 --- a/docs/data/nav.yml +++ b/docs/data/nav.yml @@ -199,6 +199,8 @@ url: /providers/mistral/ - title: Moonshot AI url: /providers/moonshot/ + - title: Nativ + url: /providers/nativ/ - title: Nebius url: /providers/nebius/ - title: NVIDIA diff --git a/docs/providers/local/index.md b/docs/providers/local/index.md index c4122a108..86f2f99cc 100644 --- a/docs/providers/local/index.md +++ b/docs/providers/local/index.md @@ -1,7 +1,7 @@ --- -title: "Local Models (Ollama, vLLM, LocalAI)" +title: "Local Models (Ollama, vLLM, LocalAI, Nativ)" description: "Run Docker Agent with locally hosted models for privacy, offline use, or cost savings." -keywords: docker agent, ai agents, model providers, llm, local models, ollama, vllm, localai, offline models +keywords: docker agent, ai agents, model providers, llm, local models, ollama, vllm, localai, nativ, offline models linkTitle: "Local Models" weight: 150 canonical: https://docs.docker.com/ai/docker-agent/providers/local/ @@ -19,6 +19,7 @@ Docker Agent can connect to any OpenAI-compatible local model server. This guide - **Ollama** — Easy-to-use local model runner - **vLLM** — High-performance inference server - **LocalAI** — OpenAI-compatible API for various backends +- [**Nativ**](../nativ/index.md) — Local MLX models on Apple Silicon > [!TIP] > **Docker Model Runner** diff --git a/docs/providers/nativ/index.md b/docs/providers/nativ/index.md new file mode 100644 index 000000000..4179528ac --- /dev/null +++ b/docs/providers/nativ/index.md @@ -0,0 +1,143 @@ +--- +title: "Nativ" +description: "Run Docker Agent with local MLX models served by Nativ on Apple Silicon." +keywords: docker agent, ai agents, model providers, local models, nativ, mlx, apple silicon +weight: 185 +canonical: https://docs.docker.com/ai/docker-agent/providers/nativ/ +--- + +_Run Docker Agent with local MLX models served by Nativ on Apple Silicon._ + +## Overview + +[Nativ](https://github.com/Blaizzy/nativ) is a macOS app for downloading and +serving MLX models locally. Docker Agent connects to its OpenAI-compatible API +through a [provider definition](../custom/index.md). No built-in `nativ` alias +or additional provider plugin is required. + +Nativ requires Apple Silicon and macOS 26 or newer. Model downloads require +network access; inference runs locally after the model is downloaded. + +## Setup + +1. Install Nativ from its [releases](https://github.com/Blaizzy/nativ/releases/latest) + or with Homebrew: + + ```bash + $ brew install --cask nativ + ``` + +2. Launch Nativ and complete its initial setup. In **Models**, download and + select a compatible language model. For a small first download, use + [`mlx-community/Qwen3-0.6B-4bit`](https://huggingface.co/mlx-community/Qwen3-0.6B-4bit), + approximately **351 MB** including tokenizer files. +3. Start the server from Nativ's **Developer** page or menu-bar controls. Keep + Nativ running while you use Docker Agent. The default OpenAI base URL is + `http://127.0.0.1:8080/v1`; the Developer page shows the configured host, + port, endpoints, and logs. +4. Check that the server responds: + + ```bash + $ curl http://127.0.0.1:8080/health + $ curl http://127.0.0.1:8080/v1/models + ``` + +If you enabled server authentication, include the configured Bearer token in +requests. See [Authentication](#authentication) below. + +## Configuration + +Save this as `nativ.yaml`: + +```yaml +providers: + nativ: + api_type: openai_chatcompletions + base_url: http://127.0.0.1:8080/v1 + provider_opts: + extra_body: + enable_thinking: false + +models: + tiny: + provider: nativ + model: mlx-community/Qwen3-0.6B-4bit + temperature: 0 + max_tokens: 256 + +agents: + root: + model: tiny + description: Local chat assistant using Nativ on Apple Silicon + instruction: You are a concise helpful assistant. Answer directly. +``` + +Run a short chat request: + +```bash +$ docker agent run nativ.yaml --exec "What is the capital of France? Answer with only the city name." +``` + +A copy is available in +[`examples/nativ.yaml`](https://github.com/docker/docker-agent/blob/main/examples/nativ.yaml). +Use the repository ID of your downloaded model in `models.tiny.model` if you +choose a different model. Update `base_url` if you change Nativ's port. + +`provider_opts.extra_body.enable_thinking: false` sends Nativ's request-level +reasoning control to the server. It avoids spending the small output budget +on reasoning. See [Extra Request Body](../../configuration/models/index.md#extra-request-body) +for other pass-through settings. + +Nativ also exposes the OpenAI Responses API. To use it, change +`providers.nativ.api_type` to `openai_responses`; keep the same base URL. +Docker Agent does not forward `provider_opts.extra_body` with Responses +requests, so the thinking override above applies only to Chat Completions. +Use Chat Completions when you need that request-level control; with Responses, +check Nativ's server defaults and allow a larger output budget if thinking is +enabled. + +## Authentication + +If you enable a server API key in Nativ, add `token_key: NATIV_API_KEY` under +`providers.nativ` and set that environment variable to the same key: + +```bash +$ export NATIV_API_KEY=your-server-api-key +$ curl http://127.0.0.1:8080/v1/models -H "Authorization: Bearer $NATIV_API_KEY" +$ docker agent run nativ.yaml --exec "Hello" +``` + +`token_key` names an environment variable, not a literal token. Omit it when +server authentication is disabled. Nativ's server key is separate from a +Hugging Face token used to download gated models. + +## Tool Calling + +The example is chat-only. Nativ's API supports tool calls, but reliability +depends on the model, instructions, and tool schemas. A 0.6B model is useful +for testing the connection, not a dependable coding agent: it can ignore a +tool or give an incorrect answer even when instructed to use one. + +Before adding filesystem or shell tools, choose a tool-capable model and test +that it calls the tool with correct arguments and uses the returned result. +Keep tool descriptions and parameter schemas clear, and retain tool approvals. +See [Tools](../../concepts/tools/index.md) for configuration. + +## Troubleshooting + +- **Connection refused:** Start Nativ's server and check the host and port in + its Developer page. Installing the app alone does not make the API available. +- **Model not found:** Confirm that the model is downloaded and use its exact + repository ID. Check `/v1/models` and Nativ's server logs. +- **Unauthorized:** Match `NATIV_API_KEY` to Nativ's configured server key, or + remove `token_key` when authentication is disabled. +- **No answer before the token limit:** Keep thinking disabled for the tiny + model, or increase `max_tokens` for a model that needs longer responses. +- **Docker Agent runs in a container:** Container-local `127.0.0.1` does not + refer to your Mac. On Docker Desktop, use + `http://host.docker.internal:8080/v1` and ensure Nativ listens on an interface + reachable from the container. Enable authentication before exposing the + server beyond loopback; changing its bind address can expose it to your LAN. + +For server controls and endpoint details, see Nativ's +[Developer documentation](https://github.com/Blaizzy/nativ/blob/main/Docs/features/developer.md). diff --git a/docs/providers/overview/index.md b/docs/providers/overview/index.md index 7771c6017..0cb49f2b7 100644 --- a/docs/providers/overview/index.md +++ b/docs/providers/overview/index.md @@ -19,6 +19,7 @@ _Docker Agent supports multiple AI model providers. Choose the right one for you - [**AWS Bedrock**](../bedrock/index.md) — access Claude, Nova, Llama, and more through AWS infrastructure. - [**Docker Model Runner**](../dmr/index.md) — run models locally with Docker. No API keys, no costs. - [**Local Models**](../local/index.md) — run Ollama, vLLM, or LocalAI locally. No API key required. +- [**Nativ**](../nativ/index.md) — run MLX models locally on Apple Silicon through a custom provider. - [**Provider Definitions**](../custom/index.md) — define reusable provider configurations with shared defaults for any provider type. ## Quick Comparison diff --git a/examples/nativ.yaml b/examples/nativ.yaml new file mode 100644 index 000000000..4b01f6227 --- /dev/null +++ b/examples/nativ.yaml @@ -0,0 +1,23 @@ +# Download this small MLX model in Nativ and start its local server first. +providers: + nativ: + api_type: openai_chatcompletions + base_url: http://127.0.0.1:8080/v1 + # If server authentication is enabled, set NATIV_API_KEY and uncomment: + # token_key: NATIV_API_KEY + provider_opts: + extra_body: + enable_thinking: false + +models: + tiny: + provider: nativ + model: mlx-community/Qwen3-0.6B-4bit + temperature: 0 + max_tokens: 256 + +agents: + root: + model: tiny + description: Local chat assistant using Nativ on Apple Silicon + instruction: You are a concise helpful assistant. Answer directly.