Skip to content

Python API and Deployment

brookie edited this page Oct 8, 2026 · 2 revisions

Python API and Deployment

This page covers using API Dock from Python (building the app, mounting it, or calling RouteMapper from another framework) and running it in production.

from api_dock import create_app            # FastAPI (same as create_fastapi_app)

app = create_app("api_dock_config/config.yaml")
# uvicorn my_module:app --host 0.0.0.0 --port 8000

Building the app

Function Returns
api_dock.create_app(config_path=None) / create_fastapi_app FastAPI app
api_dock.create_flask_app(config_path=None) Flask app (refuses configs with PostgreSQL connections)
api_dock.fast_api:app, api_dock.flask_api:app default apps, built on first access from api_dock_config/config.yaml in the working directory

Remote and database files are read from the folder that holds the main config. config_path=None means api_dock_config/config.yaml, relative to the working directory. Building the app runs the startup checks, so a bad database config raises ValueError right away.

The FastAPI app keeps its RouteMapper in app.state.route_mapper. Its lifespan calls route_mapper.start() and aclose(), which open and close PostgreSQL pools.

Mounting inside another FastAPI app

Starlette doesn't run a mounted app's lifespan. If you use PostgreSQL connections, start and stop the mapper from the parent's lifespan:

from contextlib import asynccontextmanager

from fastapi import FastAPI

from api_dock import create_app

dock = create_app("api_dock_config/config.yaml")


@asynccontextmanager
async def lifespan(app: FastAPI):
    mapper = dock.state.route_mapper
    await mapper.start()          # no-op without PostgreSQL connections
    try:
        yield
    finally:
        await mapper.aclose()


app = FastAPI(lifespan=lifespan)
app.mount("/dock", dock)          # /dock/<remote-or-database>/...

Without PostgreSQL, the lifespan isn't needed.

RouteMapper

RouteMapper is the framework-independent core. The built-in apps are thin wrappers around it.

from api_dock import RouteMapper

mapper = RouteMapper("api_dock_config/config.yaml")
Member Purpose
get_config_metadata() dict for / (name, description, authors, endpoints, remotes; databases are listed under remotes)
is_remote_name(name), is_database_name(name) decide which method handles /<name>/...
get_remote_names(), get_database_names() configured slugs
await map_route(remote_name, path, method, headers=None, body=None, query_params=None, cookies=None, multi_query_params=None) proxy to a remote, buffered → ProxyResponse
map_route_sync(...) same arguments, for sync code → ProxyResponse
await map_database_route(database_name, path, query_params=None, cookies=None, multi_query_params=None) run a database route → ProxyResponse
await prepare_remote_request(...) same arguments as map_route, but returns a PreparedRequest (or an error ProxyResponse) without calling the upstream, for streaming
await start(), await aclose() open and close PostgreSQL pools. Required when database.connections is configured

path is everything after /<name>/, including the version segment ("latest/users/1"). query_params is a single value per key. multi_query_params maps each key to its list of values (build it with api_dock.route_mapper.collect_multi_query_params(pairs)). Pass it so repeated keys like ?id=1&id=2 reach remotes and multivalue_sql templates intact.

ProxyResponse

Every map_* method returns an api_dock.types.ProxyResponse:

Field Meaning
status_code from the upstream, or from API Dock (403, 404, 500, 502, 503, ...)
content body bytes (never None)
content_type e.g. application/json
headers upstream headers safe to forward (hop-by-hop, Content-Type, Content-Length and Content-Encoding removed; Set-Cookie is in set_cookies)
set_cookies the upstream's Set-Cookie values, one per cookie (send each as its own header). Since 0.10.0
error_message set only for API Dock's own errors. Upstream 4xx/5xx leave it None

Return it as-is. Don't re-parse the body.

Example: Django (or any sync framework)

import asyncio

from django.http import HttpResponse

from api_dock import RouteMapper
from api_dock.route_mapper import collect_multi_query_params

mapper = RouteMapper("api_dock_config/config.yaml")


def api_dock_view(request, name, path=""):
    cookies = dict(request.COOKIES)
    query = request.GET.dict()
    multi = collect_multi_query_params(
        (key, value) for key in request.GET for value in request.GET.getlist(key)
    )
    if mapper.is_database_name(name):
        resp = asyncio.run(mapper.map_database_route(
            name, path, query_params=query, cookies=cookies, multi_query_params=multi
        ))
    else:
        resp = mapper.map_route_sync(
            name, path, request.method,
            headers=dict(request.headers), body=request.body or None,
            query_params=query, cookies=cookies, multi_query_params=multi,
        )
    response = HttpResponse(resp.content, status=resp.status_code,
                            content_type=resp.content_type)
    for key, value in resp.headers.items():
        response[key] = value
    for cookie in resp.set_cookies:
        response.cookies.load(cookie)   # Django sends one Set-Cookie per cookie
    return response

map_route_sync and asyncio.run each start a new event loop, so call them only from sync code. Inside a running loop, map_route_sync returns a 500. PostgreSQL pools are tied to the event loop that started them, so PostgreSQL connections need an async server with one long-running loop (as the FastAPI app does). They can't be used from per-request asyncio.run.

Example: streaming a remote response

map_route reads the whole upstream body into memory. To stream instead (what the FastAPI app does), get a PreparedRequest and send it yourself with httpx:

import httpx

from api_dock.types import ProxyResponse

prepared = await mapper.prepare_remote_request(name, path, method, headers=headers,
                                               body=body, cookies=cookies,
                                               multi_query_params=multi)
if isinstance(prepared, ProxyResponse):
    ...  # 403/404/500 from API Dock: return it as-is
client = httpx.AsyncClient(follow_redirects=prepared.follow_redirects,
                           timeout=prepared.timeout)
request = client.build_request(prepared.method, prepared.url, headers=prepared.headers,
                               params=prepared.params, cookies=prepared.cookies,
                               content=prepared.body)
upstream = await client.send(request, stream=True)
# pipe upstream.aiter_raw() to your response, then close upstream and client

Running in production

FastAPI or Flask

FastAPI (default) Flask (--backbone flask)
Remote responses streamed buffered in memory
Database queries run in worker threads, so a slow query doesn't block other requests run inside the request (asyncio.run)
PostgreSQL supported refused at startup
api-dock start server uvicorn Flask's built-in development server

Use FastAPI unless you need Flask. To run Flask in production, serve create_flask_app(...) with a WSGI server instead of api-dock start.

api-dock start

api-dock start [CONFIG_NAME] [--host 0.0.0.0] [--port 8000] [--backbone fastapi|flask] \
               [--log-level critical|error|warning|info|debug|trace]
  • CONFIG_NAME is a name, not a path: API Dock looks for api_dock_config/<name>.yaml in the working directory, then the bundled examples. The default is config.
  • If the port is taken, the next four ports are tried, and the one used is printed. In a container, make sure nothing else holds the port, so health checks hit the right one.
  • It runs a single uvicorn process. For more processes, run uvicorn yourself, e.g. uvicorn api_dock.fast_api:create_app --factory --workers 2 (this reads api_dock_config/config.yaml from the working directory). Each process has its own max_concurrent_queries limit, PostgreSQL pools and memory.

Docker and pixi

Install dependencies from a lock file so builds are reproducible, and copy the config in:

FROM debian:bookworm-slim
WORKDIR /app
RUN apt-get update && apt-get install -y --no-install-recommends curl ca-certificates \
    && rm -rf /var/lib/apt/lists/*
RUN curl -fsSL https://pixi.sh/install.sh | bash
ENV PATH="/root/.pixi/bin:${PATH}"
COPY pyproject.toml pixi.lock ./
RUN pixi install --locked
COPY api_dock_config/ api_dock_config/
EXPOSE 8080
CMD ["pixi", "run", "api-dock", "start", "--host", "0.0.0.0", "--port", "8080"]

With pip, pin the version (pip install 'api-dock==<version>', or 'api-dock[postgres]==<version>'). Pass secrets (database passwords, API_DOCK_ENCRYPTION_KEY, injected-cookie values) as environment variables, not files baked into the image.

Behind a CDN or proxy path (base_path)

When a CDN or reverse proxy forwards https://app.example.org/dock/* to API Dock without stripping /dock, set:

settings:
  base_path: /dock

/dock/birdnet/latest/detections/ is then handled as /birdnet/latest/detections/. Paths without the prefix still work, so direct calls and health checks are unaffected.

Memory and concurrency (settings.duckdb)

Each database query runs on its own DuckDB connection. Every key under settings.duckdb except max_concurrent_queries is applied as SET <key> = <value>:

settings:
  duckdb:
    memory_limit: 700MB       # per query
    threads: 2
    temp_directory: /tmp/duckdb
    max_concurrent_queries: 2 # others wait their turn
  • memory_limit applies per query. Keep memory_limit × max_concurrent_queries well below the instance's memory, and leave room for Python and the proxy. Otherwise the container can be killed for running out of memory, which clients see as 502s from the load balancer.
  • Measure your heaviest routes' peak memory. Small SQL changes matter: on one deployment, a COUNT(DISTINCT id) count route peaked near 1 GB on a 1 GB instance, while COUNT(*) (with the data already one row per id) peaked around 165 MB.

Health checks and timeouts

  • GET / returns the metadata from the main config and doesn't touch remotes or databases, so it makes a good health check. With FastAPI it stays responsive during slow database queries.
  • settings.timeout (default 10 s) applies to upstream remote calls. A timeout returns 502. Raise it for slow upstreams. Avoid disabling it, because a stalled upstream would hold the connection open. It doesn't limit database queries: use statement_timeout_ms for PostgreSQL and max_concurrent_queries and memory_limit for DuckDB.
  • Set your load balancer's idle timeout above your slowest expected request.

Clone this wiki locally