Base URL: http://localhost:15597
- Authentication
- Health & Status
- Chat & LLM
- Conversations
- Messages
- Agents
- User Profile
- Images
- Text-to-Speech
- Speech Recognition
- Tools & MCP Servers
- Skills
- Face Recognition
- Character Assets
- WebSocket
All protected endpoints require a JWT bearer token in the Authorization header:
Authorization: Bearer <token>
JWT tokens use HS256 with 30-day expiry.
Authenticate user and receive JWT token.
Request: application/x-www-form-urlencoded
| Field | Type | Required | Description |
|---|---|---|---|
| username | string | Yes | User's username |
| password | string | Yes | User's password |
Response: 200 OK
{
"access_token": "eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9...",
"token_type": "bearer"
}Error: 400 Bad Request — Incorrect username or password
Create a new user account.
Request: application/x-www-form-urlencoded
| Field | Type | Required | Description |
|---|---|---|---|
| username | string | Yes | Desired username |
| password | string | Yes | Desired password |
Response: 200 OK
{
"status": "ok"
}Error: 400 Bad Request — User already exists
Health check endpoint (no authentication required).
Response: 200 OK
{
"status": "ok",
"service": "llm-hub"
}Real-time streaming chat over WebSocket. See WebSocket section for full protocol details.
Authentication: Query parameter ?token=<JWT>
Client → Server: chat_request event with text, model_name, conversation_id?, agent_id?, images?
If model_name refers to an Ollama model that is not downloaded yet, the server now pulls it automatically on first use before starting generation.
Server → Client: stream_chunk events (sentence-by-sentence), then done event.
List available LLM models from user's Ollama instance.
Authentication: Required
Response: 200 OK
{
"models": ["llama3.2:latest", "mistral:latest", "qwen2.5:7b"]
}List available models with detailed info.
Authentication: Required
Response: 200 OK
{
"models": [
{"name": "llama3.2:latest", "size": 4109853696, "modified_at": "2025-01-15T10:30:00Z"}
]
}Pull/download a model from Ollama registry.
Use this when you want to pre-download a model before selecting it in the client. Chat requests and other server-side LLM calls will also auto-pull missing models on first use.
Authentication: Required
Request: application/json
{"name": "llama3.2:latest"}Response: 200 OK
{"status": "ok", "message": "Model pulled successfully"}Delete a downloaded model.
Authentication: Required
Response: 200 OK
{"status": "ok", "message": "Model deleted successfully"}Ensure model exists, pulling if necessary.
Authentication: Required
Response: 200 OK
{"status": "ok", "message": "Model is available"}List user's conversations. With agent_id, returns the latest conversation containing messages from that agent.
Authentication: Required
Query Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
| limit | integer | 50 | Maximum conversations to return |
| agent_id | integer | null | Filter by agent ID (returns latest conversation for that agent) |
Response: 200 OK
[
{
"id": 1,
"title": "Conversation Title",
"created_at": "2025-01-15T10:30:00Z",
"updated_at": "2025-01-15T11:45:00Z"
}
]Get conversation details with paginated messages and referenced frame metadata.
Authentication: Required
Query Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
| limit | integer | 20 | Messages per page |
| offset | integer | 0 | Pagination offset (from newest messages) |
Pagination: Messages fetched in reverse chronological order, then reversed before return (chronological order in response). Enables infinite scroll.
Response: 200 OK
{
"id": 1,
"title": "Conversation Title",
"created_at": "2025-01-15T10:30:00Z",
"messages": [
{
"id": 1,
"role": "user",
"content": "Hello",
"frame_id": 1,
"created_at": "2025-01-15T10:30:00Z",
"has_raw_data": false
},
{
"id": 2,
"role": "assistant",
"content": "Hello! How can I help you?",
"thinking": "User is greeting me...",
"name": "Assistant",
"agent_id": 2,
"agent": {"id": 2, "name": "Assistant", "avatar_uuid": null, "voice_reference": null},
"frame_id": 1,
"created_at": "2025-01-15T10:30:05Z",
"has_raw_data": true
}
],
"frames": {
"1": {
"id": 1,
"summary": "User greeted the assistant",
"created_at": "2025-01-15T10:30:00Z",
"updated_at": "2025-01-15T10:30:05Z"
}
},
"total_messages": 100,
"offset": 0,
"limit": 20,
"has_more": true
}Update conversation title.
Authentication: Required
Request: application/json
{"title": "New Conversation Title"}Response: 200 OK
{"message": "Conversation title updated successfully"}Delete conversation and all its messages.
Authentication: Required
Response: 200 OK
{"message": "Conversation deleted successfully"}List all frames in a conversation with metadata.
Authentication: Required
Response: 200 OK
{
"frames": [
{"id": 1, "summary": "Discussion about...", "created_at": "2025-01-15T10:30:00Z", "updated_at": "2025-01-15T11:00:00Z"}
]
}Note: Frames are session windows. A new frame is created when the user returns after idle time (FRAME_IDLE_THRESHOLD_MINUTES, default 30). Old frames are summarized asynchronously if User.summary_model is configured.
Get a specific message by ID.
Authentication: Required
Response: 200 OK
{
"id": 1,
"role": "user",
"content": "Message content",
"conversation_id": 1,
"created_at": "2025-01-15T10:30:00Z",
"has_raw_data": false
}Delete a message and all subsequent messages in the conversation (enables conversation branching).
Authentication: Required
Response: 200 OK
{"deleted": 5}Get raw LLM input/output for a message.
Authentication: Required
Response: 200 OK
{
"id": 1,
"raw_input": [{"role": "system", "content": "..."}, {"role": "user", "content": "..."}],
"raw_output": "Full LLM response text"
}List all agents for the current user. Auto-creates default "Administrator" and "Assistant" agents on first call.
Authentication: Required
Response: 200 OK
[
{
"id": 1,
"name": "Assistant",
"system_prompt": "You are a helpful assistant",
"voice_reference": "kurisu_ref",
"avatar_uuid": "550e8400-...",
"model_name": "llama3.2:latest",
"tools": ["play_music", "music_control"],
"think": false,
"character_config": null,
"memory": "User prefers concise answers...",
"trigger_word": "hey kurisu"
}
]Get a specific agent by ID.
Authentication: Required
Response: 200 OK — Same format as list item above.
Create a new agent.
Authentication: Required
Request: application/json
| Field | Type | Required | Description |
|---|---|---|---|
| name | string | Yes | Agent name (cannot be "Administrator" or "User") |
| model_name | string | Yes | LLM model for this agent |
| system_prompt | string | No | Agent's system prompt |
| tools | string[] | No | Opt-in tool names |
| think | boolean | No | Enable chain-of-thought (default: false) |
| trigger_word | string | No | Voice activation trigger word |
Response: 200 OK — Agent object
Update an existing agent. Cannot rename Administrator or change its system prompt/tools.
Authentication: Required
Request: application/json — Any subset of fields:
| Field | Type | Description |
|---|---|---|
| name | string | New name |
| system_prompt | string | New system prompt |
| model_name | string | New LLM model |
| tools | string[] | Updated tool list |
| think | boolean | Enable/disable thinking |
| memory | string | Agent memory text |
| trigger_word | string | Voice activation trigger word |
Response: 200 OK — Updated agent object
Update agent avatar image.
Authentication: Required
Request: multipart/form-data
| Field | Type | Required | Description |
|---|---|---|---|
| avatar | file | Yes | Avatar image file |
Response: 200 OK — Updated agent object
Upload voice reference file for agent TTS synthesis.
Authentication: Required
Request: multipart/form-data
| Field | Type | Required | Description |
|---|---|---|---|
| voice | file | Yes | Audio file (.wav, .mp3, .flac, .ogg) |
Response: 200 OK — Updated agent object
Delete an agent. Cannot delete Administrator.
Authentication: Required
Response: 200 OK
{"message": "Agent deleted successfully"}Detect faces from character pose base images and return cropped avatar candidates.
Authentication: Required
Response: 200 OK
[
{"uuid": "550e8400-...", "pose_id": "a1b2", "score": 0.95}
]Set agent avatar from an existing image UUID.
Authentication: Required
Request: application/json
{"avatar_uuid": "550e8400-..."}Response: 200 OK — Updated agent object
Get current user profile.
Authentication: Required
Response: 200 OK
{
"username": "admin",
"system_prompt": "You are a helpful assistant...",
"preferred_name": "John",
"user_avatar_uuid": "550e8400-...",
"agent_avatar_uuid": "660e8400-...",
"ollama_url": "http://localhost:11434",
"summary_model": "llama3.2:latest"
}Update user profile text fields.
Authentication: Required
Request: application/json
| Field | Type | Description |
|---|---|---|
| system_prompt | string | Custom system prompt for the LLM |
| preferred_name | string | User's preferred name |
| ollama_url | string | Ollama API URL |
| summary_model | string | Model for frame summarization and memory consolidation |
Response: 200 OK
{"status": "ok", "message": "Profile updated successfully"}Update user and/or agent avatar images.
Authentication: Required
Request: multipart/form-data
| Field | Type | Required | Description |
|---|---|---|---|
| user_avatar | file | No | User avatar image |
| agent_avatar | file | No | Agent avatar image |
Response: 200 OK
{
"status": "ok",
"user_avatar_uuid": "550e8400-...",
"agent_avatar_uuid": "660e8400-..."
}Upload an image and receive a UUID.
Authentication: Required
Request: multipart/form-data
| Field | Type | Required | Description |
|---|---|---|---|
| file | file | Yes | Image file to upload |
Response: 200 OK
{
"image_uuid": "550e8400-...",
"url": "/images/550e8400-..."
}Retrieve a user-scoped chat image (requires auth via header or ?token= query param).
Authentication: Required (Bearer token in header or token query parameter)
Response: JPEG image file with 1-year cache.
Retrieve an uploaded image (public, no auth required).
Response: Image file with 1-year cache:
Cache-Control: public, max-age=31536000, immutable
Synthesize speech from text.
Authentication: Required
Request: application/json
| Field | Type | Required | Description |
|---|---|---|---|
| text | string | Yes | Text to synthesize |
| voice | string | No | Voice name from /tts/voices |
| language | string | No | Language code (e.g., "en", "ja") |
| provider | string | No | TTS provider (default: TTS_PROVIDER env var or "vixtts") |
Provider-Specific Parameters:
GPT-SoVITS:
| Field | Type | Default | Description |
|---|---|---|---|
| max_chunk_length | integer | 200 | Maximum characters per chunk |
| text_split_method | string | "cut5" | Text splitting method |
| batch_size | integer | 20 | Batch size |
viXTTS:
| Field | Type | Default | Description |
|---|---|---|---|
| emo_audio | string | null | Voice name for emotion reference audio |
| emo_vector | float[8] | null | Emotion vector [happy, angry, sad, afraid, disgusted, melancholic, surprised, calm] |
| emo_text | string | null | Text description for emotion control |
| use_emo_text | boolean | false | Use emotion from text |
| emo_alpha | float | 1.0 | Emotion strength (0.0-1.0) |
Response: 200 OK — Audio file (audio/wav)
List available TTS voices (scans data/voice_storage/).
Authentication: Required
Query Parameters:
| Parameter | Type | Description |
|---|---|---|
| provider | string | Filter by TTS provider (optional) |
Response: 200 OK
{"voices": ["ayaka_ref", "kurisu_ref"]}Check if a TTS server is reachable.
Authentication: Required
Request: application/json
{"provider": "vixtts"}Response: 200 OK — Server health response
List available TTS backends.
Authentication: Required
Response: 200 OK
{"backends": ["gpt-sovits", "vixtts"]}Convert audio to text using faster-whisper (CTranslate2).
Authentication: Required
Request: Raw Int16 PCM bytes at 16kHz mono (application/octet-stream)
Query Parameters:
| Parameter | Type | Description |
|---|---|---|
| language | string | Language hint (e.g., "en", "zh") |
Response: 200 OK
{"text": "transcribed text content"}List all available tools (MCP + built-in).
Authentication: Required
Response: 200 OK
{
"mcp_tools": [...],
"builtin_tools": [...]
}Built-in tools (always available): search_messages, get_conversation_info, get_frame_summaries, get_frame_messages, get_skill_instructions
Opt-in tools (added to agent's tools array): play_music, music_control, get_music_queue, route_to_agent, route_to_user, MCP tools
List configured MCP servers and their status.
Authentication: Required
Response: 200 OK
{
"servers": [
{"name": "web_search", "url": "http://web-search-container:8000", "status": "available"}
]
}Status values: configured, available, unavailable
List user's skills.
Authentication: Required
Response: 200 OK
[
{"id": 1, "name": "Music Player", "instructions": "When the user asks to play music...", "created_at": "2025-01-15T10:30:00Z"}
]Create a new skill (name unique per user).
Authentication: Required
Request: application/json
{"name": "Music Player", "instructions": "When the user asks to play music..."}Response: 200 OK — Skill object
Update a skill.
Authentication: Required
Request: application/json
{"name": "Updated Name", "instructions": "Updated instructions..."}Response: 200 OK — Updated skill object
Delete a skill.
Authentication: Required
Response: 200 OK
{"deleted": true}List registered face identities with photo counts.
Authentication: Required
Response: 200 OK
[
{"id": 1, "name": "John", "photo_count": 3, "created_at": "2025-01-15T10:30:00Z"}
]Register a new face identity. Detects face in photo, computes 512-dim embedding.
Authentication: Required
Request: multipart/form-data with query param name
| Field | Type | Required | Description |
|---|---|---|---|
| name | query string | Yes | Identity name (unique per user) |
| photo | file | Yes | Photo containing the face |
Response: 200 OK
{
"id": 1,
"name": "John",
"photo": {"id": 1, "photo_uuid": "550e8400-...", "url": "/faces/1/photos/1/image"}
}Error: 400 Bad Request — No face detected in photo
Get face identity details with all photos.
Authentication: Required
Response: 200 OK
{
"id": 1,
"name": "John",
"created_at": "2025-01-15T10:30:00Z",
"photos": [
{"id": 1, "photo_uuid": "550e8400-...", "url": "/faces/1/photos/1/image", "created_at": "2025-01-15T10:30:00Z"}
]
}Delete face identity, all photos, and disk images.
Authentication: Required
Response: 200 OK
{"status": "deleted"}Add additional photo to existing identity.
Authentication: Required
Request: multipart/form-data
| Field | Type | Required | Description |
|---|---|---|---|
| photo | file | Yes | Photo containing the face |
Response: 200 OK
{"id": 2, "photo_uuid": "660e8400-...", "url": "/faces/1/photos/2/image"}Remove a specific photo from a face identity.
Authentication: Required
Response: 200 OK
{"status": "deleted"}Serve face photo image file.
Authentication: Required
Response: Image file
Upload base portrait image for character animation.
Authentication: Required
Query Parameters: agent_id (int), pose_id (string)
Request: multipart/form-data
| Field | Type | Required | Description |
|---|---|---|---|
| file | file | Yes | Base portrait image |
Response: 200 OK
{"asset_id": "2/a1b2/base.png", "image_url": "/character-assets/2/a1b2/base.png"}Upload keyframe image and compute diff patch against the pose's base image.
Authentication: Required
Query Parameters: agent_id (int), pose_id (string), part (left_eye|right_eye|mouth), index (int)
Request: multipart/form-data
| Field | Type | Required | Description |
|---|---|---|---|
| keyframe | file | Yes | Keyframe image |
Response: 200 OK
{"patch": {"image_url": "/character-assets/2/a1b2/mouth_0.png", "x": 100, "y": 200, "width": 50, "height": 30}}Upload transition video for an animation edge.
Authentication: Required
Query Parameters: agent_id (int), edge_id (string)
Request: multipart/form-data
| Field | Type | Required | Description |
|---|---|---|---|
| file | file | Yes | Video file (mp4 or webm) |
Response: 200 OK
{"asset_id": "2/edges/e1f2.mp4", "video_url": "/character-assets/2/edges/e1f2.mp4"}Update character animation config (pose tree). Auto-cleans up orphaned assets.
Authentication: Required
Request: application/json — Character config with pose_tree
Response: 200 OK
{"message": "Character config updated", "character_config": {...}}Rename asset files/folders on disk to match migrated IDs.
Authentication: Required
Request: application/json
{"id_mapping": {"old_pose_id": "new_pose_id", "old_edge_id": "new_edge_id"}}Response: 200 OK
{"message": "IDs migrated successfully"}Serve pose asset (base or patch image). No authentication, no cache.
Response: Image file
Serve transition video. No authentication, no cache.
Response: Video file (mp4 or webm)
Authentication: Query parameter ?token=<JWT>
Persistent connection for real-time chat, vision, and media control. All events are JSON objects.
chat_request — Send a message
{
"type": "chat_request",
"text": "Hello",
"model_name": "",
"conversation_id": 1,
"agent_id": 2,
"images": ["base64..."]
}model_name may reference a model that is not downloaded yet. The server will try to pull it automatically from the configured Ollama instance before generation starts.
cancel — Cancel current streaming response
{"type": "cancel"}tool_approval_response — Approve/deny a tool execution
{
"type": "tool_approval_response",
"approval_id": "abc123",
"approved": true,
"modified_args": {}
}vision_start — Start vision processing
{"type": "vision_start", "enable_face": true, "enable_pose": false, "enable_hands": false}vision_frame — Send webcam frame for processing
{"type": "vision_frame", "frame": "base64_jpeg_data"}vision_stop — Stop vision processing
{"type": "vision_stop"}Media control events: media_play (query), media_pause, media_resume, media_skip, media_stop, media_queue_add (query), media_queue_remove (index), media_volume (volume)
stream_chunk — Streaming chat content
{
"type": "stream_chunk",
"content": "Hello ",
"thinking": "",
"role": "assistant",
"agent_id": 2,
"name": "Assistant",
"voice_reference": "kurisu_ref",
"conversation_id": 1,
"frame_id": 1
}done — Streaming complete
{"type": "done", "conversation_id": 1, "frame_id": 1}error — Error occurred
{"type": "error", "error": "Error message", "code": 500}agent_switch — Agent routing change (group mode)
{
"type": "agent_switch",
"from_agent_id": 1,
"from_agent_name": "Assistant",
"to_agent_id": 3,
"to_agent_name": "Coder",
"reason": "Routing to specialized agent"
}tool_approval_request — Tool requires user approval
{
"type": "tool_approval_request",
"approval_id": "abc123",
"tool_name": "play_music",
"tool_args": {"query": "lofi beats"},
"description": "Play music from YouTube",
"risk_level": "low"
}vision_result — Vision processing results
{
"type": "vision_result",
"faces": [{"name": "John", "confidence": 0.95, "bbox": [x, y, w, h]}],
"gestures": ["wave", "thumbs_up"]
}Chat WebSocket supports replay: _accumulated_messages (complete messages) replayed on reconnect. Client filters by conversation ID.
All endpoints may return standard HTTP error responses:
| Status Code | Description |
|---|---|
| 400 | Bad Request — Invalid input |
| 401 | Unauthorized — Invalid or missing token |
| 404 | Not Found — Resource does not exist |
| 500 | Internal Server Error — Server-side error |
{"detail": "Error message describing what went wrong"}| Field | Type | Description |
|---|---|---|
| id | integer | Primary key |
| username | string | Unique identifier |
| password | string | Bcrypt-hashed password |
| system_prompt | string | Custom system prompt for the LLM |
| preferred_name | string | User's preferred name |
| user_avatar_uuid | string | UUID of user's avatar image |
| agent_avatar_uuid | string | UUID of default agent avatar |
| ollama_url | string | Ollama API URL |
| summary_model | string | Model for frame summarization and memory consolidation |
| Field | Type | Description |
|---|---|---|
| id | integer | Primary key |
| user_id | integer | Foreign key to User |
| title | string | Conversation title |
| created_at | datetime | Creation timestamp |
| updated_at | datetime | Last update timestamp |
| Field | Type | Description |
|---|---|---|
| id | integer | Primary key |
| conversation_id | integer | Foreign key to Conversation |
| summary | string | Auto-generated session summary (nullable) |
| created_at | datetime | Creation timestamp |
| updated_at | datetime | Last update timestamp |
| Field | Type | Description |
|---|---|---|
| id | integer | Primary key |
| role | string | "user", "assistant", or "tool" |
| message | string | Message content |
| thinking | string | Chain-of-thought reasoning (nullable) |
| raw_input | JSON | Raw LLM input messages (nullable) |
| raw_output | string | Raw LLM response (nullable) |
| name | string | Speaker name (nullable) |
| frame_id | integer | Foreign key to Frame |
| agent_id | integer | Foreign key to Agent (SET NULL on delete) |
| created_at | datetime | Creation timestamp |
| Field | Type | Description |
|---|---|---|
| id | integer | Primary key |
| user_id | integer | Foreign key to User |
| name | string | Agent name |
| system_prompt | string | Agent's system prompt |
| voice_reference | string | Voice file name for TTS (nullable) |
| avatar_uuid | string | UUID of agent's avatar image (nullable) |
| model_name | string | LLM model name |
| tools | JSON | Array of opt-in tool names |
| think | boolean | Enable chain-of-thought |
| memory | text | Auto-consolidated agent memory (nullable) |
| trigger_word | string | Voice activation trigger word (nullable) |
| created_at | datetime | Creation timestamp |
| Field | Type | Description |
|---|---|---|
| id | integer | Primary key |
| user_id | integer | Foreign key to User |
| name | string | Identity name (unique per user) |
| created_at | datetime | Creation timestamp |
| Field | Type | Description |
|---|---|---|
| id | integer | Primary key |
| identity_id | integer | Foreign key to FaceIdentity (CASCADE) |
| embedding | vector(512) | Face embedding (pgvector) |
| photo_uuid | string | UUID of photo image |
| created_at | datetime | Creation timestamp |
| Field | Type | Description |
|---|---|---|
| id | integer | Primary key |
| user_id | integer | Foreign key to User |
| name | string | Skill name (unique per user) |
| instructions | text | Skill instructions injected into agent prompts |
| created_at | datetime | Creation timestamp |
Conversations are auto-created on first message:
- Send a
chat_requestvia WebSocket withconversation_id: null - Backend auto-creates conversation and frame
- First
stream_chunkevent includes the newconversation_idandframe_id
Frames are session windows managed automatically:
- New frame created when user returns after idle time (default: 30 minutes)
- Old frames summarized asynchronously if
User.summary_modelis configured - LLM only sees messages from the current frame
- Built-in tools (
get_frame_summaries,get_frame_messages) let the LLM access past context
- Images sent as base64 in
chat_requestWebSocket events - Saved to per-user directory (
data/image_storage/data/users/{user_id}/) and assigned UUIDs - UUIDs stored in message's
imagesJSON column and streamed back to client viaStreamChunkEvent.images - Base64 images passed to LLM via Ollama's
imageskey for vision model support - Served via auth-required
GET /images/u/{uuid}with 1-year cache - MCP tools returning
ImageContentblocks are also saved and attached to tool result messages