Local API
Use the bdebd daemon from other tools. OpenAI-compatible chat and completion endpoints, fill-in-the-middle, and management endpoints.
The daemon listens on http://127.0.0.1:41337 (loopback only). Anything under /v1/ plus llama.cpp’s native endpoints are proxied to the engine; the engine is started on demand if it is idle.
Authentication: the API token
Every request needs the per-install token. It is a random 64-character string created on first run and stored in the daemon data directory as a file named token:
| OS | Path |
|---|---|
| Windows | %LOCALAPPDATA%\BdebTech\AI\token |
| macOS | ~/Library/Application Support/BdebTech/AI/token |
| Linux | ~/.local/share/bdebtech/ai/token |
Print it with bdebd token. Send it as X-Bdeb-Token: <token> or, for OpenAI-style clients, as the API key (Authorization: Bearer <token>). Requests without it get 401.
Why: the daemon is loopback-only, but any web page you visit can still fire requests at 127.0.0.1:41337. A browser cannot read the token file, so the token keeps web pages out while every local tool on your machine can read it. Browser-based CORS is granted only to editor webviews. If you really want an open daemon (for example inside a locked-down VM), start it with bdebd serve --no-token or set "requireToken": false in config.json.
OpenAI-compatible
TOKEN=$(bdebd token)
curl http://127.0.0.1:41337/v1/chat/completions \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{"messages":[{"role":"user","content":"Write a Python fizzbuzz"}],"max_tokens":200,"stream":false}'
Supported: /v1/chat/completions (streaming and non-streaming, including response_format with json_schema), /v1/completions, /v1/models, /v1/embeddings (when an embedding model is loaded). The model field is ignored; whatever model is active is used.
Pointing other tools at it
Use the token as the API key everywhere.
| Tool | Setting |
|---|---|
| Cline / Roo Code | Provider: OpenAI Compatible, Base URL http://127.0.0.1:41337/v1, API key = your token |
| Continue | "provider": "openai", "apiBase": "http://127.0.0.1:41337/v1", "apiKey": "<token>" |
| Aider | aider --openai-api-base http://127.0.0.1:41337/v1 --openai-api-key $(bdebd token) --model openai/local |
| Open WebUI | OpenAI API connection with the same base URL and the token as key |
Fill-in-the-middle (autocomplete)
curl http://127.0.0.1:41337/infill -H "X-Bdeb-Token: $TOKEN" -H "Content-Type: application/json" -d '{
"input_prefix": "def add(a, b):\n return ",
"input_suffix": "\n\nprint(add(2, 3))\n",
"n_predict": 32, "temperature": 0.1, "stop": ["\n"]
}'
Returns { "content": "a + b", ... }. This is llama.cpp’s native /infill; the model’s FIM tokens are inserted automatically.
Management endpoints
| Method and path | Purpose |
|---|---|
GET /bdeb/status | Version, hardware profile, engine state, installed models, downloads, license |
GET /bdeb/hw | Hardware profile only |
GET /bdeb/models | Registry with installed, recommended, active, allowed flags |
POST /bdeb/models/pull { "modelId" } | Start a download (202). Progress via events |
POST /bdeb/models/cancel { "modelId" } | Cancel a download |
POST /bdeb/models/use { "modelId" } | Switch the served model (downloads if needed, restarts engine) |
POST /bdeb/models/remove { "modelId" } | Delete model files |
POST /bdeb/engine/start · /stop · /restart | Engine lifecycle |
GET /bdeb/config · POST /bdeb/config | Read or patch daemon config |
GET /bdeb/license · POST /bdeb/license { "key" } · DELETE /bdeb/license | License management |
GET /bdeb/events | Server-sent events: status, download, engine, index, log (token also accepted as ?token= here, for EventSource) |
GET /bdeb/ping | Liveness without a token: { ok, auth } |
POST /bdeb/shutdown | Exit the daemon |
CLI
bdebd serve [--port N] [--verbose] [--no-token]
bdebd setup [--model <id>] [--no-verify]
bdebd status | hw | models | token
bdebd models add <path|url> [--id X] [--name N] [--ctx 8192]
bdebd models remove <model-id>
bdebd pull <model-id> | use <model-id>
bdebd license <key> | license --remove
bdebd stop | version
Model ids
qwen2.5-coder-0.5b-q8 · qwen2.5-coder-1.5b-q4 · qwen2.5-coder-7b-q4 · qwen2.5-coder-14b-q4 · qwen2.5-coder-32b-q4