Local API

Use the bdebd daemon from other tools. OpenAI-compatible chat and completion endpoints, fill-in-the-middle, and management endpoints.

The daemon listens on http://127.0.0.1:41337 (loopback only). Anything under /v1/ plus llama.cpp’s native endpoints are proxied to the engine; the engine is started on demand if it is idle.

Authentication: the API token

Every request needs the per-install token. It is a random 64-character string created on first run and stored in the daemon data directory as a file named token:

OSPath
Windows%LOCALAPPDATA%\BdebTech\AI\token
macOS~/Library/Application Support/BdebTech/AI/token
Linux~/.local/share/bdebtech/ai/token

Print it with bdebd token. Send it as X-Bdeb-Token: <token> or, for OpenAI-style clients, as the API key (Authorization: Bearer <token>). Requests without it get 401.

Why: the daemon is loopback-only, but any web page you visit can still fire requests at 127.0.0.1:41337. A browser cannot read the token file, so the token keeps web pages out while every local tool on your machine can read it. Browser-based CORS is granted only to editor webviews. If you really want an open daemon (for example inside a locked-down VM), start it with bdebd serve --no-token or set "requireToken": false in config.json.

OpenAI-compatible

TOKEN=$(bdebd token)
curl http://127.0.0.1:41337/v1/chat/completions \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"messages":[{"role":"user","content":"Write a Python fizzbuzz"}],"max_tokens":200,"stream":false}'

Supported: /v1/chat/completions (streaming and non-streaming, including response_format with json_schema), /v1/completions, /v1/models, /v1/embeddings (when an embedding model is loaded). The model field is ignored; whatever model is active is used.

Pointing other tools at it

Use the token as the API key everywhere.

ToolSetting
Cline / Roo CodeProvider: OpenAI Compatible, Base URL http://127.0.0.1:41337/v1, API key = your token
Continue"provider": "openai", "apiBase": "http://127.0.0.1:41337/v1", "apiKey": "<token>"
Aideraider --openai-api-base http://127.0.0.1:41337/v1 --openai-api-key $(bdebd token) --model openai/local
Open WebUIOpenAI API connection with the same base URL and the token as key

Fill-in-the-middle (autocomplete)

curl http://127.0.0.1:41337/infill -H "X-Bdeb-Token: $TOKEN" -H "Content-Type: application/json" -d '{
  "input_prefix": "def add(a, b):\n    return ",
  "input_suffix": "\n\nprint(add(2, 3))\n",
  "n_predict": 32, "temperature": 0.1, "stop": ["\n"]
}'

Returns { "content": "a + b", ... }. This is llama.cpp’s native /infill; the model’s FIM tokens are inserted automatically.

Management endpoints

Method and pathPurpose
GET /bdeb/statusVersion, hardware profile, engine state, installed models, downloads, license
GET /bdeb/hwHardware profile only
GET /bdeb/modelsRegistry with installed, recommended, active, allowed flags
POST /bdeb/models/pull { "modelId" }Start a download (202). Progress via events
POST /bdeb/models/cancel { "modelId" }Cancel a download
POST /bdeb/models/use { "modelId" }Switch the served model (downloads if needed, restarts engine)
POST /bdeb/models/remove { "modelId" }Delete model files
POST /bdeb/engine/start · /stop · /restartEngine lifecycle
GET /bdeb/config · POST /bdeb/configRead or patch daemon config
GET /bdeb/license · POST /bdeb/license { "key" } · DELETE /bdeb/licenseLicense management
GET /bdeb/eventsServer-sent events: status, download, engine, index, log (token also accepted as ?token= here, for EventSource)
GET /bdeb/pingLiveness without a token: { ok, auth }
POST /bdeb/shutdownExit the daemon

CLI

bdebd serve [--port N] [--verbose] [--no-token]
bdebd setup [--model <id>] [--no-verify]
bdebd status | hw | models | token
bdebd models add <path|url> [--id X] [--name N] [--ctx 8192]
bdebd models remove <model-id>
bdebd pull <model-id> | use <model-id>
bdebd license <key> | license --remove
bdebd stop | version

Model ids

qwen2.5-coder-0.5b-q8 · qwen2.5-coder-1.5b-q4 · qwen2.5-coder-7b-q4 · qwen2.5-coder-14b-q4 · qwen2.5-coder-32b-q4