Models and hardware
Which local model to use on which hardware, what each size is good at, and how GPU acceleration is chosen.
The model ladder
All models are Qwen2.5-Coder instruct variants in GGUF format, licensed Apache-2.0, downloaded from Hugging Face and SHA-256 verified.
| Tier | Model | Download | Needs | Good for |
|---|---|---|---|---|
| Starter (default) | 0.5B · Q8_0 | 645 MB | any CPU, 4 GB RAM | instant autocomplete, quick questions, boilerplate |
| Lite+ | 1.5B · Q4_K_M | 1.07 GB | 8 GB RAM | better multi-line completions, explanations, small refactors |
| Studio | 7B · Q4_K_M | 4.5 GB | 16 GB RAM or 6 GB VRAM | real reasoning, bug finding, tests, cross-file edits |
| Studio | 14B · Q4_K_M | 8.6 GB | 12 GB VRAM | near frontier-class local coding |
| Studio | 32B · Q4_K_M | 19 GB | 24 GB VRAM or 48 GB unified | best open coding model; matches GPT-4o on several benchmarks |
Measured on a 2021 Intel i5-1155G7 laptop (8 threads, 16 GB RAM, no discrete GPU): Starter answers at ~50 tokens/s with a first token in ~150 ms. Lite+ runs at roughly 20-25 tokens/s on the same machine.
Why the default is deliberately small
The first run has to work on the worst laptop in the room, instantly, so the default is the fastest model that still writes correct code for routine tasks. You are told when your hardware can run more: the model switcher marks the recommended size, and the setup output prints an upgrade hint.
Honest expectations
- 0.5B and 1.5B models are excellent at completing what you started and poor at inventing architecture. Give them the beginning of the line, a good function name and types, and they shine.
- They do not know your whole repository (that is coming). Attach the relevant file in chat.
- For complex reasoning, use a Studio model on a GPU, or keep using a cloud model for those tasks and BdebTech for everything you would rather keep private.
How the engine is chosen
On first run the daemon profiles your machine and picks a llama.cpp build:
| Hardware | Engine | Notes |
|---|---|---|
| NVIDIA GPU, 2 GB+ VRAM | CUDA 12.4 | fastest; the CUDA runtime is downloaded alongside |
| Discrete AMD, or Intel Arc, 4 GB+ VRAM | Vulkan | no vendor SDK needed |
| Apple Silicon | Metal | unified memory lets a 16 GB Mac run 7B comfortably |
| Everything else, including Intel/AMD integrated graphics | CPU | the safe default; uses AVX2/AVX-512 when available |
If a GPU engine fails to start, the daemon automatically falls back to the next option and finally to CPU. You can force a choice in %LOCALAPPDATA%\BdebTech\AI\config.json ("engine": "cpu" | "vulkan" | "cuda" | "metal").
Memory use
Rule of thumb: model file size plus ~0.5-1 GB for context and the engine. Starter uses about 1 GB of RAM while loaded. The daemon unloads the model after 30 minutes of inactivity (idleUnloadMinutes in config) and reloads it on the next request in about a second.
Switching models
Status bar → model name → Switch model, or the command BdebTech AI: Switch Model. Models you have not downloaded are fetched on demand with a progress bar. Delete a model via bdebd or by removing its folder under %LOCALAPPDATA%\BdebTech\AI\models.