Models and hardware

Which local model to use on which hardware, what each size is good at, and how GPU acceleration is chosen.

The model ladder

All models are Qwen2.5-Coder instruct variants in GGUF format, licensed Apache-2.0, downloaded from Hugging Face and SHA-256 verified.

TierModelDownloadNeedsGood for
Starter (default)0.5B · Q8_0645 MBany CPU, 4 GB RAMinstant autocomplete, quick questions, boilerplate
Lite+1.5B · Q4_K_M1.07 GB8 GB RAMbetter multi-line completions, explanations, small refactors
Studio7B · Q4_K_M4.5 GB16 GB RAM or 6 GB VRAMreal reasoning, bug finding, tests, cross-file edits
Studio14B · Q4_K_M8.6 GB12 GB VRAMnear frontier-class local coding
Studio32B · Q4_K_M19 GB24 GB VRAM or 48 GB unifiedbest open coding model; matches GPT-4o on several benchmarks

Measured on a 2021 Intel i5-1155G7 laptop (8 threads, 16 GB RAM, no discrete GPU): Starter answers at ~50 tokens/s with a first token in ~150 ms. Lite+ runs at roughly 20-25 tokens/s on the same machine.

Why the default is deliberately small

The first run has to work on the worst laptop in the room, instantly, so the default is the fastest model that still writes correct code for routine tasks. You are told when your hardware can run more: the model switcher marks the recommended size, and the setup output prints an upgrade hint.

Honest expectations

  • 0.5B and 1.5B models are excellent at completing what you started and poor at inventing architecture. Give them the beginning of the line, a good function name and types, and they shine.
  • They do not know your whole repository (that is coming). Attach the relevant file in chat.
  • For complex reasoning, use a Studio model on a GPU, or keep using a cloud model for those tasks and BdebTech for everything you would rather keep private.

How the engine is chosen

On first run the daemon profiles your machine and picks a llama.cpp build:

HardwareEngineNotes
NVIDIA GPU, 2 GB+ VRAMCUDA 12.4fastest; the CUDA runtime is downloaded alongside
Discrete AMD, or Intel Arc, 4 GB+ VRAMVulkanno vendor SDK needed
Apple SiliconMetalunified memory lets a 16 GB Mac run 7B comfortably
Everything else, including Intel/AMD integrated graphicsCPUthe safe default; uses AVX2/AVX-512 when available

If a GPU engine fails to start, the daemon automatically falls back to the next option and finally to CPU. You can force a choice in %LOCALAPPDATA%\BdebTech\AI\config.json ("engine": "cpu" | "vulkan" | "cuda" | "metal").

Memory use

Rule of thumb: model file size plus ~0.5-1 GB for context and the engine. Starter uses about 1 GB of RAM while loaded. The daemon unloads the model after 30 minutes of inactivity (idleUnloadMinutes in config) and reloads it on the next request in about a second.

Switching models

Status bar → model name → Switch model, or the command BdebTech AI: Switch Model. Models you have not downloaded are fetched on demand with a progress bar. Delete a model via bdebd or by removing its folder under %LOCALAPPDATA%\BdebTech\AI\models.