Repository awareness (@codebase)
Index your workspace locally so chat can answer about your whole project and autocomplete can use types and helpers from other files.
Small models are only as good as the context they see. Repository awareness gives them the right context automatically: the IDE builds a local semantic index of your workspace and pulls the most relevant snippets into every chat question and every autocomplete request.
What it does
- Chat
@codebase- with the toggle on (it is on by default once a workspace is indexed), your question is matched against every indexed chunk and the top results are attached. Ask “where do we verify license signatures?” or “how is the daemon token read?” and the answer cites the right files and lines. - Repository-level autocomplete - the code around your cursor is used to retrieve two related snippets from other files, which are passed to the model as extra context (Qwen2.5-Coder was trained for this with
<|repo_name|>/<|file_sep|>tokens). The practical effect: completions start using your function names, types and helpers instead of inventing plausible ones.
Example from this project’s own code. Cursor after return store. in a new file:
| Completion | |
|---|---|
| Without index | topHits(query); (invented) |
| With index | search(query, 10); (the real VectorStore.search method) |
How it works, locally
- The first time you open a folder, the IDE asks whether to index it. Saying yes downloads a small embedding model once (Nomic Embed Text v1.5, 146 MB, Apache-2.0).
- The daemon walks the folder, honouring
.gitignoreplus a built-in list (dependencies, build output, binaries, lockfiles, secrets like.envand keys). Files over 512 KB and data-heavy files (hashes, minified bundles) are skipped. - Files are split into chunks of roughly 60 lines at natural boundaries (function and class starts), embedded on your CPU, and stored in a SQLite database under
%LOCALAPPDATA%\BdebTech\AI\index\. Nothing leaves the machine. - Saving a file re-indexes just that file in the background.
Speed and size
Indexing is a one-time background job. On a 2021 laptop CPU it runs at about 2 chunks per second: a 150-file project (about 600 chunks) takes 4-5 minutes, a 1,000-file project 20-30 minutes. After that, each save costs well under a second. Searching is instant (about 30 ms over 30,000 chunks).
The index is capped at 30,000 chunks per workspace (indexMaxChunks in config.json); roughly 5,000-8,000 source files. Disk use is about 3 KB per chunk.
If you have a GPU the chat model uses it; the embedding model deliberately stays on CPU so it never competes for GPU memory.
Controls
| Action | Where |
|---|---|
| Index / re-index | status bar menu → Index workspace, or BdebTech AI: Index Workspace for Repository-Aware AI |
| Remove the index | BdebTech AI: Clear Workspace Index |
Toggle @codebase for a chat message | checkbox above the chat box |
| Turn off repo context for autocomplete | setting bdebtech.completions.repoContext |
| Number of snippets | bdebtech.index.chatResults (chat, default 6), bdebtech.completions.repoContextChunks (autocomplete, default 2) |
| Exclude paths | add them to .gitignore or to a .bdebignore file in the workspace root (same syntax) |
| Never ask for this folder | choose Never for this folder in the prompt |
Privacy notes
The index contains your source text and its embeddings. It lives only in your profile directory and is deleted with Clear Workspace Index or when you uninstall and choose to delete data. Secrets are excluded by default (.env*, *.pem, *.key, id_rsa*), but review .bdebignore for anything project-specific.