01 // Latency
Sub-second Latency
6.3 ms
Median semantic search over 1,219 code chunks, model warm. Queries replay a captured CUDA graph instead of launching ~2,500 kernels one by one: 43.5 ms before, 6.3 ms now.
p95 8.6 ms
Empower your workflows with Small Language Models. Zero cloud latency. Boosted productivity.
A small model for each specific job, running on your own machine, instead of a frontier model for everything.
01 // Latency
6.3 ms
Median semantic search over 1,219 code chunks, model warm. Queries replay a captured CUDA graph instead of launching ~2,500 kernels one by one: 43.5 ms before, 6.3 ms now.
p95 8.6 ms
02 // Privacy
Loopback only
The model loads from a local folder with offline mode forced. The process refuses every non-loopback connection and DNS lookup.
outbound network blocked
03 // Cost
29× less context
A 271M-parameter model on a laptop GPU finds the relevant lines, so the frontier model reads 4,039 tokens instead of 119,009.
peak VRAM 2.5 GB
codesearch: the local EmbeddingGemma2 MCP server that serves semantic code search to Claude Code today. One run on 2026-10-09.
| Step | Result | Detail |
|---|---|---|
| Cold start, once per session | 10.6 s | imports 7.2 s · model load 2.6 s · first encode 0.6 s · CUDA graph capture 0.3 s |
| Full index, 97-file repo | 58.9 s | 1,219 chunks · 804k tokens · 13.6k tokens/s (10.7k–20.4k across runs) |
| Refresh, nothing changed | 43 ms | mtime + size check per file; only changed files are re-embedded |
| Search, warm | 6.3 ms | p95 8.6 ms · query embedding 6.1 ms (eager PyTorch: 43.5 ms) |
| Context sent to Claude | 4,039 tokens | vs 119,009 reading the same hit files whole · mean of 6 queries |
| Peak VRAM | 2.5 GB | bf16 · batches capped at 16,384 tokens |
EmbeddingGemma2 text encoder (271M params, bf16, 768-d) on an RTX 5060 Laptop GPU (8 GB), Windows 11. The laptop GPU shares power with the CPU and was power-capped during runs, which is why indexing throughput varies between runs at identical settings. Tokens counted with the Gemma tokenizer. Raw results: bench.json.
01 // CodeGraph
An instant, visual representation of your entire codebase. CodeGraph maps dependencies and logical flows into a navigable structure, allowing developers to comprehend architecture at a glance and enabling AI agents to surgically retrieve context rather than guessing.
importsimported by
Code
Info
Loading graph…
Use case // Developers
Instant onboarding and visual debugging: see the entry point, the dependency layers and what a change touches before opening a file. In the preview, select config.py: a change there reaches all 8 modules that depend on it.
Use case // Agents
Structured repository context for local agents and MCP clients. codesearch already hands Claude Code ranked line ranges instead of whole files; CodeGraph adds which modules each hit depends on, so an agent pulls that slice and nothing else.
How it runs todaycodesearch MCP