← All comparisons
LemonCrow vs CodeGraph

What CodeGraph says, vs. what it scored.

Local SQLite knowledge graph of symbols, call edges, and dependencies, built via tree-sitter and queried over MCP.

What CodeGraph says about itself
“Surgical context, fewer tool calls, faster answers, 100% local.”
“CodeGraph hands the agent the exact code it needs in one call. It's a pre-built knowledge graph of every symbol, call edge, and dependency in your codebase.”
Publishes some numbers...never against another search toolView source ↗
What it actually scored — same 14 repos, same 7,213 queries as every other tool
ToolMRRp95p100
★ LemonCrow +semantic (BGE)0.727390ms1057ms
★ LemonCrow lexical (default)0.676134ms319ms
CodeGraph0.29617ms532ms

Own numbers real, fair with/without comparison. Fastest raw latency measured -- 17ms p95. 0.296 MRR vs. LemonCrow's 0.727.

By query kind -- same benchmark, broken out (no reps in this eval: one deterministic pass per query)
KindLemonCrow +semanticLemonCrow lexicalCodeGraph
definition0.873 (n=1570)0.871 (n=1570)0.766 (n=1570)
content0.873 (n=1444)0.864 (n=1444)0.000 (n=1444)
semantic0.759 (n=1800)0.576 (n=1800)0.051 (n=1800)
swebench0.500 (n=1908)0.493 (n=1908)0.364 (n=1908)
sessions0.587 (n=491)0.571 (n=491)0.299 (n=491)

n = query/gold pairs of that kind, out of 7,213 total -- every provider scored on all 5 kinds.

By repo -- all 15 repos in the corpus, same query set
RepoLemonCrow +semanticLemonCrow lexicalCodeGraph
astropy/astropy0.7720.7150.310
django/django0.6890.6520.286
lemoncrow-lab/lemoncrow-dev0.4670.4770.304
lemoncrow/lemoncrow0.5940.5570.294
matplotlib/matplotlib0.8010.7470.279
mwaskom/seaborn0.8140.7680.331
pallets/flask0.7350.6710.304
psf/requests0.8400.8030.373
pydata/xarray0.8150.7640.354
pylint-dev/pylint0.8560.7840.359
pytest-dev/pytest0.8260.7390.322
scikit-learn/scikit-learn0.7400.6690.264
sphinx-doc/sphinx0.6370.5800.225
sympy/sympy0.6940.6370.283
torvalds/linux0.7260.6680.158

MRR per repo: n-weighted blend across all 5 query kinds, same 7,213-query run.

Cost, head-to-head -- a different suite: 7 repos, 5 reps each, claude-opus-4-8
RepoCodeGraphvs baseline
Tokio$3.4428.1% pricier
Alamofire$2.4848.6% cheaper
Django$2.32even
OkHttp$1.3515.6% cheaper
VS Code$2.5616.6% cheaper
Gin$1.3624.9% pricier
Excalidraw$2.4729.6% cheaper
Total$15.9916.3% cheaper

Measured on LemonCrow's own 7-repo, 5-rep exploration-task suite (claude-opus-4-8), not CodeGraph's self-published with/without numbers. Cheaper than baseline on 5 of 7 repos (16-49%), pricier on Tokio (+28%) and Gin (+25%), and 22.7% slower overall despite the lower total cost -- a noisier result than CodeGraph's own published claims, which show it uniformly faster and never pricier on this same repo set.

Baseline total: $19.11. Wall-clock: 22.7% slower than baseline overall. Full per-repo detail, transcripts, and methodology → BENCHMARKS.md.

The true story

Same 14 repositories, same 7,213 query/gold pairs as every tool here, CodeGraph included. Full methodology, every raw number, and the other9 tools →