← All comparisons
LemonCrow vs Serena

What Serena says, vs. what it scored.

LSP-wrapped semantic code toolkit -- symbol-level navigation and refactoring via real language servers, 40+ languages.

What Serena says about itself
“Serena provides essential semantic code retrieval, editing, refactoring and debugging tools that are akin to an IDE's capabilities, operating at the symbol level and exploiting relational structure.”
“Practically, this means that your agent operates faster, more efficiently and more reliably, especially in larger and more complex codebases.”
Publishes some numbers...never against another search toolView source ↗
What it actually scored — same 14 repos, same 7,213 queries as every other tool
ToolMRRp95p100
★ LemonCrow +semantic (BGE)0.727390ms1057ms
★ LemonCrow lexical (default)0.676134ms319ms
Serena0.4013834ms269001ms

Real symbol-level LSP tooling, 40+ languages. Cold LSP spin-up shows in latency here (3834ms p95 vs. LemonCrow's 134-390ms). 0.401 MRR.

By query kind -- same benchmark, broken out (no reps in this eval: one deterministic pass per query)
KindLemonCrow +semanticLemonCrow lexicalSerena
definition0.873 (n=1570)0.871 (n=1570)0.633 (n=1570)
content0.873 (n=1444)0.864 (n=1444)0.757 (n=1444)
semantic0.759 (n=1800)0.576 (n=1800)0.005 (n=1800)
swebench0.500 (n=1908)0.493 (n=1908)0.332 (n=1908)
sessions0.587 (n=491)0.571 (n=491)0.339 (n=491)

n = query/gold pairs of that kind, out of 7,213 total -- every provider scored on all 5 kinds.

By repo -- all 15 repos in the corpus, same query set
RepoLemonCrow +semanticLemonCrow lexicalSerena
astropy/astropy0.7720.7150.490
django/django0.6890.6520.460
lemoncrow-lab/lemoncrow-dev0.4670.4770.316
lemoncrow/lemoncrow0.5940.5570.357
matplotlib/matplotlib0.8010.7470.000
mwaskom/seaborn0.8140.7680.456
pallets/flask0.7350.6710.363
psf/requests0.8400.8030.482
pydata/xarray0.8150.7640.464
pylint-dev/pylint0.8560.7840.541
pytest-dev/pytest0.8260.7390.545
scikit-learn/scikit-learn0.7400.6690.364
sphinx-doc/sphinx0.6370.5800.306
sympy/sympy0.6940.6370.392
torvalds/linux0.7260.6680.428

MRR per repo: n-weighted blend across all 5 query kinds, same 7,213-query run.

The true story

Same 14 repositories, same 7,213 query/gold pairs as every tool here, Serena included. Full methodology, every raw number, and the other9 tools →