← All comparisons
The DIY alternative, not a tool

What Caveman gets you, vs. LemonCrow's runtime.

Caveman: a one-line terse-persona system prompt, no install. Tested head-to-head against LemonCrow's runtime on the same 20-prompt Telegraphic Q&A set.

The instruction itself
“Respond terse like smart caveman. All technical substance stay. Only fluff die.”
Free, zero install 48% avg output-token cut, 18pp stdev -- consistent Compresses replies only -- costs 3.3% more than baseline this run, LemonCrow wins on both cost and tokens
All 20 prompts -- in/out tokens, turns, cost -- median of 5 reps, same run as the flagship page (LemonCrow column highlighted)

"In" = summed prompt tokens (input + cache read + cache write). Split in BENCHMARKS.md.

Prompt Baseline LemonCrow Caveman
React re-render
object prop
20,640 in / 873 out
2 turns · $ 0.065
10,383 in / 254 out
1 turns · $ 0.025 ✓
22,628 in / 268 out
2 turns · $ 0.070
Express JWT expiry bug
benchmarks
62,292 in / 1,926 out
4 turns · $ 0.148
21,263 in / 492 out
2 turns · $ 0.070 ✓
22,655 in / 694 out
2 turns · $ 0.081
Postgres connection pool setup
benchmarks
20,632 in / 1,898 out
2 turns · $ 0.091
10,823 in / 791 out
1 turns · $ 0.040 ✓
22,619 in / 1,044 out
2 turns · $ 0.089
git rebase vs merge
benchmarks
20,633 in / 1,044 out
2 turns · $ 0.070
10,863 in / 452 out
1 turns · $ 0.032 ✓
22,618 in / 675 out
2 turns · $ 0.080
Callback -> async/await refactor
benchmarks
20,789 in / 553 out
2 turns · $ 0.058
10,950 in / 133 out
1 turns · $ 0.024 ✓
45,040 in / 393 out
3 turns · $ 0.114
Split a monolith into microservices
benchmarks
20,655 in / 1,461 out
2 turns · $ 0.080
10,446 in / 646 out
1 turns · $ 0.037 ✓
22,641 in / 997 out
2 turns · $ 0.088
PR security review
benchmarks
20,726 in / 1,034 out
2 turns · $ 0.070
10,485 in / 244 out
1 turns · $ 0.025 ✓
22,713 in / 569 out
2 turns · $ 0.078
Multi-stage Dockerfile
benchmarks
41,639 in / 1,419 out
3 turns · $ 0.126
34,109 in / 885 out
3 turns · $ 0.097
22,627 in / 692 out
2 turns · $ 0.081 ✓
Postgres counter race condition
benchmarks
20,656 in / 1,277 out
2 turns · $ 0.076
10,797 in / 369 out
1 turns · $ 0.030 ✓
22,644 in / 641 out
2 turns · $ 0.080
React error boundary component
benchmarks
87,770 in / 2,649 out
5 turns · $ 0.192 ✓
75,128 in / 1,246 out
6 turns · $ 0.248
115,138 in / 2,946 out
6 turns · $ 0.233
Why re-render on parent update?
eval
20,605 in / 779 out
2 turns · $ 0.063
10,463 in / 305 out
1 turns · $ 0.029 ✓
22,592 in / 364 out
2 turns · $ 0.072
Explain connection pooling
eval
20,591 in / 1,340 out
2 turns · $ 0.077
10,801 in / 434 out
1 turns · $ 0.030 ✓
22,575 in / 246 out
2 turns · $ 0.069
TCP vs UDP
eval
20,597 in / 823 out
2 turns · $ 0.044
10,460 in / 133 out
1 turns · $ 0.023 ✓
22,583 in / 325 out
2 turns · $ 0.071
Node.js memory leak
eval
20,612 in / 1,281 out
2 turns · $ 0.075
10,640 in / 447 out
1 turns · $ 0.032 ✓
22,599 in / 715 out
2 turns · $ 0.081
SQL EXPLAIN
eval
20,602 in / 963 out
2 turns · $ 0.068
10,948 in / 266 out
1 turns · $ 0.028 ✓
22,589 in / 477 out
2 turns · $ 0.075
Hash table collisions
eval
20,595 in / 890 out
2 turns · $ 0.066
10,631 in / 225 out
1 turns · $ 0.025 ✓
22,582 in / 453 out
2 turns · $ 0.074
CORS errors
eval
20,603 in / 884 out
2 turns · $ 0.066
10,903 in / 326 out
1 turns · $ 0.029 ✓
22,591 in / 545 out
2 turns · $ 0.077
Debounce a search input
eval
20,605 in / 704 out
2 turns · $ 0.061
10,946 in / 217 out
1 turns · $ 0.027 ✓
22,592 in / 323 out
2 turns · $ 0.071
git rebase vs merge (eval)
eval
20,597 in / 780 out
2 turns · $ 0.063
10,459 in / 182 out
1 turns · $ 0.025 ✓
22,584 in / 384 out
2 turns · $ 0.073
Queue vs topic
eval
20,606 in / 918 out
2 turns · $ 0.066
10,465 in / 251 out
1 turns · $ 0.029 ✓
22,594 in / 381 out
2 turns · $ 0.073
Average, all 20
29,643 in / 1,187 out
2.4 turns · $0.084
16,031 in / 464 out
1.43 turns · $0.045 ✓
27,442 in / 695 out
2.21 turns · $0.087

marks the cheapest arm on that exact prompt. LemonCrow wins 18 of 20. Full per-suite cost/turn breakdown in the flagship page.

The true story

Same prompts, same model, same 5-rep methodology as every other number on this site -- baseline, LemonCrow, caveman, one run. Raw data, all 20 prompts, every arm →