{"measured_at":"2026-07-10","dataset":{"rugs":19,"legit_tokens":26,"note":"N is small — see limitations."},"route_under_test":"GET /v1/solana/full-scan/{mint}","metrics":{"recall_flag":"rug bucketed CAUTION or AVOID","recall_avoid":"rug bucketed AVOID","fp_loose":"live legitimate token bucketed CAUTION or AVOID","fp_strict":"live legitimate token bucketed AVOID"},"results":[{"tool":"soldefi","recall_flag_pct":94.7,"recall_avoid_pct":57.9,"fp_loose_pct":34.6,"fp_strict_pct":3.8},{"tool":"GoPlus","recall_flag_pct":94.7,"recall_avoid_pct":5.3,"fp_loose_pct":34.6,"fp_strict_pct":19.2},{"tool":"RugCheck","recall_flag_pct":68.4,"recall_avoid_pct":26.3,"fp_loose_pct":76.9,"fp_strict_pct":38.5}],"headline":"Same rugs caught as GoPlus (94.7%), but soldefi condemns ONE live blue-chip where RugCheck condemns ten and GoPlus five.","limitations":["This measures POST-HOC detection, not prediction. Every rug in the dataset had already rugged when it was scanned, so all three tools observe the same drained-pool chain state. High recall means 'would warn an agent scanning this token today', NOT 'would have warned you before the rug'.","The false-positive rate does NOT carry that caveat: the 26 legitimate tokens are alive and were scanned live.","N is small. At N=19, one rug is worth 5.3 percentage points — treat small differences as noise.","Runs are not perfectly reproducible: LP depth and holder distribution are measured live. Re-run before quoting.","soldefi's single strict false positive is RENDER, which genuinely has both mint and freeze authority active (RugCheck and GoPlus condemn it too). QUANT is missed by every tool — the dev dumped on a webcam stream and the chain looks clean afterwards."],"reproduce":"npx tsx benchmark/run.ts in the soldefi repo — costs $0 (local scorers + free public APIs).","full_writeup":"benchmark/README.md in the repo, including the per-token disagreements."}