Skip to content
meirlabs

tool-eval

Pick tools with a multi-agent evaluation, not a hunch

v0.1.0

Runs "what platform should I use for X" as a five-phase agent harness: parallel research with pricing fetched from official pages, an opus shortlist that names the decision-flipping claims, adversarial verification of exactly those claims, a judge panel scored on the buyer's real priorities, and a decision memo with a runner-up and switch triggers. Born from a real run that caught the planned pick silently missing a hard requirement — and reversed it.

npx @meir-labs/skill-tool-eval
GitHub

What it does

  • researches 8-10 candidates in parallel, pricing verified against official pages, never from memory
  • adversarially fact-checks the load-bearing claims of every finalist before judging
  • scores finalists through judge lenses derived from the buyer's ranked priorities
  • delivers a decision memo with growth path, switch triggers, and a setup plan