1---
2name: qa-army
3description: Run a full QA regression of an app with an army of agents — the strongest model writes the test plan, then automated, browser, and code-reading testers run rounds in parallel and file every issue as a ticket (one master regression ticket + prioritized subtickets, each with a plain-English TL;DR). Finds issues only, never fixes them. Use when the user says "qa army", "full regression", "QA the app", "test everything", "find all the bugs", or wants a pre-release QA pass.
4---
5
6# QA Army
7
8A full regression of one app, run by a team of agents, ending in tickets. The strongest model writes the test plan; testers matched to their strengths execute it in rounds; every new confirmed issue becomes a subticket under one master regression ticket, each opening with a TL;DR anyone can read. Existing tickets are reused; reopened tickets have no parent.
9
10**This skill finds issues. It never fixes them.** No code edits, no commits, no PRs, no "quick fix while I'm here". The only things it writes are the test plan, evidence files, and tickets.
11
12## Operating rules (read first)
13
141. **Find, don't fix.** Testers may read code, run tests, and drive the app. They may not edit source, config, migrations, or data the app depends on. A tester that spots the fix writes it as a *hint* in the ticket and moves on.
152. **Safe target only.** Run against local, preview, or staging. Production only if the user says so, and then read-only: no sign-ups with real emails, no payments, no sends, no deletes. Use test accounts and seed data the user provides or approves.
163. **Nothing outward-facing.** Never send real emails/SMS/WhatsApp, never charge a card, never post publicly. If a flow needs that to complete, test up to the last safe step and note the gap.
174. **Evidence or it didn't happen.** Every issue carries reproduction steps and at least one artifact: a screenshot, console/network log, failing test output, or `file:line`.
185. **Confirm before filing.** A second agent reproduces each candidate before it becomes a ticket (Phase 4). Flaky = filed as flaky, not as a bug.
196. **Dedupe before filing.** Search the tracker (open and recently closed) for the same issue; comment on an existing ticket instead of opening a duplicate.
207. **Reopened tickets lose their parent.** Whenever reopening a ticket in any project or team, including returning it for further work after failed QA, automatically clear its actual parent relationship without asking. Reopened tickets stay parentless unless the user explicitly requests a parent. Read back and verify the reopened status and no parent; report any unlinking failure. See `references/tickets.md` for tracker handling.
21
22## Inputs to gather (ask once, batched)
23
24- **Target:** URL(s) and/or repo path. Which environment.
25- **Access:** test accounts per role (anonymous, user, admin, other tenant), seed data, feature flags.
26- **Scope:** whole app (default) or named areas; what changed recently (git log since last release is a good proxy).
27- **Tracker:** Linear (team + project), GitHub Issues (repo), or local markdown (default fallback: `qa-army/<date>/tickets/`).
28- **Depth:** `quick` (one round, critical paths only) / `full` (default, rounds until dry) / `deep` (adds exploratory personas and cross-browser).
29
30If something is missing and has a sensible default, use the default and say so in the master ticket.
31
32## The roles
33
34Match each job to the cheapest model that does it well. Anything that reasons about intended behavior stays on the strongest model.
35
36| Role | Model tier | Job |
37|------|-----------|-----|
38| **Planner** | strongest | Maps the app, writes the test plan, assigns lanes |
39| **Automated runner** | small/fast | Runs existing suites (unit, integration, e2e, typecheck, lint, build) and reports failures verbatim |
40| **Scripted browser tester** | mid | Executes the plan's step-by-step cases in a headless browser, screenshot per case |
41| **Exploratory browser tester** | strongest or mid | Uses the app like a real persona — wrong inputs, back button, refresh mid-flow, double-clicks, slow network, mobile width |
42| **Code inspector** | strongest | Reads routes, handlers, data access, and edge-case logic for bugs the UI can't reveal (unchecked states, silent catches, wrong permissions, race conditions) |
43| **Verifier** | strongest | Re-reproduces each candidate from its steps alone, rejects what doesn't reproduce |
44| **Ticket writer** | mid | Turns verified issues into tickets in the house format |
45
46Full lane checklists and prompt templates: `references/lanes.md`.
47
48## The method
49
50### Phase 0 — Recon (Planner)
51
52Build a map before writing a single test case:
53
54- Stack, scripts, and existing test suites (`package.json`, `playwright.config.*`, `vitest.config.*`, CI files).
55- Every route/screen and every API entry point. Every role and what it should and shouldn't see.
56- The **critical journeys**: the 3–7 flows that, if broken, make the product useless or lose money/data (sign-up, sign-in, the core action, payment, data export...).
57- Recent changes: `git log --since=<last release>` — weight testing toward them.
58- Known issues already in the tracker, so testers don't refile them.
59
60### Phase 1 — Test plan (Planner, strongest model)
61
62Write `qa-army/<date>/TEST-PLAN.md` using `references/test-plan.md`. It holds:
63
64- Scope and out-of-scope, environment, accounts.
65- A coverage matrix: area × lane (which areas get automated, scripted browser, exploratory, code inspection).
66- Numbered test cases (`TC-001`…) — preconditions, steps, expected result, priority, lane.
67- Personas for exploratory testing.
68- Viewports/browsers for the depth chosen.
69
70Show the plan to the user as message text before running anything heavier than recon, unless they said to go end-to-end without stopping.
71
72### Phase 2 — Rounds (all testers, parallel)
73
74Fan out one agent per lane per area. Every tester returns **candidates** in this shape (not tickets yet):
75
76```
77id, title, area, lane, test_case (TC-xxx or "exploratory"),
78steps[], expected, actual, evidence[], suspected_severity, environment
79```
80
81Rounds:
82
83- **Round 1 — Plan coverage.** Every test case in the plan gets executed once. Automated runner goes first; its failures seed the others.
84- **Round 2+ — Gaps and hotspots.** The Planner reviews round results, adds cases where bugs clustered (bugs cluster), and reassigns. Exploratory testers get new personas.
85- **Stop** when two consecutive rounds surface nothing new (dedupe against everything *seen*, not just confirmed), or at the round cap for the depth (`quick` 1, `full` 4, `deep` 6).
86
87### Phase 3 — Dedupe (Planner)
88
89Merge candidates that are the same root symptom seen from different lanes (a failing e2e test + a broken button + a 500 in the logs may be one issue). Keep all evidence on the merged candidate.
90
91### Phase 4 — Verify (Verifier, strongest model)
92
93For each candidate, a fresh agent with **only** the steps and environment tries to reproduce it.
94
95- Reproduces → confirmed.
96- Reproduces sometimes → confirmed, labeled `flaky`, with the hit rate (e.g. 3/5).
97- Doesn't reproduce → dropped, listed in the master ticket's "not reproduced" section.
98- Works as designed → dropped unless the design itself hurts users; then filed as a UX issue at Low.
99
100### Phase 5 — Rate
101
102Rate each confirmed issue on four separate axes (scales in `references/tickets.md`):
103
104- **Priority** — Urgent / High / Medium / Low: how soon it should be fixed.
105- **Severity** — Blocker / Major / Minor / Cosmetic: how bad it is when it happens.
106- **Risk** — High / Medium / Low: blast radius × likelihood — how many users hit it, and what's lost (data, money, trust, security).
107- **Effort** — XS / S / M / L / XL (map to the tracker's estimate field): rough size of the fix, judged from the code, without fixing.
108
109### Phase 6 — File tickets
110
111Use `references/tickets.md` for exact templates.
112
1131. **Master regression ticket** first: `QA regression — <app> — <date>`. Holds the TL;DR of the whole run, coverage, counts by priority, the top 3 to fix first, what wasn't tested and why, and links to every subticket.
1142. **One new subticket per confirmed issue that has no existing ticket**, linked as a child of the master. Reuse existing tickets; if reopened, remove their parent and link them from the master's report as references only. Every new subticket opens with a **TL;DR** in plain, non-technical language — what's broken, who notices, why it matters — before any technical detail.
1153. Set priority, estimate, labels (`qa-army`, area, `flaky` when relevant), and project in the tracker's real fields, not only in text. Read each ticket back and confirm the fields stuck.
1164. Set status to the team's to-do state. Assign no one unless asked.
117
118### Phase 7 — Report
119
120Reply to the user with: the master ticket link, counts by priority, the top 3 issues in one line each, and what wasn't covered. Nothing else.
121
122## Harness notes
123
124The method is the same everywhere; only the fan-out mechanism changes.
125
126- **Claude Code:** spawn testers with the Agent tool (one per lane × area, in a single message so they run in parallel), passing a model per the roles table. For a large app and a user who opted into multi-agent orchestration, use a Workflow: `pipeline(lanes, run, verify)`. Browser lanes use the Playwright MCP server (isolated headless browser), never the user's own logged-in browser unless they ask.
127- **Codex:** use subagents if the session supports them; otherwise run the lanes sequentially in the same order (automated → scripted → exploratory → code inspection), and run verification in a fresh session or as a separate pass that reads only the candidate list. Browser lanes use Playwright (`npx playwright` scripts or a Playwright MCP server).
128- **Trackers:** Linear via its MCP server; GitHub via `gh issue create` with a sub-issue/task-list link to the master; otherwise local markdown files in `qa-army/<date>/tickets/` (`000-master.md`, `001-<slug>.md`, …).
129
130## Anti-patterns
131
132- Fixing something. Even one line. File it.
133- Filing a candidate nobody re-reproduced.
134- A TL;DR that uses jargon ("hydration mismatch on the RSC boundary"). Say what the person using the app sees.
135- One giant ticket listing 30 bugs. One issue per subticket.
136- Severity inflation: a typo is not High because it's on the home page. Use the scales.
137- Raw test-runner dumps pasted as tickets. Group failures by root cause, one ticket each.
138- Testing against production with real side effects.
139- Silent gaps: anything not tested goes in the master ticket's "not covered" list.
140
141## References
142
143- `references/test-plan.md` — the test plan template the Planner fills in.
144- `references/lanes.md` — per-lane checklists and the prompt template for each tester.
145- `references/tickets.md` — master and subticket templates, rating scales, and tracker field mapping.
146