Two independent critics from different families beat one — and beat five from one family.
Before finalizing anything, deploy a non-biased critic — and research the optimal number. The literature answered: 2, from different model families, independent. One critic carries single-model bias; three+ adds cost with non-monotonic quality ("debate fatigue"); round-trips risk "collective delusion."
The user's instruction
"I think before proceeding and finalizing anything or saying anything is done, we should adopt a non-bias critic agent deploy against it and lock that in from now on... We should figure out what the optimal amount of critics to deploy against each per result."
The research (what the literature said)
"Multi-agent debate shows diminishing, sometimes negative returns... 2025 evaluations of five MAD frameworks found they fail to consistently outperform simpler single-agent strategies at equal compute. 'Debate fatigue': quality is non-monotonic. Collective delusion: ~65% of debate failures. Diversity beats count."
Conclusion: 2 independent critics from different model families is the optimal balance.
The first deployment (it worked — and then some)
"Round 1 — REVISE: primary found no design tokens, missing focus states. Adversary found two real layout bugs. Round 2 — REVISE: the adversary caught my own fix introducing a regression (opaque bar) plus an invisible keyboard focus (WCAG failure). Round 3 — PASS (8.8 + 8.8): all issues resolved."
The adversary caught bugs a single same-family review missed — including a bug in my own fix. That's the proof the protocol works.
critic.sh — blind context, deterministic numeric gate, full archive in critiques/.