☑️ Antiagree [Claude Code Plugin - Agent Skill - Evals]
Antiagree reviews an idea, plan, architecture, benchmark or decision the way an outside reviewer with no stake in it would. It checks the claims against the code and the data, looks for what already exists, tries to kill the proposal with the strongest argument it can find, and keeps only what survives. Every review ends in one decision: GO, ITERATE, RETHINK or KILL. It doesn’t fake harshness: when the plan is sound, it says GO. It ships as both a Claude Code plugin and an Agent Skill. Instructions only, no telemetry, MIT.
View on GitHub DocumentationAgreement is the default:
Language models lean toward agreeing with the person in front of them. Ask “is this a good plan?” and you often get a polished yes, a few minor tweaks and an estimate nobody measured. The opposite failure is just as common: ask for brutal honesty and you get invented problems, because that is what the role seems to want. Both tell you what you asked to hear instead of what is true.
Antiagree’s loyalty is to your goal, not to your proposal, your earlier decisions or anything it said earlier in the conversation. Intellectual brutality, not tonal brutality.
How a review works:
- Check the ledger. If
ANTIAGREE.mdexists, it reads it first: a kill criterion that was met, or a deadline that passed, leads the review. - Investigate before judging. It verifies the three to five claims the decision depends on, in the code, data, benchmarks and git history. With no artifacts, it says the review is reasoning-only.
- Check whether it already exists. For anything you plan to build, it searches for prior art and asks what is actually different.
- Find the real goal. What result is wanted, which variable matters, and whether the question itself is framed wrong.
- Attack, then try to kill it. Assumptions, activity versus impact, causality, the measurement, hidden costs, second-order effects, people and incentives, sunk cost and simpler paths. Then the strongest case against the proposal, not a straw man.
- Rebuild. The bottleneck, the few moves with real leverage, and experiments with GO and KILL thresholds.
Every claim that matters carries a label, in the user’s language: FACT (with its source), STRONGLY SUPPORTED, HYPOTHESIS or SPECULATION. No invented percentages: if nobody knows, the answer is “we don’t know yet”.
The verdict:
Every review opens with a card: what is being reviewed, the verdict, the biggest risk, the cheapest next step that settles the question, and the one fact that would change its mind.
- GO: sound; proceed. The default when the diagnosis is supported and the plan can be verified and rolled back cheaply.
- ITERATE: right goal and approach, with a must-fix the plan doesn’t already cover.
- RETHINK: right goal, wrong approach.
- KILL: drop it.
It closes with the uncomfortable truth, what to keep, what to change, what not to build and what’s still unknown.
Reviewing the session, and holding you to it:
- Session review. Run
/antiagreewith nothing after it and it audits what you and Claude already decided in the conversation, from the first message on. It says up front that it helped make those decisions, and opens with a table of every number and claim the session accepted: who introduced it, whether it was verified, and what depends on it. - Kill criteria. When a review produces experiments with thresholds, it offers to record them in
ANTIAGREE.md. The next review reads them first, so the bar can’t move after the result is in. - Calibrated. It names what the plan gets right before attacking it, credits the safeguards already built in, and changes position on new facts or better arguments, not on displeasure.
Install:
# Claude Code
/plugin marketplace add antiagree/antiagree
/plugin install antiagree@antiagree
# Use
/antiagree # review this session's decisions
/antiagree our plan to split the monolith # review anything
/antiagree quick bench/RESULTS.md # verdict card + top three findings
Other agents with Agent Skills (Codex, Cursor, Gemini CLI) install it by copying skills/antiagree/ into their skills directory. It also triggers on plain requests like “attack my plan” or “no me des la razón”.
Checked, not assumed:
The repository ships five eval cases for claude plugin eval: a benchmark with planted flaws, a sound plan with “be brutal” in the prompt, a Spanish case where the confidence interval includes zero, a kill criterion that the latest results met, and a session that agreed its way into a plan. Four of them run with and without the plugin.
On those cases, with three runs each and Opus 5.5 as the judge, Sonnet 5.5 and Opus 5.5 passed all 15 runs with Antiagree, against 0 of 12 and 3 of 12 without it. Without the plugin, both still found the planted flaws, but invented problems in the sound plan and gave no decision rule for the benchmark. The cases were written for this project, so the numbers are a signal, not proof.
Technical Stack:
- Format: a single Agent Skill, one
SKILL.md, with no hooks, no MCP servers and no scripts. The tools it pre-approves are read-only (Read, Grep, Glob); web searches for prior art ask for permission. - Distribution: a Claude Code plugin marketplace and the Agent Skills standard.
- Evals:
claude plugin evalwith a with-and-without ablation, LLM graders with written pass criteria, a replayed session history and scaffolded workspaces. - CI: strict plugin validation, manifest and eval case checks, and plugin directory rules (file sizes, file names, wording), with no API key and no cost.
- Privacy: no network code and no telemetry. It reads files in your project, writes
ANTIAGREE.mdonly when you agree, and its searches go through Claude’s own WebSearch tool.
Conclusion:
Antiagree started from my own long prompts asking Claude to stop agreeing with me, turned into a skill anyone can install. The project demonstrates:
- Shaping model behavior with instructions alone: a skill that changes how Claude reviews, with no code running on the user’s machine.
- Treating calibration as a feature: saying GO on a sound plan is tested as hard as catching a flawed one.
- Measuring a prompt like a product: ablation evals with and without the plugin, graders in the repository, and results published with their limits.
- Making decisions durable: a ledger of kill criteria that the next review enforces.
Tags
Share