Sakana AI has published Beyond Imitation, a TMLR research paper on LLM-assisted peer review built around error detection. Most AI reviewers are graded on how closely they copy human reviews. This work asks a harder question: can an AI reviewer find a planted mistake? The research team ships two pieces: a Contradiction Benchmark and a Multi-Layered Review (MLR) system. For developers building research agents, the lesson is practical. Both system design and model choice move error detection.
TL;DR
- Size: 1,164 inserted contradictions across 257 papers from 5 venues. MLR reads up to 10 pages of main text.
- Runs on: Off-the-shelf API models (Claude Sonnet 4, Claude Haiku 3.5). No GPU, no fine-tuning. About $0.47 per review.
- Performance: Highest error detection of all 4 systems tested, with human-aligned scores.
- Best: Caught 73.43% of core-claim errors with 4 reviews, versus 14.81% for the best baseline.
- Worst: Only 16.11% exact matches on real retracted arXiv papers.
- Bottom line:
- Best: reads before judging, and finds far more serious errors.
- Worst: still falls for hidden prompt injection.
What is Multi-Layered Review?
Multi-Layered Review is an agentic AI review system from Sakana AI that understands a research paper before critiquing it. It uses 3 agents on off-the-shelf Claude models:
- Appendix Agent (Claude Haiku 3.5): summarizes experiments and implementation details from the appendix.
- Literature Review Agent (Claude Sonnet 4): uses web search to place the paper in prior work. It is optional.
- Review Agent (Claude Sonnet 4): runs a 3-pass prompt chain inspired by Keshav’s Three-Pass Approach.
Pass 1 writes a high-level outline. Pass 2 reads in detail and flags weaknesses, assumptions and gaps. Pass 3 merges all agent outputs into Strengths, Weaknesses, Questions, Recommendation, Score and a To-Do list. The PDF is passed directly, so figures and equations survive.
How does the Contradiction Benchmark work?
The benchmark plants errors into real papers and checks whether reviewers catch them. The research team collected 257 CC-licensed papers from ACL, AISTATS, CVPR and ICML 2025, plus NeurIPS 2024.
Gemini 2.5 Pro builds a knowledge graph of each paper’s claims, evidence and methods. Node distance from a “main claim” sets severity. Distance 0 hits a core claim; larger distances hit details. GPT-4.1 then rewrites 1 node per distance into a contradiction, yielding 1,164 data points.
An o3 judge scores each review 10 times. On clean papers it reached 99.9% accuracy. It showed 86.8% sensitivity on manually confirmed catches, so reported scores may be conservative.
How well does MLR detect errors?
MLR led every baseline on the benchmark. With 4 reviews, it caught 73.43% of distance-0 contradictions and 40.95% overall. The best baseline, AgentReview, caught 14.81% at distance 0. A single MLR review still caught 60.79%.
An ablation separates model from design. Swapping GPT-4.1 for Claude Sonnet 4 inside LLM-Review lifted distance-0 detection from 14.56% to 35.40%. MLR’s design added about 25 more points on a single review. Accuracy falls as node distance grows, which supports the severity scoring.
On real retracted papers from WithdrarXiv-Check (211 papers), gains shrink. MLR scored 26.07% on ‘similar’ matches and 16.11% on ‘exact’ matches. The strongest baselines scored 18.48% and 9.00%.
Does MLR agree with human reviewers?
On scores, mostly yes. On ICLR 2025 submissions, MLR’s predicted scores reached a Pearson correlation of 0.586 with human scores. The human-to-human reference was 0.742. On ICML 2025, the AI Reviewer edged it, 0.439 versus 0.429.
On focus, no. MLR stresses validity and experiments, while humans weigh clarity and novelty more. The authors frame this as a complementary perspective, not a replacement.
What does it cost to run?
MLR costs about $0.47 per review, excluding the optional literature agent. It uses 189,062 input tokens, about half of the AI Reviewer’s 403,654. A single-prompt variant cut cost by about two-thirds. Its detection dropped about 3.5 points on a subset of the benchmark.
How does MLR compare with other AI reviewers?
| Feature | MLR (Sakana AI) | LLM-Review | AI Reviewer | AgentReview |
|---|---|---|---|---|
| LLM used in this study | Claude Sonnet 4 + Haiku 3.5 | GPT-4.1 | o4-mini | GPT-4o |
| Design | 3 agents, 3-pass chain | Single prompt, text truncated | 5-review ensemble, meta-review, reflection | Reviewer, author, area chair roles |
| Contradiction Benchmark, full | 40.95% (4 reviews) | 6.39% | 6.50% | 5.95% |
| Core-claim errors (distance 0) | 73.43% | 14.56% | 11.17% | 14.81% |
| WithdrarXiv-Check, similar / exact | 26.07% / 16.11% | 5.21% / 2.37% | 13.74% / 9.00% | 18.48% / 5.69% |
| ICLR 2025 Pearson vs human | 0.586 | -0.013 | 0.538 | 0.195 |
| Input tokens per review | 189,062 | 6,517 | 403,654 | 310,964 |
| Cost per review | ~$0.47 | ~$0.01 | ~$0.49 | ~$0.81 |
| Open code | On request | Yes | Yes | Yes |
All benchmark, correlation, token and cost figures come from the Beyond Imitation paper.
Key Takeaways
- Sakana AI scores AI reviewers on catching errors, not copying humans.
- MLR caught 73.43% of core-claim errors, about 5x the best baseline.
- Model swap and 3-pass design each add large detection gains.
- Real retracted-paper errors remain hard: 16.11% exact matches.
- Hidden prompt injection still sways every AI reviewer tested.
Check out the Paper here. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Asif Razzaq is the CEO of Marktechpost AI Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.








