Towards an AI Co-Scientist
A Multi-Agent System for Scientific Discovery
Gottweis, Weng, Daryin, Tu et al.
Google Cloud AI Research · Google Research · Google DeepMind
+ Houston Methodist · Sequome · Imperial College London · Stanford
arXiv:2502.18864 · February 26, 2025
🔬 The Problem
- Science faces a breadth vs. depth conundrum
- Exponential growth in publications — no single researcher can keep up
- Trans-disciplinary insights drive breakthroughs
(CRISPR, AlphaFold, AI itself)
- But these connections are increasingly hard to find
What is the AI Co-Scientist?
- Multi-agent system built on Gemini 2.0
- Not automating science — augmenting scientists
- "Scientist-in-the-loop" paradigm
- Input: natural language research goals
- Output: novel, testable hypotheses + ranked research proposals
System Design Overview
🧑🔬 Scientist
Research Goal
→
⚙️ Generate
Hypotheses
→
💬 Debate
Review & Rank
→
🧬 Evolve
Refine & Combine
→
📋 Ranked
Proposals
Key: asynchronous task execution · test-time compute scaling
The Six Specialized Agents (1/2)
🔍 Generation Agent — Literature search, simulated scientific debates, assumption identification
📝 Reflection Agent — Peer review simulation: initial, full, deep verification, observation, and simulation reviews
🏆 Ranking Agent — Elo-based tournament via pairwise scientific debates between hypotheses
The Six Specialized Agents (2/2)
🔗 Proximity Agent — Builds similarity graphs for clustering and deduplication of hypotheses
🧬 Evolution Agent — Refines top hypotheses: grounding, simplification, out-of-box thinking, combination
📊 Meta-review Agent — Synthesizes feedback, creates research overviews, identifies potential collaborators
Generate → Debate → Evolve
- Hypotheses compete in Elo-ranked tournaments
- Winners get refined and recombined
- Losers get eliminated
- Meta-review feeds lessons back into generation
- Self-improving without fine-tuning or RL
📈 Scaling Test-Time Compute
- More compute = better hypotheses
- Key finding: no saturation observed
- Continued improvement with more iterations
- Not just "run it longer" — structured scaling
via tournament + evolution
🧑🔬 Expert-in-the-Loop Design
Scientists can:
- Refine research goals iteratively
- Provide manual reviews at any stage
- Inject their own hypotheses into the tournament
- Direct follow-up explorations
Not a black box.
Evaluation: Elo Concordance
Elo ratings correlate with accuracy on GPQA benchmark
78.4%
Top-1 accuracy on GPQA Diamond set
Internal ranking mechanism validated against external ground truth
Evaluation: vs. Other Models
On 15 expert-curated research goals, AI Co-Scientist outperformed:
- Gemini 2.0 Pro & Flash Thinking
- OpenAI o1, o3-mini-high
- DeepSeek R1
- Human expert "best guesses"
⚠️ Caveat: auto-evaluation via Elo
Expert Evaluation
11 research goals evaluated by domain experts (PhD+, postdocs, faculty)
Most preferred system vs. all baselines
🧪 Validation 1: Drug Repurposing for AML
- Combinatorial search across drug candidates
- 78 hypotheses formatted as NIH Specific Aims pages
- Reviewed by 6 board-certified oncologists
- Rated highly across 15 evaluation axes
🔬 AML: Wet-Lab Results
Known drugs validated:
- Binimetinib (IC₅₀ = 7 nM)
- Pacritinib
- Cerivastatin
Tumor inhibition in MOLM-13 cells
Novel candidate:
- KIRA6 (IRE1α inhibitor)
- No prior AML evidence
- KG-1: IC₅₀ = 13 nM
- Also active in MOLM-13, HL-60
Real wet-lab validation, not benchmarks.
🫁 Validation 2: Liver Fibrosis
- Goal: find epigenetic modifiers for liver fibrosis
- 3 of 15 top-ranked hypotheses selected for testing
- 2 of 3 epigenetic targets showed significant anti-fibrotic activity
- Tested in human hepatic organoids
- One target drug is already FDA-approved
→ immediate repurposing opportunity
🦠 Validation 3: Antimicrobial Resistance
- cf-PICIs: capsid-forming phage-inducible chromosomal islands
- Researchers had unpublished results on gene transfer mechanism
- Fed the system the question without revealing the answer
- System independently proposed the same hypothesis as top-ranked
~2 days vs. ~10 years
AI Co-Scientist vs. conventional research
What Makes This Different?
- Actual wet-lab validation
Not just benchmarks or ranking known drugs
- Novel predictions
Not recapitulation of existing knowledge
- Scientist-in-the-loop
Not fully automated — augmentation by design
⚠️ Limitations
- Relies on open-access literature only — misses paywalled papers
- No access to negative results or failed experiments
- Limited multimodal reasoning (can't fully parse figures/charts)
- Inherited LLM limitations: hallucinations, factuality gaps
- Elo is auto-evaluated — not objective ground truth
- No clinical trial design capability
- In vitro ≠ clinical efficacy
🛡️ Safety & Ethics
- 1,200 adversarial goals across 40 topics — all rejected
- Multi-layer safety pipeline:
goal review → hypothesis review → continuous monitoring → logging
- Trusted Tester Program for controlled access
- Dual-use risk acknowledged — needs ongoing work
Implications for Research
- Drug repurposing at scale
- Hypothesis generation as a service
- Accelerating cross-disciplinary discovery
- Augmenting, not replacing, scientists
- Test-time compute as a new dimension of AI capability
🤔 Critical Questions
- Can it work beyond biomedicine?
Claimed general-purpose, only tested in bio
- How much does expert pre-screening bias results?
- Is Elo self-evaluation reliable for truly novel hypotheses?
- What happens when wrong hypotheses pass all reviews?
- Cost and compute requirements remain unclear
Adapting for Our Organization
What can we take from the AI Co-Scientist architecture?
- 5 actionable adaptations mapped to our products
- Focus: SchemaForge, The Auditor, The Updater
- Principle: steal the architecture, not the compute bill
1. Tournament Evolution for SchemaForge
Research
Question
→
Generate
3–5 designs
→
Debate
Pairwise critique
→
Evolve
Best design
- Current: SchemaForge generates one study design per query
- Adapted: Generate 3–5 competing designs, rank via pairwise debate
- User gets the winner, not the first draft
- Same principle as their Ranking Agent tournaments
2. Reflection Agent → Automated Peer Review
Their 6 review types, adapted to study design:
Initial Review — Quick check: is the design coherent? Correct outcome type? Right comparator?
Methodological Review — Confounders, selection bias, measurement error, assumption violations
Deep Verification — Decompose into sub-assumptions. Challenge each one independently
Feasibility Review — Sample size realistic? Data available? Timeline achievable?
Productize what Anas does manually as a reviewer
3. Meta-Review Feedback Loop
- Their Meta-review Agent finds recurring weaknesses across all hypotheses
- Feeds patterns back → generation improves over time
Our adaptation:
- Track common design mistakes across SchemaForge users
- Examples: immortal time bias, wrong estimand, collider conditioning
- Feed patterns into generation prompts automatically
- System gets smarter without retraining
4. Elo Ranking for Evidence Quality
The Auditor — Evidence Quality Engine
Paper A
vs
Paper B
→
🏆 Winner
- Absolute scoring is unreliable — "7.2 out of 10" means nothing
- Pairwise comparison: "Which study has stronger methodology?"
- Derive Elo ratings from tournament results
- More robust, more interpretable, same approach as the Co-Scientist
- Enables: living systematic reviews ranked by evidence strength
5. Proximity Graph for Literature
Their Proximity Agent → Our literature intelligence layer
- Cluster papers by methodology, findings, population
- Identify research gaps — areas with few or weak studies
- Deduplicate overlapping evidence
- Map the boundary of current knowledge for any research question
Feeds into: The Updater (Living Guidelines Platform)
Operator: Helicase (CDO)
⏭️ What We Skip
- Heavy compute scaling — hours/days of runtime
Too expensive for $29/mo SaaS. Be smart about compute budget.
- Complex Supervisor orchestration
Our sessions_spawn + clear task descriptions is simpler and works.
- Building on a single model family
They're locked to Gemini 2.0. Our Cortex board uses multi-model consensus.
Steal the architecture, not the compute bill.
📋 Implementation Roadmap
This month — Tournament evolution in SchemaForge (3 designs → debate → winner)
Next month — Automated peer review layer (Reflection Agent lite)
Q2 — Elo ranking engine for The Auditor MVP
Q2 — Meta-review feedback loop (track user patterns)
Q3 — Proximity graph for literature mapping (The Updater)
Each builds on the previous. Ship incrementally.
"The wet-lab validation is what separates this from every other AI-for-science paper. Most stop at 'our model ranked known drugs correctly.' These people actually grew organoids and ran assays."
arXiv:2502.18864
Presentation prepared for academic discussion