Towards an AI Co-Scientist

A Multi-Agent System for Scientific Discovery

Gottweis, Weng, Daryin, Tu et al.
Google Cloud AI Research · Google Research · Google DeepMind
+ Houston Methodist · Sequome · Imperial College London · Stanford

arXiv:2502.18864 · February 26, 2025

🔬 The Problem

  • Science faces a breadth vs. depth conundrum
  • Exponential growth in publications — no single researcher can keep up
  • Trans-disciplinary insights drive breakthroughs
    (CRISPR, AlphaFold, AI itself)
  • But these connections are increasingly hard to find

What is the AI Co-Scientist?

  • Multi-agent system built on Gemini 2.0
  • Not automating science — augmenting scientists
  • "Scientist-in-the-loop" paradigm
  • Input: natural language research goals
  • Output: novel, testable hypotheses + ranked research proposals

System Design Overview

🧑‍🔬 Scientist
Research Goal
⚙️ Generate
Hypotheses
💬 Debate
Review & Rank
🧬 Evolve
Refine & Combine
📋 Ranked
Proposals

Key: asynchronous task execution · test-time compute scaling

The Six Specialized Agents (1/2)

🔍 Generation Agent — Literature search, simulated scientific debates, assumption identification
📝 Reflection Agent — Peer review simulation: initial, full, deep verification, observation, and simulation reviews
🏆 Ranking Agent — Elo-based tournament via pairwise scientific debates between hypotheses

The Six Specialized Agents (2/2)

🔗 Proximity Agent — Builds similarity graphs for clustering and deduplication of hypotheses
🧬 Evolution Agent — Refines top hypotheses: grounding, simplification, out-of-box thinking, combination
📊 Meta-review Agent — Synthesizes feedback, creates research overviews, identifies potential collaborators

Generate → Debate → Evolve

  • Hypotheses compete in Elo-ranked tournaments
  • Winners get refined and recombined
  • Losers get eliminated
  • Meta-review feeds lessons back into generation
  • Self-improving without fine-tuning or RL

📈 Scaling Test-Time Compute

  • More compute = better hypotheses
  • Key finding: no saturation observed
  • Continued improvement with more iterations
  • Not just "run it longer" — structured scaling
    via tournament + evolution

🧑‍🔬 Expert-in-the-Loop Design

Scientists can:

  • Refine research goals iteratively
  • Provide manual reviews at any stage
  • Inject their own hypotheses into the tournament
  • Direct follow-up explorations

Not a black box.

Evaluation: Elo Concordance

Elo ratings correlate with accuracy on GPQA benchmark

78.4%

Top-1 accuracy on GPQA Diamond set

Internal ranking mechanism validated against external ground truth

Evaluation: vs. Other Models

On 15 expert-curated research goals, AI Co-Scientist outperformed:

  • Gemini 2.0 Pro & Flash Thinking
  • OpenAI o1, o3-mini-high
  • DeepSeek R1
  • Human expert "best guesses"

⚠️ Caveat: auto-evaluation via Elo

Expert Evaluation

11 research goals evaluated by domain experts (PhD+, postdocs, faculty)

Novelty

3.64/5

Impact

3.09/5

Most preferred system vs. all baselines

🧪 Validation 1: Drug Repurposing for AML

  • Combinatorial search across drug candidates
  • 78 hypotheses formatted as NIH Specific Aims pages
  • Reviewed by 6 board-certified oncologists
  • Rated highly across 15 evaluation axes

🔬 AML: Wet-Lab Results

Known drugs validated:

  • Binimetinib (IC₅₀ = 7 nM)
  • Pacritinib
  • Cerivastatin

Tumor inhibition in MOLM-13 cells

Novel candidate:

  • KIRA6 (IRE1α inhibitor)
  • No prior AML evidence
  • KG-1: IC₅₀ = 13 nM
  • Also active in MOLM-13, HL-60

Real wet-lab validation, not benchmarks.

🫁 Validation 2: Liver Fibrosis

  • Goal: find epigenetic modifiers for liver fibrosis
  • 3 of 15 top-ranked hypotheses selected for testing
  • 2 of 3 epigenetic targets showed significant anti-fibrotic activity
  • Tested in human hepatic organoids
  • One target drug is already FDA-approved
    → immediate repurposing opportunity

🦠 Validation 3: Antimicrobial Resistance

  • cf-PICIs: capsid-forming phage-inducible chromosomal islands
  • Researchers had unpublished results on gene transfer mechanism
  • Fed the system the question without revealing the answer
  • System independently proposed the same hypothesis as top-ranked
~2 days vs. ~10 years

AI Co-Scientist vs. conventional research

What Makes This Different?

  1. Actual wet-lab validation
    Not just benchmarks or ranking known drugs
  2. Novel predictions
    Not recapitulation of existing knowledge
  3. Scientist-in-the-loop
    Not fully automated — augmentation by design

⚠️ Limitations

  • Relies on open-access literature only — misses paywalled papers
  • No access to negative results or failed experiments
  • Limited multimodal reasoning (can't fully parse figures/charts)
  • Inherited LLM limitations: hallucinations, factuality gaps
  • Elo is auto-evaluated — not objective ground truth
  • No clinical trial design capability
  • In vitro ≠ clinical efficacy

🛡️ Safety & Ethics

  • 1,200 adversarial goals across 40 topics — all rejected
  • Multi-layer safety pipeline:
    goal review → hypothesis review → continuous monitoring → logging
  • Trusted Tester Program for controlled access
  • Dual-use risk acknowledged — needs ongoing work

Implications for Research

  • Drug repurposing at scale
  • Hypothesis generation as a service
  • Accelerating cross-disciplinary discovery
  • Augmenting, not replacing, scientists
  • Test-time compute as a new dimension of AI capability

🤔 Critical Questions

  • Can it work beyond biomedicine?
    Claimed general-purpose, only tested in bio
  • How much does expert pre-screening bias results?
  • Is Elo self-evaluation reliable for truly novel hypotheses?
  • What happens when wrong hypotheses pass all reviews?
  • Cost and compute requirements remain unclear

Adapting for Our Organization

What can we take from the AI Co-Scientist architecture?

  • 5 actionable adaptations mapped to our products
  • Focus: SchemaForge, The Auditor, The Updater
  • Principle: steal the architecture, not the compute bill

1. Tournament Evolution for SchemaForge

Research
Question
Generate
3–5 designs
Debate
Pairwise critique
Evolve
Best design
  • Current: SchemaForge generates one study design per query
  • Adapted: Generate 3–5 competing designs, rank via pairwise debate
  • User gets the winner, not the first draft
  • Same principle as their Ranking Agent tournaments

2. Reflection Agent → Automated Peer Review

Their 6 review types, adapted to study design:

Initial Review — Quick check: is the design coherent? Correct outcome type? Right comparator?
Methodological Review — Confounders, selection bias, measurement error, assumption violations
Deep Verification — Decompose into sub-assumptions. Challenge each one independently
Feasibility Review — Sample size realistic? Data available? Timeline achievable?

Productize what Anas does manually as a reviewer

3. Meta-Review Feedback Loop

  • Their Meta-review Agent finds recurring weaknesses across all hypotheses
  • Feeds patterns back → generation improves over time

Our adaptation:

  • Track common design mistakes across SchemaForge users
  • Examples: immortal time bias, wrong estimand, collider conditioning
  • Feed patterns into generation prompts automatically
  • System gets smarter without retraining

4. Elo Ranking for Evidence Quality

The Auditor — Evidence Quality Engine

Paper A
vs
Paper B
🏆 Winner
  • Absolute scoring is unreliable — "7.2 out of 10" means nothing
  • Pairwise comparison: "Which study has stronger methodology?"
  • Derive Elo ratings from tournament results
  • More robust, more interpretable, same approach as the Co-Scientist
  • Enables: living systematic reviews ranked by evidence strength

5. Proximity Graph for Literature

Their Proximity Agent → Our literature intelligence layer

  • Cluster papers by methodology, findings, population
  • Identify research gaps — areas with few or weak studies
  • Deduplicate overlapping evidence
  • Map the boundary of current knowledge for any research question

Feeds into: The Updater (Living Guidelines Platform)
Operator: Helicase (CDO)

⏭️ What We Skip

  • Heavy compute scaling — hours/days of runtime
    Too expensive for $29/mo SaaS. Be smart about compute budget.
  • Complex Supervisor orchestration
    Our sessions_spawn + clear task descriptions is simpler and works.
  • Building on a single model family
    They're locked to Gemini 2.0. Our Cortex board uses multi-model consensus.

Steal the architecture, not the compute bill.

📋 Implementation Roadmap

This month — Tournament evolution in SchemaForge (3 designs → debate → winner)
Next month — Automated peer review layer (Reflection Agent lite)
Q2 — Elo ranking engine for The Auditor MVP
Q2 — Meta-review feedback loop (track user patterns)
Q3 — Proximity graph for literature mapping (The Updater)

Each builds on the previous. Ship incrementally.

"The wet-lab validation is what separates this from every other AI-for-science paper. Most stop at 'our model ranked known drugs correctly.' These people actually grew organoids and ran assays."

arXiv:2502.18864
Presentation prepared for academic discussion