Agentic Mechanistic Interpretability

Seesaw

A multi-agent system that automates the hypothesis → experiment → critique loop in mechanistic interpretability research — plan, run, and critique, with a human approving every step.

Overview

A human researcher provides a research question. Three agents handle the rest — planning, execution, and review — with human-in-the-loop checkpoints between every stage, so nothing runs without approval.

Scout research question → Research Plan → Lens Research Plan → ExperimentBundle → Quill ExperimentBundle → CritiqueReport
Seesaw v0 architecture diagram

The Three Agents

🔭

Scout

The Research Planner

A ReAct agent that searches arXiv and the web, then commits to a falsifiable hypothesis and a concrete experiment sequence for Lens to run.

LangGraph ReAct · Claude · Firecrawl
🔬

Lens

The Experiment Runner

A six-node LangGraph state machine that executes TransformerLens experiments from the plan, interprets each result, and can queue adaptive follow-ups.

LangGraph StateGraph · Claude · TransformerLens
🪶

Quill

The Reviewer

Critiques the experiment bundle like a peer reviewer — flags methodological gaps and unsupported conclusions, and generates executable follow-up specs.

LangGraph StateGraph · Claude Opus

Lens's Tools

Five TransformerLens-based diagnostic and causal tools, each independently callable.

ToolAnswers
logit_lensWhere across layers does the model commit to an answer?
attention_patternWhich tokens does each head attend to?
direct_logit_attributionWhich components write the final prediction?
ablationIs this component causally necessary?
activation_patchingWhich layer and position carries the signal?

Evaluation

A three-layer eval suite: graders, fixtures that validate the graders themselves, and tasks built from published ground-truth papers plus adversarial probes.

29
grading functions
56
grader fixtures
30/30
offline fixtures pass
15
ground-truth + adversarial tasks

Deployment

Nothing always-on. Scout, Quill, and Lens's graph run on a Modal CPU app; Lens's five experiment tools run on a Modal GPU app with per-second billing. A thin Next.js dashboard on Vercel queues jobs and reads state from a Postgres (Neon) job store — the dashboard never executes an agent directly.