Fromrawdatatoamodelyoucandefend.
An agent runs the machine learning loop inside your environment, one idea per sandboxed session. An orchestrator it cannot override scores, checks and judges every result.
best CV ROC AUC
0.8412+0.0391 over the baseline
best so far, by iteration
- adopted
- not significant
- rejected
- baseline
LLM spend
$1.84 / $5.00
compute
2.1 h / 6 h
iterations
12 / 20
wall clock
1 h 52 / 4 h
event river
CH.01·The problem
A model that grades itself is a model nobody can check.
A coding agent can build a model in an afternoon. Left alone, it also picks the seed, reports its own score and can read the test set.
CH.02·Follow one run
One run, from a data source to a held-out score.
Each experiment is a fresh agent session in a sandbox. Everything between the sessions belongs to the orchestrator.
CH.03·The judge
Agents propose. The orchestrator judges.
The agent writes the model. The harness grades it, on rules the agent cannot see past and cannot change.
Every idea, on the same folds
MOCKUP · SAMPLE DATAThe best score moves only when a gain is adopted. Ideas inside the noise stay on the chart, muted, so nothing tried is hidden.
- adopted
- not significant
- rejected
- baseline
Seven checks before anything counts
- schemathe result is valid result/v1
- contractthe entrypoint keeps the contract
- gitthe commit is HEAD and the tree is clean
- mlflowthe run is in this run's experiment
- featuresno forbidden feature was used
- seed_envseed and environment hash match
- isolationeach fold was fitted in isolation
A gain has to clear a test
Nadeau–Bengio corrected resampled paired t-test, one-sided, α 0.05. The correction is for folds that share training rows.
It stops by rule
Patience on adopted gains, plus budgets for LLM spend, compute and time. Before it stops, it can ask itself for a different direction.
MOCKUP · SAMPLE DATA
It looks for where it is wrong
Segments where the best model errs significantly go back into the next brief, as the first thing the agent reads.
CH.04·Control
You steer it. It asks before it spends.
Three autonomy modes, and a console that can pause, steer, veto or stop a run at any point.
Autonomy
- AutonomousRuns the loop inside its budgets.
- GatedAsks before every remote job and shows exactly what data leaves.
- SupervisedAsks before every hypothesis.
Steer note
“Try monotonic constraints on tenure before adding more trees.”
Ana · operator · 10:42
- pause
- stop
- veto an idea
- edit budgets
- promote with a reason
Drive it from Claude Code
An operator MCP server with 23 tools, behind the same sign-in and the same audit log as the cockpit.
❯ claude mcp add --transport http \
auto-ml-scientist \
https://<your-host>/operator/mcp/
✓ 23 tools
Every decision, timed and attributed
Approvals, vetoes, notes and promotions land in an append-only log, with the person and the minute.
- 10:42Anasteer note added
- 10:47Anaremote job approved · cap $4.00
- 11:03Ruihypothesis H-09 vetoed
- 11:58Ruiiteration 11 promoted · reason given
CH.05·Results
Eight runs on public data, with the intervals left in.
Four public datasets, two runs each, on claude-opus-5-5 in October 2026. The reference is a default LightGBM scored on the same holdout.
- Auto ML Scientist, with its 95% interval
- its own baseline
- reference: default LightGBM
california-housing
RMSE · lower is better
0.47310.3883 · 0.38300.3552bank-marketing
PR AUC · higher is better
0.60960.6648 · 0.64720.7033adult
ROC AUC · higher is better
0.92340.9295 · 0.92980.9359credit-g
ROC AUC · higher is better
0.69430.7798 · 0.78170.8608
- LLM spend per run
- $2.27 to $3.34
- wall clock per run
- 31 to 141 min
- iterations per run
- 5 to 7
Show the numbers as a table
| dataset | metric | baseline | reference | run | 95% interval | LLM spend |
|---|---|---|---|---|---|---|
| california-housing | RMSE | 0.4650 | 0.4635 | 0.3883 | 0.3698 – 0.4069 | $2.57 |
| 0.4650 | 0.4635 | 0.3830 | 0.3633 – 0.4032 | $2.58 | ||
| bank-marketing | PR AUC | 0.6330 | 0.6304 | 0.6648 | 0.6329 – 0.6968 | $3.34 |
| 0.6330 | 0.6304 | 0.6472 | 0.6161 – 0.6783 | $2.51 | ||
| adult | ROC AUC | 0.9263 | 0.9270 | 0.9295 | 0.9243 – 0.9347 | $2.27 |
| 0.9263 | 0.9270 | 0.9298 | 0.9245 – 0.9350 | $2.72 | ||
| credit-g | ROC AUC | 0.7690 | 0.7229 | 0.7798 | 0.7058 – 0.8493 | $3.19 |
| 0.7690 | 0.7229 | 0.7817 | 0.7127 – 0.8487 | $2.49 |
Every run beat the reference on its point estimate. Only on California housing does the interval clear it.
On adult and credit-g the gains are small, and credit-g has 1,000 rows, so its interval is wide.
No image, forecast or text benchmark has run yet. They are defined in the suite and will be published when they have.
CH.06·What you keep
A model you can hand over, with its reasons attached.
A report a client can read
Held-out results with intervals, where the model is weak, what was tried and the provenance to reproduce it. It also exports to PDF and to a slide deck.
A package an engineer can run
The model, a FastAPI service with health and predict routes, a batch script, a pinned Dockerfile, a model card and a reproduction test.
A verdict after it ships
Schema checks and drift per feature, then one of three answers. A retrain is one more run.
The next run starts from this one
Up to six adopted ideas and six failed ones carry into later runs of the same or a similar project. An operator can retire any of them.
What it is not
- 01
A SaaS
It is deployed once per client, in your cloud or on your servers. Your data does not come to us.
- 02
A promise of big gains
On a problem that is already well tuned the gain can be small, and the report says so in those words.
- 03
Unattended spending
Remote compute is off by default, capped on price and approved one job at a time.
- 04
A black box
Every hypothesis, check, decision and person is written to a log nobody can edit.
Bring one dataset and the question you want it to answer
We deploy into your environment, run the loop on your data and walk you through the report, starting with where the model is weak.
The holdout never reaches an agent.