Skip to content
The Agent Builder idea, applied to ML

Fromrawdatatoamodelyoucandefend.

An agent runs the machine learning loop inside your environment, one idea per sandboxed session. An orchestrator it cannot override scores, checks and judges every result.

See how it works
Runs inside your environmentA gain counts when it is significantIt asks before it spends
demo-churn · run 4MOCKUP · SAMPLE DATA

best CV ROC AUC

0.8412

+0.0391 over the baseline

best so far, by iteration

  • adopted
  • not significant
  • rejected
  • baseline
  • LLM spend

    $1.84 / $5.00

  • compute

    2.1 h / 6 h

  • iterations

    12 / 20

  • wall clock

    1 h 52 / 4 h

event river

CH.01·The problem

A model that grades itself is a model nobody can check.

A coding agent can build a model in an afternoon. Left alone, it also picks the seed, reports its own score and can read the test set.

The agent alone: Reports its own cross-validation score
With the orchestrator: The harness scores every candidate on the same fixed folds
The agent alone: Picks the random seed
With the orchestrator: The seed is derived from the run and never chosen
The agent alone: Can open the test set
With the orchestrator: The holdout is mounted once, at the very end
The agent alone: Keeps any gain, however small
With the orchestrator: A gain is adopted only when a paired test says it is real
The agent alone: Stops when it decides it is done
With the orchestrator: A stop rule ends the run on patience or on budget
Who grades the work. Same agent, same data, two answers to that question.

CH.02·Follow one run

One run, from a data source to a held-out score.

Each experiment is a fresh agent session in a sandbox. Everything between the sessions belongs to the orchestrator.

CH.03·The judge

Agents propose. The orchestrator judges.

The agent writes the model. The harness grades it, on rules the agent cannot see past and cannot change.

Every idea, on the same folds

MOCKUP · SAMPLE DATA

The best score moves only when a gain is adopted. Ideas inside the noise stay on the chart, muted, so nothing tried is hidden.

1 · 0.8021 · baseline2 · 0.8109 · adopted3 · 0.8087 · not significant4 · 0.8190 · adopted5 · 0.8152 · not significant6 · 0.8236 · rejected7 · 0.8297 · adopted8 · 0.8281 · not significant9 · 0.8344 · adopted10 · 0.8330 · not significant11 · 0.8412 · adopted12 · 0.8398 · not significant
  • adopted
  • not significant
  • rejected
  • baseline

Seven checks before anything counts

  • schemathe result is valid result/v1
  • contractthe entrypoint keeps the contract
  • gitthe commit is HEAD and the tree is clean
  • mlflowthe run is in this run's experiment
  • featuresno forbidden feature was used
  • seed_envseed and environment hash match
  • isolationeach fold was fitted in isolation

A gain has to clear a test

Nadeau–Bengio corrected resampled paired t-test, one-sided, α 0.05. The correction is for folds that share training rows.

It stops by rule

Patience on adopted gains, plus budgets for LLM spend, compute and time. Before it stops, it can ask itself for a different direction.

MOCKUP · SAMPLE DATA

It looks for where it is wrong

Segments where the best model errs significantly go back into the next brief, as the first thing the agent reads.

CH.04·Control

You steer it. It asks before it spends.

Three autonomy modes, and a console that can pause, steer, veto or stop a run at any point.

Autonomy

  • AutonomousRuns the loop inside its budgets.
  • GatedAsks before every remote job and shows exactly what data leaves.
  • SupervisedAsks before every hypothesis.

Steer note

“Try monotonic constraints on tenure before adding more trees.”

Ana · operator · 10:42

  • pause
  • stop
  • veto an idea
  • edit budgets
  • promote with a reason

Drive it from Claude Code

An operator MCP server with 23 tools, behind the same sign-in and the same audit log as the cockpit.

❯ claude mcp add --transport http \

auto-ml-scientist \

https://<your-host>/operator/mcp/

✓ 23 tools

Every decision, timed and attributed

Approvals, vetoes, notes and promotions land in an append-only log, with the person and the minute.

  1. 10:42Anasteer note added
  2. 10:47Anaremote job approved · cap $4.00
  3. 11:03Ruihypothesis H-09 vetoed
  4. 11:58Ruiiteration 11 promoted · reason given

CH.05·Results

Eight runs on public data, with the intervals left in.

Four public datasets, two runs each, on claude-opus-5-5 in October 2026. The reference is a default LightGBM scored on the same holdout.

  • Auto ML Scientist, with its 95% interval
  • its own baseline
  • reference: default LightGBM
  • california-housing

    RMSE · lower is better

    0.47310.3883 · 0.38300.3552
  • bank-marketing

    PR AUC · higher is better

    0.60960.6648 · 0.64720.7033
  • adult

    ROC AUC · higher is better

    0.92340.9295 · 0.92980.9359
  • credit-g

    ROC AUC · higher is better

    0.69430.7798 · 0.78170.8608
LLM spend per run
$2.27 to $3.34
wall clock per run
31 to 141 min
iterations per run
5 to 7
Show the numbers as a table
datasetmetricbaselinereferencerun95% intervalLLM spend
california-housingRMSE0.46500.46350.38830.3698 – 0.4069$2.57
0.46500.46350.38300.3633 – 0.4032$2.58
bank-marketingPR AUC0.63300.63040.66480.6329 – 0.6968$3.34
0.63300.63040.64720.6161 – 0.6783$2.51
adultROC AUC0.92630.92700.92950.9243 – 0.9347$2.27
0.92630.92700.92980.9245 – 0.9350$2.72
credit-gROC AUC0.76900.72290.77980.7058 – 0.8493$3.19
0.76900.72290.78170.7127 – 0.8487$2.49

Every run beat the reference on its point estimate. Only on California housing does the interval clear it.

On adult and credit-g the gains are small, and credit-g has 1,000 rows, so its interval is wide.

No image, forecast or text benchmark has run yet. They are defined in the suite and will be published when they have.

CH.06·What you keep

A model you can hand over, with its reasons attached.

A report a client can read

Held-out results with intervals, where the model is weak, what was tried and the provenance to reproduce it. It also exports to PDF and to a slide deck.

A package an engineer can run

The model, a FastAPI service with health and predict routes, a batch script, a pinned Dockerfile, a model card and a reproduction test.

A verdict after it ships

Schema checks and drift per feature, then one of three answers. A retrain is one more run.

The next run starts from this one

Up to six adopted ideas and six failed ones carry into later runs of the same or a similar project. An operator can retire any of them.

What it is not

  • 01

    A SaaS

    It is deployed once per client, in your cloud or on your servers. Your data does not come to us.

  • 02

    A promise of big gains

    On a problem that is already well tuned the gain can be small, and the report says so in those words.

  • 03

    Unattended spending

    Remote compute is off by default, capped on price and approved one job at a time.

  • 04

    A black box

    Every hypothesis, check, decision and person is written to a log nobody can edit.

Bring one dataset and the question you want it to answer

We deploy into your environment, run the loop on your data and walk you through the report, starting with where the model is weak.

The holdout never reaches an agent.