Aero autoresearch

An agent is given a set of car shapes whose drag has been measured by computational fluid dynamics, and an entire body style is held back that it never sees. It gets a sandbox, a shell and a time budget. It decides its own method, trains whatever it likes, and pushes a submission when it thinks it has improved.

The reference data

velocity magnitude, centreline
Velocity through the centreline. The blue region behind the car is the wake, and it is most of the drag.
pressure coefficient, centreline
Pressure on the same slice. High at the nose, low over the roof.
velocity magnitude, wake cross-section
A cross-section cut through the wake, three metres behind the nose.
local drag contribution
Where the drag actually comes from. Red adds drag, blue removes it. This is the quantity the benchmark asks a model to get right.

Results

resolution limit ±0.9695 ct (paired, n=97) y_s trivial atlas mean · 47.65 ct best honest free rung (b1) · 32.47 ct DoMINO (external, other split) · 3.36 ct 0.0 10.8 21.6 32.3 43.1 53.9 0.00 0.13 0.27 0.40 0.53 0.67 Agent runtime, orchestrator clock (hours) Drag MAE from the submitted field (counts, 1 ct = 0.001 Cd) MODEL gpt-5.6-sol 5.71 ct · best of 9 runs claude-opus-5 5.92 ct · best of 9 runs