Research

The state of the art in fluid-dynamics foundation models, in common units

August 2026 · 46 published rows from 13 owners across 20 evaluation corpora · every number is tagged MEASURED, ESTIMATED or UNKNOWN · 7 rows admit a drag-count conversion; the other 39 say MISSING · the same table is published as data

For anyone deciding whether a fluid-dynamics surrogate is good enough to make a decision with - and for anyone building a benchmark that has to state what the bar already is.

Conventions, restated on every row. An automotive drag count is 0.001 Cd. An aircraft drag count is 0.0001 Cd - ten times smaller. A 2D airfoil count is 0.0001 Cd. Never convert across those domains without saying so; we quoted an aircraft figure against automotive numbers once and it was an 11x error in our own file. Every number here is MEASURED, ESTIMATED (under a stated assumption) or UNKNOWN. Missing data is MISSING, never 0, and that applies to denominators too. "Not reported" means we did not find it after a targeted search of the sources we hold.


1. What this is, and what it is not

Every vendor in this field publishes a real number. They publish it in a different unit, on a different dataset, against a different evaluation set, under a different normalisation. So the field has a state of the art and no way to read it. This page is the assembly: what each model actually reports, converted to a common unit where a conversion is defensible, and placed beside the ground truth's own repeatability.

It is not a ranking. Section 6 says why, and the reason is the finding.

Three things stop these rows being comparable, and each one is independently sufficient.

  1. Units. Only one corpus in this table publishes both a force scale and a label-noise figure. That is why 7 of 46 rows carry a drag-count conversion and 39 say MISSING.
  2. The evaluation set. A rank-correlation ceiling is a property of the subset you scored on, not of the dataset. Same DrivAerML labels, same sigma: 0.9989 over all 484 designs, 0.6370 over the tightest 48-design window, -0.0100 over the twelve designs inside a one-count band. No row in this table reports its evaluation set's Cd spread. A vendor can move a headline from 0.99 to 0.57 without touching the model.
  3. The normalisation. "Relative L2" is a family, not a metric. AB-UPT reports surface-pressure rel L2 of 0.0142% on Luminary SHIFT-SUV; NVIDIA reports 0.10 for the same physical quantity on DrivAerML. Up to three orders of magnitude, and nobody states the normalisation in comparable terms.

2. The one conversion that exists

On DrivAerML, 1 automotive drag count = 1.9756 N (MEASURED: U = 38.889 m/s, A = 2.17 m2, paper-implied rho = 1.2041). The dataset states its own statistical convergence as ±1.5 drag counts at 95% CI, so sigma = 0.75 counts and the mean-absolute-error floor a perfect model would still show against these labels is sigma·sqrt(2/pi) = 0.598 counts.

Those two constants are the whole of the conversion column. Every row that has them is below; every row that does not says MISSING.

ModelOwnerAs publishedDrag counts (automotive, 1 ct = 0.001 Cd)Multiple of the labels' own MAE floor (0.598 ct)Eval nTag
X-MeshGraphNetNVIDIAMAE 15.23 N; max 58.16 N7.71 ct MAE; 29.44 ct max12.88x MAE floor; max 39.2 sigma48MEASURED
OLS on 16 design numbers (ours)oursR2 0.8882; rho 0.9426; MAE 4.65 ctMAE 4.65 ct7.77x MAE floor484 (CV)MEASURED (ours)
FIGConvNetNVIDIAMAE 8.86 N; max 25.72 N4.48 ct MAE; 13.02 ct max7.49x MAE floor; max 17.4 sigma48MEASURED
DoMINONVIDIAMAE 6.64 N; max 23.08 N3.36 ct MAE; 11.68 ct max5.62x MAE floor; max 15.6 sigma48MEASURED
UniversalAGI SUV-PTUniversalAGIMAPE 1.1%; max APE 3.4%3.08 ct (band 2.61-3.74)5.15x MAE floor (4.36-6.25x)UNKNOWNESTIMATED

The three NVIDIA rows are internally comparable - one paper, one 48-case test set, one metric - so the ordering among them is the source's own. The other two rows are not on that evaluation set and are not comparable to them or to each other: UniversalAGI does not publish its evaluation set at all, and our linear-regression bar is cross-validated over the complete 484.

Read the last column, not the fourth. Rank correlation says this field is finished: 0.96 to 0.99 everywhere. Absolute drag says the best published surrogate sits at 5.6x the labels' own repeatability, and the worst at 12.9x. Both statements are true about the same models.

3. The line the industry actually draws

The thresholds are external to us and published. OEM internal target: 1-2 drag counts on delta-Cd, "since every single count directly impacts the energy efficiency of a car". Regulatory (UN ECE R154 / WLTP): |delta(CdA)_CFD - delta(CdA)_EXP| <= 0.015 m2. Reference-data noise: ±1.5 counts, the dataset's own 95% CI.

Against that line, the best published absolute-drag error in this table is 3.36 counts - 1.7 to 3.4x the decision threshold, and 5.6x the labels' own MAE floor. No published aero surrogate is inside the engineering-acceptance band for absolute Cd. That is not a criticism of any model; it is the size of the remaining gap, and it is why the vendors' headline metric is rank correlation.

4. The table

Sorted by dataset, then by the published metric within each dataset. Lettered notes are in Section 9; every row carries its source.

DrivAerML - public, CC BY-SA 4.0, label sigma = 0.75 drag counts

The only corpus in this table with a published label-noise floor, so it is the only one where a "distance from the noise floor" column can be filled honestly.

ModelOwnerWhat was scoredAs publishedCount conventionIn drag countsvs the labels' own noiseEval nOpen?Train costTagSource
DoMINONVIDIAdrag MAE + max abs err (N)MAE 6.64 N; max 23.08 Nautomotive 1 ct = 0.001 Cd3.36 ct MAE; 11.68 ct max5.62x MAE floor; max 15.6 sigma48code YES~33 H100-hMEASUREDarXiv 2507.10747 [a]
FIGConvNetNVIDIAdrag MAE + max abs err (N)MAE 8.86 N; max 25.72 Nautomotive 1 ct = 0.001 Cd4.48 ct MAE; 13.02 ct max7.49x MAE floor; max 17.4 sigma48code YES~32 GPU-hMEASUREDarXiv 2507.10747 [a]
X-MeshGraphNetNVIDIAdrag MAE + max abs err (N)MAE 15.23 N; max 58.16 Nautomotive 1 ct = 0.001 Cd7.71 ct MAE; 29.44 ct max12.88x MAE floor; max 39.2 sigma48code YESUNKNOWNMEASUREDarXiv 2507.10747 [a]
OLS on 16 shipped design parameters - OUR trivial bar, not a model releaseoursdrag from 16 design numbersR2 0.8882; rho 0.9426; MAE 4.65 ctautomotive 1 ct = 0.001 CdMAE 4.65 ct7.77x MAE floor484 (CV)reproducibleseconds of CPUMEASURED (ours)our audit note [s]
DoMINONVIDIASpearman + R2 over designsdrag 0.99 / lift 0.98; R2 0.98 / 0.97no count convention (not a force metric)MISSINGUNDER ceiling 0.997748code YES~33 H100-hMEASUREDarXiv 2507.10747 [b]
FIGConvNetNVIDIASpearman + R2 over designsdrag 0.99 / lift 0.98; R2 0.97 / 0.95no count convention (not a force metric)MISSINGUNDER ceiling 0.997748code YES~32 GPU-hMEASUREDarXiv 2507.10747 [b]
X-MeshGraphNetNVIDIASpearman + R2 over designsdrag 0.96 / lift 0.81; R2 0.92 / 0.52no count convention (not a force metric)MISSINGUNDER ceiling 0.997748code YESUNKNOWNMEASUREDarXiv 2507.10747 [b]
DoMINONVIDIAsurface + volume rel L2aw-L2 p 0.08; raw 0.10; vol p 0.1042no count convention (not a force metric)MISSINGMISSING (no field floor)48code YES~33 H100-hMEASUREDarXiv 2507.10747 [c]
FIGConvNetNVIDIAsurface rel L2aw-L2 p 0.14; raw 0.21; wss-x 0.32no count convention (not a force metric)MISSINGMISSING (no field floor)48code YES~32 GPU-hMEASUREDarXiv 2507.10747 [c]
X-MeshGraphNetNVIDIAsurface rel L2aw-L2 p 0.14; raw 0.14; wss-x 0.17no count convention (not a force metric)MISSINGMISSING (no field floor)48code YESUNKNOWNMEASUREDarXiv 2507.10747 [c]
SUV-PT (LIFT architecture)UniversalAGICdA MAPE + max APEMAPE 1.1%; max APE 3.4%automotive 1 ct = 0.001 Cd3.08 ct (band 2.61-3.74)5.15x MAE floor (4.36-6.25x)UNKNOWNNOUNKNOWNESTIMATEDuniversalagi.com [d]
SUV-PT (LIFT architecture)UniversalAGISpearman over designs0.97no count convention (not a force metric)MISSINGUNDER ceiling 0.9989 / 0.9977UNKNOWNNOUNKNOWNMEASUREDuniversalagi.com [e]

DrivAerNet++ - public, CC BY-NC 4.0 - the CarBench leaderboard, verbatim

Surface-pressure regression at 10,000 sampled points. No drag, no lift: CarBench says so itself. Ranked here worst-to-best by relative L2 exactly as the source page ranks them - that ordering is the source's, not ours.

ModelOwnerWhat was scoredAs publishedCount conventionIn drag countsvs the labels' own noiseEval nOpen?Train costTagSource
AB-UPTCarBench (MIT + TRI)surface-p rel L2 @ 10k ptsrel L2 0.1358 +/-0.0024; R2 0.9675; RMSE 23.6no count convention (not a force metric)MISSING (no drag scored)MISSING (no field floor)1,154NO (0 code files)UNKNOWNMEASUREDdecode.mit.edu/carbench [f]
TransolverLargeCarBench (MIT + TRI)surface-p rel L2 @ 10k ptsrel L2 0.1457 +/-0.0025; R2 0.9595; RMSE 24.1no count convention (not a force metric)MISSING (no drag scored)MISSING (no field floor)1,154NO (0 code files)UNKNOWNMEASUREDdecode.mit.edu/carbench [f]
TransolverCarBench (MIT + TRI)surface-p rel L2 @ 10k ptsrel L2 0.1503 +/-0.0024; R2 0.9577; RMSE 24.6no count convention (not a force metric)MISSING (no drag scored)MISSING (no field floor)1,154NO (0 code files)UNKNOWNMEASUREDdecode.mit.edu/carbench [f]
Transolver++CarBench (MIT + TRI)surface-p rel L2 @ 10k ptsrel L2 0.1573 +/-0.0023; R2 0.9543; RMSE 25.6no count convention (not a force metric)MISSING (no drag scored)MISSING (no field floor)1,154NO (0 code files)UNKNOWNMEASUREDdecode.mit.edu/carbench [f]
TripNetCarBench (MIT + TRI)surface-p rel L2 @ 10k ptsrel L2 0.1608 +/-0.0024; R2 0.9590; RMSE 24.9no count convention (not a force metric)MISSING (no drag scored)MISSING (no field floor)1,154NO (0 code files)UNKNOWNMEASUREDdecode.mit.edu/carbench [f]
PointTransformerCarBench (MIT + TRI)surface-p rel L2 @ 10k ptsrel L2 0.1909 +/-0.0024; R2 0.9359; RMSE 30.3no count convention (not a force metric)MISSING (no drag scored)MISSING (no field floor)1,154NO (0 code files)UNKNOWNMEASUREDdecode.mit.edu/carbench [f]
RegDGCNNCarBench (MIT + TRI)surface-p rel L2 @ 10k ptsrel L2 0.2006 +/-0.0016; R2 0.9327; RMSE 30.9no count convention (not a force metric)MISSING (no drag scored)MISSING (no field floor)1,154NO (0 code files)UNKNOWNMEASUREDdecode.mit.edu/carbench [f]
PointNetLargeCarBench (MIT + TRI)surface-p rel L2 @ 10k ptsrel L2 0.2436 +/-0.0013; R2 0.9025; RMSE 37.2no count convention (not a force metric)MISSING (no drag scored)MISSING (no field floor)1,154NO (0 code files)UNKNOWNMEASUREDdecode.mit.edu/carbench [f]
PointMAECarBench (MIT + TRI)surface-p rel L2 @ 10k ptsrel L2 0.2713 +/-0.0016; R2 0.8791; RMSE 41.4no count convention (not a force metric)MISSING (no drag scored)MISSING (no field floor)1,154NO (0 code files)UNKNOWNMEASUREDdecode.mit.edu/carbench [f]
NeuralOperatorCarBench (MIT + TRI)surface-p rel L2 @ 10k ptsrel L2 0.3016 +/-0.0019; R2 0.8503; RMSE 46.2no count convention (not a force metric)MISSING (no drag scored)MISSING (no field floor)1,154NO (0 code files)UNKNOWNMEASUREDdecode.mit.edu/carbench [f]
PointNetCarBench (MIT + TRI)surface-p rel L2 @ 10k ptsrel L2 0.3803 +/-0.0020; R2 0.7639; RMSE 57.9no count convention (not a force metric)MISSING (no drag scored)MISSING (no field floor)1,154NO (0 code files)UNKNOWNMEASUREDdecode.mit.edu/carbench [f]

DrivAerNet++ - public, CC BY-NC 4.0 - everything else

The vendor rows here are the field's own honest report of what a proper hold-out costs.

ModelOwnerWhat was scoredAs publishedCount conventionIn drag countsvs the labels' own noiseEval nOpen?Train costTagSource
PXTransolver (~28M params)PhysicsXdrag Spearman ID -> OOD93.50% -> 69.10%no count convention (not a force metric)MISSINGUNDER 0.9475 by 0.0125 - closest callUNKNOWNNOUNKNOWNMEASUREDphysicsx.ai [g]
PXTransolver (~28M params)PhysicsXsurface-p rel L2 ID -> OOD13.47% -> 19.89%no count convention (not a force metric)MISSINGMISSING (no field floor)UNKNOWNNOUNKNOWNMEASUREDphysicsx.ai [g]
PXTransolver (~28M params)PhysicsXcross-simulator drag Spearman83.45% -> 8.35%no count convention (not a force metric)MISSINGNOT ASSESSABLE (no cross-code floor)UNKNOWNNOUNKNOWNMEASUREDphysicsx.ai [h]
Gradient-boosted trees on shipped design parameters - OUR trivial baroursdrag from design numbersMAE 9.9-14.4 ct (grouped CV)automotive 1 ct = 0.001 CdMAE 9.9-14.4 ct3.8-7.7x floor4,165 (grouped CV)reproducibleseconds of CPUESTIMATEDour audit note [s]

Luminary SHIFT-SUV - gated CC BY-NC 4.0; we obtained ZERO bytes (HTTP 401)

Four models, one corpus, and no label-noise statement anywhere - so no row here can be placed against a floor.

ModelOwnerWhat was scoredAs publishedCount conventionIn drag countsvs the labels' own noiseEval nOpen?Train costTagSource
AB-UPTEmmi AI (Mistral)surface/volume MAE + rel L2MAE p 6.47; rel L2 p 0.0142%, wss 5.97%, vel 2.83%no count convention (not a force metric)MISSINGNOT ASSESSABLEUNKNOWNcode YES13.5-22.3 H100-hMEASUREDarXiv 2510.15808 [j]
Transolver (as AB-UPT's baseline)re-run by Emmi AIsurface pressure MAEMAE 8.04no count convention (not a force metric)MISSINGNOT ASSESSABLEUNKNOWNarch. publishedUNKNOWNMEASUREDarXiv 2510.15808 [j]
DoMINO (as AB-UPT's baseline, dim 32)re-run by Emmi AIsurface pressure MAEMAE 42.85 (dim 32)no count convention (not a force metric)MISSINGNOT ASSESSABLEUNKNOWNcode YES35 GPU-h at dim 32MEASUREDarXiv 2510.15808 [k]
PXTransolver (~28M params)PhysicsXdrag Spearman + rel L2, ID -> OOD97.32% -> 83.45%; RL2 5.21% -> 9.57%no count convention (not a force metric)MISSINGNOT ASSESSABLEUNKNOWNNOUNKNOWNMEASUREDphysicsx.ai [g]

Luminary SHIFT-Wing - AIRCRAFT. Read the convention column before comparing anything here to a car

An aircraft drag count is 0.0001 Cd: ten times smaller than an automotive count. This is the trap that already caught us once.

ModelOwnerWhat was scoredAs publishedCount conventionIn drag countsvs the labels' own noiseEval nOpen?Train costTagSource
SHIFT-Wing reference model (NVIDIA PhysicsNeMo DoMINO / GeoTransolver architecture)Luminary Cloud% error on CL / CD / CMCD 1.73% median @M0.85; 0.81% @M0.50aircraft 1 ct = 0.0001 CdMISSING (no reference CD)NOT ASSESSABLE (0 bytes obtained)UNKNOWNdata YES (gated)UNKNOWNMEASUREDluminary.ai [l]

PXNetCar and SUV-Bench - proprietary corpora, no public labels

Nothing here can be checked by a third party. We print the numbers because the vendors published their own worst cases, which is more than most.

ModelOwnerWhat was scoredAs publishedCount conventionIn drag countsvs the labels' own noiseEval nOpen?Train costTagSource
PXTransolver (~28M params)PhysicsXdrag Spearman + rel L2, ID -> OOD81.02% -> 76.72%; RL2 15.12% -> 39.10%no count convention (not a force metric)MISSINGNOT ASSESSABLEUNKNOWNNOUNKNOWNMEASUREDphysicsx.ai [g]
SUV-PT (LIFT architecture)UniversalAGICdA MAPE / max APE / SpearmanMAPE 2.0%; max 5.6%; rho 0.96automotive 1 ct = 0.001 CdMISSING (no reference Cd)NOT ASSESSABLEUNKNOWNNOUNKNOWNMEASUREDuniversalagi.com [m]
SUV-PT (LIFT architecture)UniversalAGICdA MAPE / max APE / SpearmanMAPE 16.8%; max 49.1%; rho 0.54automotive 1 ct = 0.001 CdMISSING (no reference Cd)NOT ASSESSABLEUNKNOWNNOUNKNOWNMEASUREDuniversalagi.com [m]

AirfRANS - 2D airfoils, public, ODbL. 2D convention: 1 count = 0.0001 Cd

Included because the dataset's own published baseline anti-predicts drag, and a 1980s panel code beats it for $4.7e-7 a case.

ModelOwnerWhat was scoredAs publishedCount conventionIn drag countsvs the labels' own noiseEval nOpen?Train costTagSource
GraphSAGE (the dataset's own reference baseline)AirfRANS authorsSpearman on dragrho = -0.303 +/- 0.1242D airfoil 1 ct = 0.0001 CdMISSINGUNKNOWN (no published sigma)UNKNOWNdataset ODbLUNKNOWNMEASUREDarXiv 2212.07564 [o]
XFOIL 6.99 (a 1980s panel code) - OUR reference bar, not a foundation modeloursSpearman on drag; residual vs RANSrho +0.8794 -> +0.9949 (Ncrit 0.05); MAPE 2.3%2D airfoil 1 ct = 0.0001 Cdresidual 36.8 -> 1.66 ctUNKNOWN (no published sigma)AirfRANS test setXFOIL is publicnone - it is a solverMEASURED (ours)arXiv 2212.07564 [o]

Models with NO accuracy number we could source

A row with no number is still a row. We print it rather than leave the model out.

ModelOwnerWhat was scoredAs publishedCount conventionIn drag countsvs the labels' own noiseEval nOpen?Train costTagSource
LGM-Aero (~100M params)PhysicsXnothing sourceableMISSINGno count convention (not a force metric)MISSINGMISSINGUNKNOWNNOUNKNOWNUNVERIFIEDphysicsx.ai [i]
GeoTransolverAnsys (Synopsys)efficiency only (a rival's measurement)P50 23.6 s; 0.41 samples/s; 23.0 GiBno count convention (not a force metric)MISSING (no accuracy number)MISSINGUNKNOWNarch. publishedUNKNOWNUNVERIFIEDuniversalagi.com [n]
SimAIAnsys (Synopsys)drag error vs CFD'<0.5% (5 to 10 drag counts)'convention NOT STATED (ambiguous)MISSING (halves disagree)NOT ASSESSABLEUNKNOWNNOUNKNOWNUNVERIFIEDlinkedin (UNVERIFIED) [r]

PDE and geophysical foundation models - the aero null

These are real models with real results on their own axes. None of them reports an integrated drag or lift error on a car or a wing in anything we hold. That absence is MEASURED, and it is why they cannot be compared to the rows above.

ModelOwnerWhat was scoredAs publishedCount conventionIn drag countsvs the labels' own noiseEval nOpen?Train costTagSource
Poseidon (T / B / L)ETH Zurichno aero-relevant numberMISSING for any aero quantityno count convention (not a force metric)MISSINGMISSINGn/a to an aero claimYES176-1,320 GPU-hMEASURED (null)arXiv 2405.19101 [p]
DPOT (up to 1B params)Tsinghuano aero-relevant numberMISSING for any aero quantityno count convention (not a force metric)MISSINGMISSINGn/a to an aero claimYESUNKNOWNMEASURED (null)arXiv 2403.03542 [p]
MPP (Multiple Physics Pretraining)Polymathic AIno aero-relevant numberMISSING for any aero quantityno count convention (not a force metric)MISSINGMISSINGn/a to an aero claimYESUNKNOWNMEASURED (null)arXiv 2310.02994 [p]
Walrus (1.3B params)Polymathic AIno aero-relevant numberMISSING for any aero quantityno count convention (not a force metric)MISSINGMISSINGn/a to an aero claimYESUNKNOWNMEASURED (null)arXiv 2511.15684 [p]
AuroraMicrosoft Researchno aero-relevant numberMISSING for any aero quantityno count convention (not a force metric)MISSINGMISSINGn/a to an aero claimYESUNKNOWNMEASURED (null)nature.com [p]
FNO / TFNO / U-net / CNextU-net baselinesPolymathic AIVRMSE (predict-the-mean = 1.0)8 of 66 cells worse than the mean; worst 3.447no count convention (not a force metric)MISSINGthe trivial floor IS the metric66 cellsYES12 GPU-h capMEASUREDarXiv 2412.00568 [q]

5. The null, stated plainly

We built an instrument to flag any published claim that exceeds its dataset's own label-noise ceiling - a claim that cannot be true, because a perfect model scored against noisy labels cannot beat the noise. We ran it over every claim we had collected. It fires zero times. We hoped for a positive and did not get one, so we publish the null.

ClaimDatasetEval nPublishedCeiling usedVerdict
DoMINODrivAerML48Spearman 0.990.9977 (random-48) / 0.9882 (extreme-48)UNDER
FIGConvNetDrivAerML48Spearman 0.99sameUNDER
X-MeshGraphNetDrivAerML48Spearman 0.96sameUNDER
DoMINO / FIGConvNet / X-MeshGraphNetDrivAerML48R2 0.98 / 0.97 / 0.920.9981UNDER
UniversalAGI SUV-PTDrivAerMLUNKNOWNSpearman 0.970.9989 full / 0.9977 at n=48UNDER
PhysicsX (in-distribution)DrivAerNet++UNKNOWNSpearman 0.93500.9475 worst case (N_eff = 1)UNDER - closest
PhysicsX (out-of-distribution)DrivAerNet++UNKNOWNSpearman 0.6910as aboveUNDER
UniversalAGI SUV-Bench Med / Hard; PhysicsX on Luminary SHIFT-SUVproprietaryUNKNOWN0.96 / 0.54 / 0.9732none existsNOT ASSESSABLE

The closest call: PhysicsX's in-distribution drag Spearman of 0.9350 sits 1.25 points under 0.9475, the most pessimistic bar for that dataset - and that bar needs DrivAerNet++'s shipped scatter column (Std Cd) to be a standard error of the mean, which it is not (it is iteration scatter over a 1000-iteration window). On DrivAerML, DoMINO and FIGConvNet sit within 0.008 of the labels' own self-consistency, so optimising that metric further fits solver noise.

6. We do not rank the vendors, and that is the point

A cross-vendor ranking is not supportable from these numbers. Six independent reasons, all visible in the table above:

What is supportable: the three NVIDIA models against each other on their shared 48 cases, the eleven CarBench models against each other on their shared 1,154, and every absolute-drag row against the ground truth's own repeatability. Those are the comparisons this page makes. It makes no others.

7. What it costs to train one

Where a training budget is published. Note the last line: on these corpora the ground truth costs more than the model.

What was trainedOnCostBasis
DoMINO (surface + volume)DrivAerML~33 H100-GPU-hESTIMATED from 4.2 h x 8 H100, NVIDIA docs
FIGConvNetDrivAerNet / DrivAerML~32 GPU-hMEASURED wall: 16 h x 2 A100, stated by its authors
X-MeshGraphNetDrivAerMLUNKNOWNthe paper defers to the PhysicsNeMo repo
AB-UPTLuminary SHIFT13.5-22.3 GPU-h on one H100MEASURED
DoMINO as a third-party baselineLuminary SHIFT-SUV35 GPU-h at dim 32MEASURED - and an unequal budget, see note (k)
The Well baselines (FNO / TFNO / U-net / CNextU-net)The Well, 17 systems12 GPU-h cap per model x systemMEASURED: their published protocol
Poseidon-T / -B / -L pretrainPDEgym176 / 944 / 1,320 GPU-hMEASURED wall on 8x RTX 4090; GPU-hours ESTIMATED. CNO-FM 1,424
Transolver / Transolver++ / GeoTransolvervariousUNKNOWNnot stated in any source we hold
CarBench's 11 modelsDrivAerNet++UNKNOWNno training code or configs released
UniversalAGI, PhysicsX, Luminary, Ansys, SiemensproprietaryUNKNOWNno vendor states a training budget
one WindsorML CFD case, for scaleWindsorML~224 GPU-hMEASURED: 28 h on 8x A10G - the ground truth costs more than most of the models

8. Corrections, and what we could not source

The aircraft-versus-automotive trap, and an ambiguity in our own correction

Our earlier note recorded a cross-code spread of "4-5 drag counts" from a DPW-6 statistical analysis and compared it against automotive numbers. Those are aircraft counts. The correction we published says the automotive equivalent is "~11x larger". Both cannot be read the same way: converting the absolute Cd gives 0.4-0.5 automotive counts (ten times smaller as a number), while converting the relative error gives roughly 11x more, because a road car's Cd is about eleven times a transport aircraft's cruise CD. Neither our file nor the source prints the reference CD that reconciles them, so we do not publish a cross-code floor in automotive counts at all. Where a cross-simulator row needs a floor, this page says NOT ASSESSABLE and quotes instead the measured automotive cross-code spread on DrivAer Cd: 49, 54 and 56 counts at AutoCFD 2, 3 and 4.

An attribution error in our own audit note

Its null table labels the 93.50% → 69.10% and 83.45% → 8.35% rows "PhysicsX LGM-Aero". The source post attributes those numbers to PXTransolver (~28M parameters), a different model from LGM-Aero (~100M parameters, pre-trained on 25M+ meshes), for which no accuracy number is published at all. This page uses the source's attribution. The audit note's row labels are wrong and we are saying so rather than quietly fixing them.

A vendor claim that disagrees with itself

Ansys SimAI is quoted as "drag error compared to CFD is less than 0.5% (5 to 10 drag counts)". On a road car at Cd 0.28, 0.5% is 1.4 automotive counts, not 5-10. The two halves are self-consistent only under the aircraft convention at Cd 0.10-0.20, or for a body with Cd of 1.0-2.0. We do not guess which was meant, so that row's conversion is MISSING. The figure is also UNVERIFIED: it appears in a LinkedIn post, not on a vendor-owned page.

A rounding difference in our own constant

Recomputing the force scale from the published inputs gives 1.9758 N per automotive count; our audit note prints 1.9756 N. The gap is 0.011% and moves no digit on this page. Both values are recorded.

What we could not source, listed rather than guessed

X-MeshGraphNet's training cost. The evaluation-set size and spread for every commercial row - UniversalAGI, PhysicsX, Luminary, Ansys. The reference CD behind Luminary's SHIFT-Wing percentages, without which those percentages cannot become counts. Any accuracy number for PhysicsX LGM-Aero or for GeoTransolver. Training budgets for Transolver, Transolver++, GeoTransolver and all eleven CarBench models. A field-level label-noise floor for any corpus in this table. WindsorML's uncertainty, which does not exist: a full-text search of its paper for "drag count", "95%", "confidence" and "standard deviation" returns zero hits. And an aero-relevant integrated-force number for Poseidon, DPOT, MPP, Walrus or Aurora - searched, not found, reported as a null rather than as a zero.

9. Notes

(a) NVIDIA's 48-case test mixes in-distribution cases with the lowest- and highest-drag cases, so the applicable rank ceiling sits between random-48 (0.9977) and extreme-48 (0.9882). The conversion is ours: on DrivAerML 1 automotive drag count = 1.9756 N (U = 38.889 m/s, A = 2.17 m2, rho = 1.2041), and at sigma = 0.75 counts the label-noise MAE floor is sigma*sqrt(2/pi) = 0.598 counts. Recomputed independently on a clean machine: 1.9758 N, a 0.011% difference that moves no digit reported here.

(b) A rank statistic carries no force units, so its drag-count cell is MISSING, never 0. And no row in this whole table reports the Cd spread of the set it was scored on. Same DrivAerML labels, same sigma: the ceiling is 0.9989 over all 484 designs, 0.6370 over the tightest 48-design window, and -0.0100 once only the twelve designs inside a one-count band remain. Without the evaluation set's spread, two Spearman numbers are not comparable.

(c) No field-level label-noise floor is published for any corpus in this table. DrivAerML does ship a pressure-variance map: area-weighted Cp_rms is 3.93x larger on the rear-facing base than on the upper surface, and a single whole-surface relative L2 averages that away.

(d) A percentage error has no units until a Cd is chosen. We convert at Cd 0.28 (3.08 counts) and print the band over the corpus's measured Cd range, 0.2370 to 0.3401 (2.61 to 3.74 counts). That assumption is why the row is tagged ESTIMATED and not MEASURED.

(e) Evaluation set unknown, so note (b) applies at full strength: this number is not comparable to any other Spearman in the table.

(f) Copied verbatim from the leaderboard page's own hardcoded model-data array (const modelData). The +/- is a test-set bootstrap CI, not seed variance: "each model in this study was trained only once". CarBench scores no drag and no lift - its own Limitations call them future extensions. MSE and RMSE are in kinematic-pressure units (m2 s-2), not Pa, because the solver is simpleFoam. MEASURED 2026-08-25: the benchmark repo holds 15 entries and 0 code files, README.md is 84 bytes, there is no licence, and the page has no submission path.

(g) Vendor post: no split IDs, no code, no evaluation-set spread. The ID to OOD pair is the vendor's own report of what a proper hold-out does to its headline.

(h) A cross-simulator claim needs a cross-code floor and we do not have a defensible one in automotive counts - see the corrections section. The automotive cross-code spread on DrivAer Cd is 49, 54 and 56 counts at AutoCFD 2, 3 and 4.

(i) No accuracy number is published for this model itself.

(j) Relative L2 is a family of normalisations, not one metric. AB-UPT's surface-pressure rel L2 of 0.0142% and NVIDIA's 0.10 on DrivAerML differ by up to three orders of magnitude because the normalisation differs. Do not read them as the same axis. AB-UPT is also the only entry in this table that reports a median over five seeds.

(k) Unequal budget, and it is documented: DoMINO was run at dim 32 for 35 GPU-h while the transformer baselines ran at dim 192. NVIDIA's own paper reports DoMINO as the best of three on DrivAerML. Both can be true; neither is comparable.

(l) AIRCRAFT, not automotive. An aircraft drag count is 0.0001 Cd, ten times smaller than an automotive count. The reference CD is not published in any source we hold, so a percentage cannot become counts here: MISSING.

(m) Proprietary evaluation set, no published reference Cd, no label-noise statement. We quote the Hard split because it is a vendor publishing its own collapse, which is rarer than it should be.

(n) No accuracy number found on disk. The latency and memory figures are a competitor's measurement at matched parameter count, not the authors'.

(o) AirfRANS publishes no convergence or label-noise specification, so there is no floor to divide by: UNKNOWN, not 0. 2D airfoil convention throughout: 1 count = 0.0001 Cd. The XFOIL row is ours and it is a solver, not a foundation model.

(p) MEASURED absence, not zero. We searched our on-disk corpus and found no aero-relevant integrated-force number for these models. They report real results on their own axes; those axes are not drag.

(q) The Well builds the trivial baseline into the metric: VRMSE = 1.0 is exactly predict-the-mean. 8 of 66 published cells are worse than that. Two further cells are MISSING and are excluded from the denominator rather than counted as 0.

(r) The claim's two halves do not agree under one convention - see the corrections section for the arithmetic.

(s) Ours, not a vendor claim. Published so that every table has a floor to read against: on DrivAerML, ordinary least squares on 16 shipped design numbers lands at 4.65 counts, between FIGConvNet (4.48) and X-MeshGraphNet (7.71).


Sources and artifacts

The table as data: research/artifacts/fluid-foundation-models.csv (46 rows, 14 columns, one row per published result, every row carrying its full-length caveat text and its source URL). Published numbers are transcribed from the sources linked in each row; the CarBench rows are copied verbatim from that same hardcoded array. The drag-count conversions and the noise-floor multiples are ours and were recomputed from the published inputs on a clean 4-CPU machine in about 10 seconds; the script is spike/sota_verify.py. The label-noise ceilings, the 1.9756 N force scale and the trivial-baseline rows come from our audit note, What CFD-surrogate benchmarks actually measure, whose own source map is published at audit-source-map.csv.

Licences of the corpora referenced

DrivAerML, AhmedML and WindsorML CC BY-SA 4.0; DrivAerNet++ CC BY-NC 4.0 (non-commercial); Luminary SHIFT CC BY-NC 4.0 and gated; AirfRANS ODbL; The Well public. SUV-Bench, PXNetCar and every Ansys, Siemens and Neural Concept figure rest on data no third party can obtain.

The audit note this table is the constructive half of · the scalar-drag instrument · all research