What CFD-surrogate benchmarks actually measure: convention, spread, offset, and mesh
For teams publishing on DrivAerML, DrivAerNet++, AhmedML and WindsorML, and vendors quoting Spearman against them.
Two drag-count conventions appear in this note and each is restated at every use. For road cars, 1 count = 0.001 Cd. For 2D airfoils, 1 count = 0.0001 Cd. Every number carries a tag: MEASURED, ESTIMATED (under a stated assumption), or UNKNOWN. Missing data is written as MISSING and never as 0, including in denominators. "Not reported" means we did not find it after a targeted search. These are easy mistakes, and several of them are naming choices rather than science. Four of our own numbers were wrong; they are retracted in Section 4.
1. The claim
Published numbers on these benchmarks are not comparable across papers. Three reasons are each sufficient on their own, and a fourth, numerical one is now measured in 2D.
- Convention. Each dataset ships two ground-truth force files that differ only in the reference area used to normalise drag, and switching files reorders the designs. On DrivAerML 24.85% of design pairs flip, and the two truths agree on 0 of the 5 best designs (MEASURED, 484 runs).
- Spread. A rank-correlation ceiling belongs to the evaluation set, not to the dataset. Same sigma, same labels, a different subset: the DrivAerML ceiling moves 0.9989 → 0.6370 (MEASURED, complete 484). No paper reports its evaluation set's Cd spread.
- Offset. The out-of-distribution collapse is about 86% a constant level error. An oracle mean shift restores R2 from 0.2078 to 0.745 (MEASURED).
- Mesh. Every ceiling here is statistical. The discretisation term is larger, and it is not purely common-mode (A7).
Two further findings are reported here for the first time. The corpora publish a 4x noise map that the metric everyone uses throws away (A8), and the newest official out-of-distribution split is drag-confounded (A9).
We built an instrument to flag claims that exceed a dataset's own label-noise ceiling. It fires zero times (Section 3). Four of our own conclusions were falsified; they are published, not edited away.
2. The findings
A1. The reference-area convention flips ranks; frontal-area spread sets how badly
Re-derived from each repository's authoritative aggregate label files, joined on run id. MEASURED, road-car convention, 1 count = 0.001 Cd.
| dataset | n | pairs | pairs flipped (%) | Spearman A-B | Kendall | top-5 agreement | frontal-area spread |
|---|---|---|---|---|---|---|---|
| DrivAerML | 484 | 116886 | 24.85 | 0.6912 | 0.5030 | 0 of 5 | 48.2% (1.779-2.636 m2 vs const 2.17) |
| AhmedML | 499 | 124251 | 30.34 | 0.5588 | 0.3932 | 1 of 5 | const 0.112 m2 |
| WindsorML | 355 | 62835 | 5.34 | 0.9828 | 0.8932 | 5 of 5 | 5.4% (0.1127-0.1188 m2 vs const 0.1120) |
Top-5 agreement is agreement on the 5 lowest-drag designs, which is the engineering decision. The mean absolute difference between the two conventions is 17.6 counts on DrivAerML, with a maximum of 61.4, and the relation between them is algebraic: Cd_constref = Cd_pergeom x (aRef/aRefRef) holds to a maximum residual of 9.99e-08 over 484 runs. WindsorML gives the general law. Its frontal area varies by 5.4% and it flips 5.34% of pairs; DrivAerML's varies by 48.2% and it flips 24.85%. The hazard is set by frontal-area spread in the design space.
The naming inversion is the part that bites. The default per-run force file (force_mom_i.csv) uses per-geometry area in DrivAerML but constant area in AhmedML and WindsorML: same author, same layout, opposite meaning, confirmed verbatim in the three READMEs. Taking the default file everywhere is the natural thing to do and it is wrong for two of the three. Inside a single dataset, "Spearman over designs" is defined only to about ± 0.45 by a choice of filename.
A2. The Spearman ceiling belongs to the evaluation set, not the dataset
A label-noise ceiling is what a perfect surrogate scores against a noisy test label, so it depends on label noise relative to the spread of the scored subset. Method: y = t + e, with e drawn from N(0, sigma^2); t by empirical-Bayes shrinkage; the ceiling is E[Spearman(t, y_new)] over 600-1000 Monte-Carlo repetitions. DrivAerML's sigma of 0.75 counts is MEASURED from its stated 95% confidence interval of ± 1.5 counts. Road-car convention, 1 ct = 0.001 Cd.
| dataset | subset | n | sd of Cd (ct, 1ct=1e-3) | sigma (ct) | Spearman ceiling | tag |
|---|---|---|---|---|---|---|
| DrivAerML | full | 484 | 17.56 | 0.75 | 0.9989 | MEASURED (complete) |
| DrivAerML | random 48 | 48 | 16.43 | 0.75 | 0.9977 | MEASURED (complete) |
| DrivAerML | extreme 48 = 24 lowest + 24 highest | 48 | 36.30 | 0.75 | 0.9882 | MEASURED (complete) |
| DrivAerML | tightest 48 window | 48 | 0.98 | 0.75 | 0.6370 | MEASURED (complete) |
| DrivAerML | all within a 1 ct band of the median | 12 | 0.25 | 0.75 | -0.0100 | MEASURED (complete) |
| DrivAerNet++ | full | 4165 | 22.02 | 2.01 | 0.9952 | ESTIMATED (N_eff = 10) |
| DrivAerNet++ | middle 200 band | 200 | 0.81 | 0.66 | 0.4258 | ESTIMATED (N_eff = 100) |
These are recomputed on the complete 484 runs after correction C7, with 4,000 Monte-Carlo repetitions; the earlier 442-row values were 0.9988, 0.9878 and 0.5720. The correction moves nothing in direction and the effect is slightly larger: 0.9989 across the whole drag range, 0.6370 over the tightest 48-design window, and -0.0100 once only the twelve designs inside a one-count band remain. Same labels, same solver, same sigma. The ceiling is a property of the subset you chose to score on.
A3. Correlation says the field is finished; absolute error says 5-13x
On DrivAerML, 1 drag count = 1.9756 N (MEASURED: U = 38.889 m/s, A = 2.17 m^2, paper-implied rho = 1.2041; the paper states T = 293.15 K and never prints rho). At sigma = 0.75 counts the label-noise MAE floor is sigma*sqrt(2/pi), or 0.598 counts. Published MAE and maximum AE are from NVIDIA arXiv:2507.10747; the conversion to counts is ours. Road-car convention, 1 ct = 0.001 Cd.
| model | published MAE (N) | MAE (ct, 1ct=1e-3) | x above noise floor | published max AE (N) | max AE (in sigma) | tag |
|---|---|---|---|---|---|---|
| X-MeshGraphNet | 15.23 | 7.70 | 12.87x | 58.16 | 39.2 | MEASURED (conversion) |
| FIGConvNet | 8.86 | 4.48 | 7.49x | 25.72 | 17.3 | MEASURED (conversion) |
| DoMINO | 6.64 | 3.36 | 5.61x | 23.08 | 15.6 | MEASURED (conversion) |
| our GBM on DrivAerNet++, grouped CV | not published | 9.9-14.4 | 3.8-7.7x | not published | not published | ESTIMATED (N_eff = 100-10) |
Rank correlation saturates while absolute drag, the metric that decides anything, sits 5-13x above the labels' own repeatability.
A4. Ranking the top designs is at the edge of what the labels support
Road-car convention, 1 count = 0.001 Cd.
| dataset | n | median adjacent gap (ct) | best-to-2nd gap (ct) | smallest resolvable difference, 95% (ct) | tag |
|---|---|---|---|---|---|
| DrivAerNet++ | 4165 | 0.0144 | 1.0757 | 17.65 at N_eff = 1; 5.58 at N_eff = 10 (central); 1.77 at N_eff = 100 | MEASURED (gaps); ESTIMATED (resolution) |
| DrivAerML | 484 | 0.0943 | 0.5204 | 2.08 | MEASURED (complete) |
| AhmedML | 499 | 0.1280 | 3.7349 | 3.60 | MEASURED (complete); ESTIMATED (sigma) |
| WindsorML | 355 | 0.1882 | 3.9624 | UNKNOWN | MEASURED (complete); no published sigma |
The probability that the true best design is ranked first is 0.5295 on DrivAerNet++ at the central N_eff of 10, falling to 0.1003 at N_eff = 1 and rising to 0.9513 at N_eff = 100 (MEASURED, 4,000 Monte-Carlo repetitions). In the dataset's favour: using the top two designs' own shipped scatter column (Std Cd), which reads 4.218 and 2.830 counts, the 95% least significant difference at N_eff = 100 is 1.00 count against a gap of 1.0757 counts.
DrivAerML here is a retraction (C7). On complete labels that same probability is 0.6893, not the 0.998 we published. The true second-best is run 159, one of the rows our own download had dropped, and the top-two gap collapses from 3.23 to 0.52 counts. AhmedML is the one dataset that separates its top two, and only just: 3.7349 counts against a least significant difference of 3.60 counts. Top-1 selection difficulty grows with dataset size, so any paper claiming to identify the optimal design should report the top-two gap against label noise on the complete label set.
A5-A6. The default file leaks frontal area, and one dataset is solved by linear regression
GBM and OLS on the shipped design parameters, predicting Cd, using the per-geometry-area file in every row and corrected label counts. Two keys were used: a random five-fold split, which is not leak-robust for a design of experiments, and a grouped five-fold split on KMeans clusters of the standardised parameters (k = 5). MEASURED.
| dataset | n | parameters | rho, random 5-fold | rho, leak-robust | our first draft |
|---|---|---|---|---|---|
| WindsorML | 350 | 7 | 0.5944 | 0.4397 | 0.593 |
| DrivAerNet++ | 4165 | 23 | 0.7886 | 0.5593 (body-style holdout) | 0.789 |
| AhmedML | 499 | 8 | 0.8148 | 0.3029 | 0.917 RETRACTED |
| DrivAerML | 484 | 16 | 0.9503 | 0.9457 | 0.952 |
On DrivAerML, 16 numbers and ordinary least squares reach R2 0.8882, rho 0.9426 and MAE 4.65 counts (1 ct = 0.001 Cd), and this survives the leak-robust key at 0.8745 and 0.9353.
The headroom ranking inverts, and the error was ours (C7). We first put AhmedML at rho 0.917 and called it "nearly solved". That came from the leaky constant-area file: normalising by a constant area leaves the per-geometry frontal area inside the target. Spearman between area and Cd is +0.7755 on AhmedML with constant area, against +0.0059 with per-geometry area; the same pair on DrivAerML is +0.8230 against +0.1893, and DrivAerNet++ is clean at -0.0340. On the per-geometry file AhmedML scores 0.8148, and under the leak-robust key it collapses to rho 0.3029 and R2 -0.141, worse than predicting the mean. That makes AhmedML the hardest of the four, not the easiest. WindsorML's 0.5944 needs a seventh parameter, the front-to-rear length ratio (ratio_length_front_rear), which is absent from the published aggregate file; with the published six it scores 0.5362.
Trivial-rule bars, all MEASURED: the body-style label alone scores 0.5057 on DrivAerNet++; body height times body width scores 0.7642 on AhmedML with its default file; linear regression scores 0.9426 on DrivAerML.
A7. The two-mesh design sweep: we ran it in 2D. In 3D it is still open.
Every ceiling above is statistical, so it is a lower bound on label error. Published grid sensitivity is 20-30x larger: DrivAerNet++'s relative Cd error against the TUM reference runs 8.21% at 6M cells, 6.55% at 12M and 2.17% at 24M (arXiv:2406.09624, T7). Grid error is usually assumed to be common-mode, that is, a bias rather than a scrambling. We tested that in 2D and it is half true. The study covers 12 cases (NACA 4-digit, Re 2-6e6, alpha -2 to 12 deg) on 3 systematically refined 6-block C-grids of 32,256, 72,576 and 163,296 cells, with a refinement ratio of exactly r = 1.5000 and grading frozen across levels, solved with simpleFoam and k-omega SST wall-resolved (y+ max 1.2-2.6): 66 runs, about 46 CPU-hours. The 2D convention applies in this block: 1 count = 0.0001 Cd.
| estimator | cases | median (ct, 2D: 1ct=1e-4) | p90 (ct) |
|---|---|---|---|
| U_C assumption-free (1.25 x half-spread over 3 levels) | 6 | 2.20 | 3.83 |
| U_B Roache fallback (2-grid; assumed p = 2; Fs = 3) | 7 | 2.93 | 6.25 |
| U_A textbook 3-grid GCI with OBSERVED p | 2 | 14.59 | 19.64 |
The observed order of convergence p was computable in only 2 of 12 cases, at 0.97 and at 0.19, the second flagged as unphysical. 4 cases are oscillatory and 6 never reached a usable iterative band in 3,000-3,500 iterations. Those are MISSING, never 0: the denominators are the cases column.
Two results follow. First, the label-noise floor the field quotes is the wrong quantity. Iterative convergence on this solver family measures 0.001 counts, while discretisation uncertainty is 2.20 counts assumption-free and 2.93 counts on the Roache fallback, about 2,200x larger. Quoting iterative convergence as a label-noise floor therefore understates numerical uncertainty by more than three orders of magnitude. Second, mesh error is not purely common-mode. The coarse-to-fine mean shift is -3.94 counts, but the standard deviation of the per-design shift is 1.55-2.81 counts, with one genuine rank swap: Spearman between levels is 1.0000 from L1 to L2 and then 0.9643, at n = 7. So the design-dependent term that bounds every ceiling here is real and, in 2D, larger than the 1.66-count effect we had proposed to model.
Threats this lane states against itself: maximum mesh non-orthogonality is 65.7 deg, so a smoother C-grid could show cleaner p and smaller uncertainty; and three medium-to-fine deltas, of 0.001-0.13 counts, sit below the study's own iterative band of 0.30-0.63 counts, so those may be underestimates. The defensible claim is the order of magnitude: a few drag counts, not a thousandth of one. For the four road-car datasets the design-dependent mesh term stays UNKNOWN.
The method finding may be worth more than the number. 18 solver configurations were probed. SIMPLEC at 0.9 relaxation, and even at 0.5/0.5, holds a plausible Cd plateau for about 1,200 iterations and then collapses into a ±50-count limit cycle (2D convention), at the coarse and the medium level, so it is the relaxation and not the grid. Only 0.4/0.4/0.4 with cellLimited gradients and one non-orthogonal corrector survived.
Any grid study that does not print the full drag history is not checkable.
A8. DrivAerML publishes its own noise map, and a whole-surface relative-L2 hides it
DrivAerML ships the pressure-variance field itself (pPrime2MeanTrim) on the surface, at 35.3 MB per case, and in the volume. We did not find it used as an error weight in any surrogate paper after a targeted search. Area-weighted Cp_rms on run_1, MEASURED:
| corpus | cases | base/upper Cp_rms ratio | spread |
|---|---|---|---|
| DrivAerML | 484 | 3.093 (median) | IQR 0.67; range 2.03-4.52; only 5.0% of cases exceed 3.93 |
| WindsorML | 349 | 1.74 (median) | not stated |
| run_1 alone, the case first examined | 1 | 3.93 | a p95 outlier, not the typical case |
A single whole-surface relative-L2 hides a 3.1x heteroscedasticity, ranging 2.0-4.5x across the corpus, that the dataset itself publishes. This was corrected before publication: an earlier draft quoted 3.93x from a single case, and measuring all 484 showed that case is a p95 outlier while the corpus median is 3.093. AhmedML ships no pressure variance at all.
Two limitations we state rather than bury. This is a lower bound on the ratio of standard errors of the mean, because the integral time scale is longer in the wake. And a variance map without a correlation length under-reports field-integral uncertainty by more than 6x: the independent-face bound gives 0.25 counts (1 count = 0.001 Cd) against the dataset's own ±1.5-count tolerance, so that tolerance is set by coherent wake unsteadiness. The artifact that would close this is a per-case force-coefficient history (forceCoeffs) of about 100 KB per case. There is none, and the convergence history ships only as a PNG. 3D discretisation error is likewise unmeasurable from the shipped data, because there is one mesh per case.
A9. The new official splits are good, and one of them is drag-confounded
DrivAerML published 8 deterministic splits on 2026-08-17, ungated, under CC BY-SA 4.0, and they hold. Counts match the README exactly: 400/34/50 for the full split, and 339/48/97 each for the geometry, high-drag, low-drag and rear-separation splits. Train/test, train/validation and validation/test overlap is zero in all 8 families, the nesting from super-scarce through scarce and medium to full is true, and the published test set reproduces exactly from the committed score. The splits cover 484 cases, not 500, independently confirming C7. The rear-separation split is genuinely physics-defined, taken from wake area in images rather than from an integrated coefficient.
But it is drag-confounded, and the README does not say so. Spearman between score and Cd is -0.5594, the AUC of Cd for test against train is 0.2619, and 41 of its 97 test cases are also in the low-drag test set.
A model that merely under-predicts drag will look as though it fails on separated wakes.
Use the geometry split, whose AUC is 0.538 and which is drag-neutral, as the primary out-of-distribution split; or publish the rear-separation split beside a drag-matched control. This is an easy mistake on splits published days ago.
And the fix does not transfer, which we only learned by measuring all three corpora. Re-measuring the AUC of Cd for test against train across all 24 split families: DrivAerML's geometry split reproduces at 0.5381 and is drag-neutral, but AhmedML's is 0.7164, with 41 of its 100 test cases also in the high-drag test set, and WindsorML's is 0.6046. The image-wake split is worse still, at 0.7713 and 0.6529. The geometry split is drag-neutral only on DrivAerML. A benchmark that adopts it as a universal primary split inherits a drag confound on two corpora out of three.
Separately, on all three corpora the super-scarce training set is drag-shifted high by 19 to 37 counts against its own test set, so the data-efficiency ladder confounds less data with different data. We did not find this reported.
Two further errata were found while extracting the surfaces. WindsorML runs 350-354 have no directory on the host, and run_354 appears in six test folds, so every published WindsorML test denominator is one too large (full test 36 to 35; high-drag test and low-drag test 71 to 70). And the per-run force files are variable-reference on DrivAerML but constant-reference on AhmedML, at 0.112032, and on WindsorML, at 0.112 — the same naming inversion as A1, now confirmed on the per-run files as well as the aggregates.
Supporting corrections we also owe the field
All MEASURED, from primary sources or first-hand fetches.
- Air density: two values, one real hazard, and not the one we first published. DrivAerNet++ states a density of 1.184 kg/m3 at 298.15 K; PhysicsX use 1.204 for Luminary SHIFT-SUV at 293.15 K. The ideal-gas relation p/RT reproduces both, at 1.1839 and 1.2041, so this is 25 C air versus 20 C air. Cd is dimensionless, so density cannot shift it: rescaling every DrivAerML Cd by 1.016892 leaves Spearman at exactly 1.0 and flips 0.00% of pairs. The 4.31-count figure, at Cd 0.255 with 1 count = 0.001 Cd, is only the error from converting a dimensional result with the wrong density, which is what PhysicsX describe while stating both densities. No CAE-ML dataset paper prints density at all, and the velocity term in the same sentence, 30 versus 31.298 m/s, is 6x larger.
- The DrivAerNet++ file named "Updated" carries no correction: it is bit-identical to the Average Cd column on all 4,165 parametric rows, agreeing to better than 1e-9 counts. See Section 6.
- Label counts match the papers, with one gap: DrivAerML 484 of 500, AhmedML 500 of 500, WindsorML 355 of 355. Our earlier lower counts were our own bug (C7). The S3 route printed in the papers is dead and returns AllAccessDisabled; use the HuggingFace mirrors listed in Sources.
- Smaller errata, first reported here: WindsorML's aggregate parameter file drops the front-to-rear length ratio; AhmedML run 500 is in one force file only; WindsorML runs 350-354 are aggregate-only and return 404 per run. WindsorML uses y as the vertical axis and is yawed -2.5 deg, and AhmedML sets reference density to 1.
3. The null, stated plainly
We built an instrument to flag any published claim that exceeds its dataset's own label-noise ceiling, and ran it over every claim we had collected. The verdict is that it fires zero times. We had hoped for a positive.
| claim | dataset | eval n | published | ceiling used | verdict |
|---|---|---|---|---|---|
| DoMINO (arXiv:2507.10747) | DrivAerML | 48 | Spearman 0.99 | 0.9977 random-48 / 0.9882 extreme-48 | UNDER |
| FIGConvNet (arXiv:2507.10747) | DrivAerML | 48 | Spearman 0.99 | same | UNDER |
| X-MeshGraphNet (arXiv:2507.10747) | DrivAerML | 48 | Spearman 0.96 | same | UNDER |
| DoMINO / FIGConvNet / X-MeshGraphNet | DrivAerML | 48 | R2 0.98 / 0.97 / 0.92 | 0.9981 | UNDER |
| UniversalAGI SUV-PT | DrivAerML | UNKNOWN | Spearman 0.97 | 0.9989 full / 0.9977 at n = 48 | UNDER |
| PhysicsX LGM-Aero | DrivAerNet++ | UNKNOWN | Spearman 0.9350 (in-distribution) | 0.9475 worst case (N_eff = 1) | UNDER - closest |
| PhysicsX LGM-Aero | DrivAerNet++ | UNKNOWN | Spearman 0.6910 (OOD) | as above | UNDER |
| UniversalAGI SUV-Bench-Med / -Hard; PhysicsX on Luminary SHIFT-SUV | proprietary | UNKNOWN | 0.96 / 0.54 / 0.9732 | none exists | NOT ASSESSABLE |
The closest call is PhysicsX's in-distribution drag Spearman of 0.9350, which sits 1.25 points under 0.9475, the most pessimistic bar for that dataset. That bar in turn needs the shipped scatter column to be a standard error of the mean, which it is not (C5). On DrivAerML, DoMINO and FIGConvNet sit within 0.008 of the labels' own self-consistency, so optimising that metric further fits solver noise.
4. Our own corrections
C1. We called a published baseline "huge headroom"; it is a strawman
AirfRANS's own GraphSAGE reference reports a drag Spearman of rho = -0.303 ± 0.124 (arXiv:2212.07564, Table 5), so drag is anti-predicted. Raw XFOIL, a 1980s panel code costing a MEASURED 0.0378 s and $4.7e-7 per case, scores +0.8794; with one configuration change, moving Ncrit from 9 to 0.05, it scores +0.9949 at a MAPE of 2.3%, which is 1.30 rank-correlation units above the published baseline. We had written the strawman-baseline attack down ourselves, and then walked into it. The same change moves the median XFOIL-versus-RANS residual from +36.8 to 1.66 counts (2D convention, 1 count = 0.0001 Cd): the "physics gap" was the transition model.
C5. Our first noise ceiling misread the dataset's own column
We treated DrivAerNet++'s shipped scatter column as a standard error of the mean. arXiv:2406.09624, Appendix E.3.3, verbatim: "... forces averaged over the last 1000 iterations. We provide both the mean and standard deviation for uncertainty quantification." It is iteration scatter over a 1000-iteration SIMPLE window. A free check: the ratio of the rear-lift scatter to rear lift has a median of 0.436, and a 44% standard error on rear lift is impossible. Do not quote our earlier 0.9153 ceiling.
C6. Two statistical errors in our own calculation
First, we computed the expected Spearman between two repeat runs, which is test-retest reliability; the bar for a model-versus-label Spearman is the expected Spearman between truth and label, roughly its square root: at N_eff = 1 that is 0.9475, not 0.9153. Second, using the noisy label as truth double-counts noise into signal; empirical-Bayes shrinkage gives an R2 ceiling of 0.8986, not 0.9061.
C7. Three "missing label" counts were our own download bug, and it moved a headline by 31 points
Our first draft said these datasets ship fewer labelled runs than their papers claim, at DrivAerML 442/500, AhmedML 466/500 and WindsorML 329/355, and treated that as a defect biasing every published number. It was our defect. Our own download log records 336 + 243 + 141 exhausted-retry failures, which is our network and not HTTP 404, and the table builder wrote them as NaN, turning our failures into "missing labels". Re-probing all 403 locally absent files, 323 return HTTP 200 and were re-downloaded in 10.2 s. Corrected: DrivAerML 484 of 500, AhmedML 500 of 500, WindsorML 355 of 355 (MEASURED).
The only genuine gap is DrivAerML's 16 runs, whose design parameters are published, so clustering is testable: the mean pairwise distance among the 16 in standardised design space is 5.7739 against a 5,000-draw permutation null of 5.5971, giving p = 0.8516, not clustered, and nothing significant after Bonferroni. A caveat we do not bury: 3 of 16 parameters land at p < 0.05 where 0.8 are expected, and the probability of 3 or more under the null is 0.0429, which we do not claim at n = 16. So "the missing runs bias every published number" is NOT SUPPORTED, and two of our own numbers die instead: A6's AhmedML rho falls from 0.917 to 0.8148, and to 0.3029 under the leak-robust key, inverting our headroom ranking; and A4's DrivAerML probability that the best design is ranked first falls from 0.998 to 0.6893. The lesson: missing is MISSING, never 0 was obeyed, because the cells really were NaN, but nobody asked whether the missingness was ours. A download log is part of the dataset.
What survived an independent re-derivation: A1 reproduced from the authoritative aggregates, with DrivAerML at 24.85% flips, rho 0.6912 and top-5 agreement 0 of 5, and AhmedML at 30.34% and 0.5588, with an algebraic residual of 9.99e-08, tighter than the 3.99e-07 we first published — and it generalised into a better claim, that frontal-area spread governs the hazard. A6's DrivAerML kill held on the corrected 484 rows and under leak-robust grouping. A4's DrivAerNet++ half held.
5. Recommendations
| # | recommendation | closes |
|---|---|---|
| R1 | Pin the reference-area convention IN THE TASK DEFINITION and name the file | A1 |
| R2 | Ship the per-geometry force file as default; or rename the default | A6 |
| R3 | Ship a per-run error estimate WITH ITS DEFINITION (SEM vs window scatter) | C5+A4 |
| R4 | Ship per-case force histories (~100 KB/case) so tau_int/N_eff and a correlation length are measurable | A4+A8 |
| R5 | Publish the EVALUATION SET's Cd spread and implied ceiling beside every Spearman | A2 |
| R6 | Report drag counts and rank correlation SEPARATELY; state the count convention | A3 |
| R7 | Include a TRIVIAL-BASELINE row (linear regression; frontal area; body-style label) | A6 |
| R8 | Report raw AND offset-corrected scores on any OOD split | offset |
| R9 | Print the FULL force history in any grid study; run one design sweep on two meshes | A7 |
| R10 | Encode unscoreable as absent (never as a number) | Section 4 |
| R11 | Weight any field metric by the shipped variance; report the quiet/noisy ratio beside one rel-L2 | A8 |
| R12 | Publish every OOD split beside the AUC of the obvious confounder (drag) or a drag-matched control | A9 |
If only three are adopted, pick R1, R5 and R7. R9 is free.
6. Limitations
ESTIMATED. DrivAerNet++'s N_eff is 3-30, central 10, inferred from scatter of 2.4% of Cd under RMS residuals below 1e-5; every DrivAerNet++ ceiling depends on it. AhmedML's sigma of 1.3 counts, range 1.0-1.8, with 1 ct = 0.001 Cd, is read off the instantaneous-Cd trace in Figure 11(a) of arXiv:2407.20801, and AhmedML uses a fixed 80-CTU budget rather than an error target. DrivAerML's sigma of 0.75 counts is MEASURED but modal, not maximal.
UNKNOWN. WindsorML uncertainty: a full-text search of arXiv:2407.19320 for "drag count", "95%", "confidence", "standard deviation" and "Meancalc" returns zero hits, so every WindsorML row here is a sensitivity analysis. The design-dependent mesh term for the four road-car datasets is UNKNOWN (A7; measured only in 2D). For Luminary SHIFT-SUV and SHIFT-Wing, the listings are public but the bytes return HTTP 401 and we obtained zero bytes, so every Luminary row in Section 3 is NOT ASSESSABLE.
Settled since the first draft: the wall-shear-stress erratum does not reach scalar Cd. PhysicsX report two wall-shear-stress defects in DrivAerNet++ — a sign convention differing from common use, on all vehicles, and an incorrect symmetry-plane reflection on the estate and notchback half-car cases — and state verbatim that they "have not ablated the effect of these corrections". We had marked scalar Cd UNVERIFIED. It is REFUTED, on four lines, with 1 count = 0.001 Cd. Scope: the reflection happens "during post-processing", while DrivAerNet++ Appendix E.4 computes forces in-solver over the last 1000 iterations, so the two pipelines do not touch. Magnitude: skin friction is 10-15% of road-car drag, or 24-36 counts (ESTIMATED, external), so a sign-flipped viscous term would put Cd 48-72 counts under the tunnel, while its own validation table shows 7. Falsification: the reflection hits the estate and the notchback, so both must carry the same offset, but OLS inside the shared WWC wheel stratum gives beta_E = +24.28 ± 0.46 counts against beta_N = +1.09 ± 0.54, a 22.8-count difference at z = 32. And there is no noise signature: both indicted styles show lower residual standard deviation, and a lower shipped scatter column, than the style that was not indicted.
Why nobody could ablate it: the file named "Updated" is bit-identical to the Average Cd column, so there is nothing to diff against. Named residual UNKNOWN: integrating pressure times normal plus wall shear stress over the shipped boundary field, for one estate, one notchback and one fastback, would settle it — 39 TB, CC BY-NC, Globus-only in practice, and not obtainable at $0. Confidence HIGH, not certain.
Scope. Everything here is scalar drag from shipped design tables, about 1.8 MB standing in for about 69 TB, plus our own 2D mesh study. No claim is made about field prediction.
Sources and artifacts
Every number above resolves to a file and a line in the published source map (scan/AUDIT_SOURCE_MAP.csv).
| artifact | identifier or hash |
|---|---|
| DrivAerNet++ parametric design table | DrivAerNet_ParametricData.csv, 1.441 MB, 4,165 rows, sha256 9570826656321b9a7f468c8ef38e71709392b830917353aed97839170f26a7bd |
| DrivAerNet++ updated drag file | DrivAerNetPlusPlus_Cd_8k_Updated.csv, 273,749 B |
| Aggregate label and parameter files | force_mom_all.csv, force_mom_{constref,varref}_all.csv, geo_parameters_all.csv, and the per-run force_mom_<n>.csv |
| Dataset mirrors used | HuggingFace neashton/{drivaerml,ahmedml,windsorml}. The S3 route printed in the papers, s3://caemldatasets, is dead |
| A7 solver image | opencfd/openfoam-run:2312 |
| A7 per-case results | mesh_results.csv, sha256 72fea952a7c1188525b41918da7d7ce0c7854c784a551e8ef6bc5809369ce9f3 |
| C1 panel code | XFOIL 6.99, built from source |
| Licences | DrivAerML, AhmedML and WindsorML CC BY-SA 4.0; DrivAerNet++ CC BY-NC 4.0 (non-commercial); AirfRANS ODbL |
The scalar-drag instrument these rules are enforced on · all research