Research

What CFD-surrogate benchmarks actually measure: convention, spread, offset, and mesh

August 2026 · audit note on four public CFD-surrogate corpora · every number is tagged MEASURED, ESTIMATED or UNKNOWN and resolves to a file and a line in the published source map · four of our own numbers are wrong and are retracted in section 4, not edited away

For teams publishing on DrivAerML, DrivAerNet++, AhmedML and WindsorML, and vendors quoting Spearman against them.

Two drag-count conventions appear in this note and each is restated at every use. For road cars, 1 count = 0.001 Cd. For 2D airfoils, 1 count = 0.0001 Cd. Every number carries a tag: MEASURED, ESTIMATED (under a stated assumption), or UNKNOWN. Missing data is written as MISSING and never as 0, including in denominators. "Not reported" means we did not find it after a targeted search. These are easy mistakes, and several of them are naming choices rather than science. Four of our own numbers were wrong; they are retracted in Section 4.


1. The claim

Published numbers on these benchmarks are not comparable across papers. Three reasons are each sufficient on their own, and a fourth, numerical one is now measured in 2D.

  1. Convention. Each dataset ships two ground-truth force files that differ only in the reference area used to normalise drag, and switching files reorders the designs. On DrivAerML 24.85% of design pairs flip, and the two truths agree on 0 of the 5 best designs (MEASURED, 484 runs).
  2. Spread. A rank-correlation ceiling belongs to the evaluation set, not to the dataset. Same sigma, same labels, a different subset: the DrivAerML ceiling moves 0.9989 → 0.6370 (MEASURED, complete 484). No paper reports its evaluation set's Cd spread.
  3. Offset. The out-of-distribution collapse is about 86% a constant level error. An oracle mean shift restores R2 from 0.2078 to 0.745 (MEASURED).
  4. Mesh. Every ceiling here is statistical. The discretisation term is larger, and it is not purely common-mode (A7).

Two further findings are reported here for the first time. The corpora publish a 4x noise map that the metric everyone uses throws away (A8), and the newest official out-of-distribution split is drag-confounded (A9).

We built an instrument to flag claims that exceed a dataset's own label-noise ceiling. It fires zero times (Section 3). Four of our own conclusions were falsified; they are published, not edited away.


2. The findings

A1. The reference-area convention flips ranks; frontal-area spread sets how badly

Re-derived from each repository's authoritative aggregate label files, joined on run id. MEASURED, road-car convention, 1 count = 0.001 Cd.

datasetnpairspairs flipped (%)Spearman A-BKendalltop-5 agreementfrontal-area spread
DrivAerML48411688624.850.69120.50300 of 548.2% (1.779-2.636 m2 vs const 2.17)
AhmedML49912425130.340.55880.39321 of 5const 0.112 m2
WindsorML355628355.340.98280.89325 of 55.4% (0.1127-0.1188 m2 vs const 0.1120)

Top-5 agreement is agreement on the 5 lowest-drag designs, which is the engineering decision. The mean absolute difference between the two conventions is 17.6 counts on DrivAerML, with a maximum of 61.4, and the relation between them is algebraic: Cd_constref = Cd_pergeom x (aRef/aRefRef) holds to a maximum residual of 9.99e-08 over 484 runs. WindsorML gives the general law. Its frontal area varies by 5.4% and it flips 5.34% of pairs; DrivAerML's varies by 48.2% and it flips 24.85%. The hazard is set by frontal-area spread in the design space.

The naming inversion is the part that bites. The default per-run force file (force_mom_i.csv) uses per-geometry area in DrivAerML but constant area in AhmedML and WindsorML: same author, same layout, opposite meaning, confirmed verbatim in the three READMEs. Taking the default file everywhere is the natural thing to do and it is wrong for two of the three. Inside a single dataset, "Spearman over designs" is defined only to about ± 0.45 by a choice of filename.

A2. The Spearman ceiling belongs to the evaluation set, not the dataset

A label-noise ceiling is what a perfect surrogate scores against a noisy test label, so it depends on label noise relative to the spread of the scored subset. Method: y = t + e, with e drawn from N(0, sigma^2); t by empirical-Bayes shrinkage; the ceiling is E[Spearman(t, y_new)] over 600-1000 Monte-Carlo repetitions. DrivAerML's sigma of 0.75 counts is MEASURED from its stated 95% confidence interval of ± 1.5 counts. Road-car convention, 1 ct = 0.001 Cd.

datasetsubsetnsd of Cd (ct, 1ct=1e-3)sigma (ct)Spearman ceilingtag
DrivAerMLfull48417.560.750.9989MEASURED (complete)
DrivAerMLrandom 484816.430.750.9977MEASURED (complete)
DrivAerMLextreme 48 = 24 lowest + 24 highest4836.300.750.9882MEASURED (complete)
DrivAerMLtightest 48 window480.980.750.6370MEASURED (complete)
DrivAerMLall within a 1 ct band of the median120.250.75-0.0100MEASURED (complete)
DrivAerNet++full416522.022.010.9952ESTIMATED (N_eff = 10)
DrivAerNet++middle 200 band2000.810.660.4258ESTIMATED (N_eff = 100)

These are recomputed on the complete 484 runs after correction C7, with 4,000 Monte-Carlo repetitions; the earlier 442-row values were 0.9988, 0.9878 and 0.5720. The correction moves nothing in direction and the effect is slightly larger: 0.9989 across the whole drag range, 0.6370 over the tightest 48-design window, and -0.0100 once only the twelve designs inside a one-count band remain. Same labels, same solver, same sigma. The ceiling is a property of the subset you chose to score on.

A3. Correlation says the field is finished; absolute error says 5-13x

On DrivAerML, 1 drag count = 1.9756 N (MEASURED: U = 38.889 m/s, A = 2.17 m^2, paper-implied rho = 1.2041; the paper states T = 293.15 K and never prints rho). At sigma = 0.75 counts the label-noise MAE floor is sigma*sqrt(2/pi), or 0.598 counts. Published MAE and maximum AE are from NVIDIA arXiv:2507.10747; the conversion to counts is ours. Road-car convention, 1 ct = 0.001 Cd.

modelpublished MAE (N)MAE (ct, 1ct=1e-3)x above noise floorpublished max AE (N)max AE (in sigma)tag
X-MeshGraphNet15.237.7012.87x58.1639.2MEASURED (conversion)
FIGConvNet8.864.487.49x25.7217.3MEASURED (conversion)
DoMINO6.643.365.61x23.0815.6MEASURED (conversion)
our GBM on DrivAerNet++, grouped CVnot published9.9-14.43.8-7.7xnot publishednot publishedESTIMATED (N_eff = 100-10)

Rank correlation saturates while absolute drag, the metric that decides anything, sits 5-13x above the labels' own repeatability.

A4. Ranking the top designs is at the edge of what the labels support

Road-car convention, 1 count = 0.001 Cd.

datasetnmedian adjacent gap (ct)best-to-2nd gap (ct)smallest resolvable difference, 95% (ct)tag
DrivAerNet++41650.01441.075717.65 at N_eff = 1; 5.58 at N_eff = 10 (central); 1.77 at N_eff = 100MEASURED (gaps); ESTIMATED (resolution)
DrivAerML4840.09430.52042.08MEASURED (complete)
AhmedML4990.12803.73493.60MEASURED (complete); ESTIMATED (sigma)
WindsorML3550.18823.9624UNKNOWNMEASURED (complete); no published sigma

The probability that the true best design is ranked first is 0.5295 on DrivAerNet++ at the central N_eff of 10, falling to 0.1003 at N_eff = 1 and rising to 0.9513 at N_eff = 100 (MEASURED, 4,000 Monte-Carlo repetitions). In the dataset's favour: using the top two designs' own shipped scatter column (Std Cd), which reads 4.218 and 2.830 counts, the 95% least significant difference at N_eff = 100 is 1.00 count against a gap of 1.0757 counts.

DrivAerML here is a retraction (C7). On complete labels that same probability is 0.6893, not the 0.998 we published. The true second-best is run 159, one of the rows our own download had dropped, and the top-two gap collapses from 3.23 to 0.52 counts. AhmedML is the one dataset that separates its top two, and only just: 3.7349 counts against a least significant difference of 3.60 counts. Top-1 selection difficulty grows with dataset size, so any paper claiming to identify the optimal design should report the top-two gap against label noise on the complete label set.

A5-A6. The default file leaks frontal area, and one dataset is solved by linear regression

GBM and OLS on the shipped design parameters, predicting Cd, using the per-geometry-area file in every row and corrected label counts. Two keys were used: a random five-fold split, which is not leak-robust for a design of experiments, and a grouped five-fold split on KMeans clusters of the standardised parameters (k = 5). MEASURED.

datasetnparametersrho, random 5-foldrho, leak-robustour first draft
WindsorML35070.59440.43970.593
DrivAerNet++4165230.78860.5593 (body-style holdout)0.789
AhmedML49980.81480.30290.917 RETRACTED
DrivAerML484160.95030.94570.952

On DrivAerML, 16 numbers and ordinary least squares reach R2 0.8882, rho 0.9426 and MAE 4.65 counts (1 ct = 0.001 Cd), and this survives the leak-robust key at 0.8745 and 0.9353.

The headroom ranking inverts, and the error was ours (C7). We first put AhmedML at rho 0.917 and called it "nearly solved". That came from the leaky constant-area file: normalising by a constant area leaves the per-geometry frontal area inside the target. Spearman between area and Cd is +0.7755 on AhmedML with constant area, against +0.0059 with per-geometry area; the same pair on DrivAerML is +0.8230 against +0.1893, and DrivAerNet++ is clean at -0.0340. On the per-geometry file AhmedML scores 0.8148, and under the leak-robust key it collapses to rho 0.3029 and R2 -0.141, worse than predicting the mean. That makes AhmedML the hardest of the four, not the easiest. WindsorML's 0.5944 needs a seventh parameter, the front-to-rear length ratio (ratio_length_front_rear), which is absent from the published aggregate file; with the published six it scores 0.5362.

Trivial-rule bars, all MEASURED: the body-style label alone scores 0.5057 on DrivAerNet++; body height times body width scores 0.7642 on AhmedML with its default file; linear regression scores 0.9426 on DrivAerML.

A7. The two-mesh design sweep: we ran it in 2D. In 3D it is still open.

Every ceiling above is statistical, so it is a lower bound on label error. Published grid sensitivity is 20-30x larger: DrivAerNet++'s relative Cd error against the TUM reference runs 8.21% at 6M cells, 6.55% at 12M and 2.17% at 24M (arXiv:2406.09624, T7). Grid error is usually assumed to be common-mode, that is, a bias rather than a scrambling. We tested that in 2D and it is half true. The study covers 12 cases (NACA 4-digit, Re 2-6e6, alpha -2 to 12 deg) on 3 systematically refined 6-block C-grids of 32,256, 72,576 and 163,296 cells, with a refinement ratio of exactly r = 1.5000 and grading frozen across levels, solved with simpleFoam and k-omega SST wall-resolved (y+ max 1.2-2.6): 66 runs, about 46 CPU-hours. The 2D convention applies in this block: 1 count = 0.0001 Cd.

estimatorcasesmedian (ct, 2D: 1ct=1e-4)p90 (ct)
U_C assumption-free (1.25 x half-spread over 3 levels)62.203.83
U_B Roache fallback (2-grid; assumed p = 2; Fs = 3)72.936.25
U_A textbook 3-grid GCI with OBSERVED p214.5919.64

The observed order of convergence p was computable in only 2 of 12 cases, at 0.97 and at 0.19, the second flagged as unphysical. 4 cases are oscillatory and 6 never reached a usable iterative band in 3,000-3,500 iterations. Those are MISSING, never 0: the denominators are the cases column.

Two results follow. First, the label-noise floor the field quotes is the wrong quantity. Iterative convergence on this solver family measures 0.001 counts, while discretisation uncertainty is 2.20 counts assumption-free and 2.93 counts on the Roache fallback, about 2,200x larger. Quoting iterative convergence as a label-noise floor therefore understates numerical uncertainty by more than three orders of magnitude. Second, mesh error is not purely common-mode. The coarse-to-fine mean shift is -3.94 counts, but the standard deviation of the per-design shift is 1.55-2.81 counts, with one genuine rank swap: Spearman between levels is 1.0000 from L1 to L2 and then 0.9643, at n = 7. So the design-dependent term that bounds every ceiling here is real and, in 2D, larger than the 1.66-count effect we had proposed to model.

Threats this lane states against itself: maximum mesh non-orthogonality is 65.7 deg, so a smoother C-grid could show cleaner p and smaller uncertainty; and three medium-to-fine deltas, of 0.001-0.13 counts, sit below the study's own iterative band of 0.30-0.63 counts, so those may be underestimates. The defensible claim is the order of magnitude: a few drag counts, not a thousandth of one. For the four road-car datasets the design-dependent mesh term stays UNKNOWN.

The method finding may be worth more than the number. 18 solver configurations were probed. SIMPLEC at 0.9 relaxation, and even at 0.5/0.5, holds a plausible Cd plateau for about 1,200 iterations and then collapses into a ±50-count limit cycle (2D convention), at the coarse and the medium level, so it is the relaxation and not the grid. Only 0.4/0.4/0.4 with cellLimited gradients and one non-orthogonal corrector survived.

Any grid study that does not print the full drag history is not checkable.

A8. DrivAerML publishes its own noise map, and a whole-surface relative-L2 hides it

DrivAerML ships the pressure-variance field itself (pPrime2MeanTrim) on the surface, at 35.3 MB per case, and in the volume. We did not find it used as an error weight in any surrogate paper after a targeted search. Area-weighted Cp_rms on run_1, MEASURED:

corpuscasesbase/upper Cp_rms ratiospread
DrivAerML4843.093 (median)IQR 0.67; range 2.03-4.52; only 5.0% of cases exceed 3.93
WindsorML3491.74 (median)not stated
run_1 alone, the case first examined13.93a p95 outlier, not the typical case

A single whole-surface relative-L2 hides a 3.1x heteroscedasticity, ranging 2.0-4.5x across the corpus, that the dataset itself publishes. This was corrected before publication: an earlier draft quoted 3.93x from a single case, and measuring all 484 showed that case is a p95 outlier while the corpus median is 3.093. AhmedML ships no pressure variance at all.

Two limitations we state rather than bury. This is a lower bound on the ratio of standard errors of the mean, because the integral time scale is longer in the wake. And a variance map without a correlation length under-reports field-integral uncertainty by more than 6x: the independent-face bound gives 0.25 counts (1 count = 0.001 Cd) against the dataset's own ±1.5-count tolerance, so that tolerance is set by coherent wake unsteadiness. The artifact that would close this is a per-case force-coefficient history (forceCoeffs) of about 100 KB per case. There is none, and the convergence history ships only as a PNG. 3D discretisation error is likewise unmeasurable from the shipped data, because there is one mesh per case.

A9. The new official splits are good, and one of them is drag-confounded

DrivAerML published 8 deterministic splits on 2026-08-17, ungated, under CC BY-SA 4.0, and they hold. Counts match the README exactly: 400/34/50 for the full split, and 339/48/97 each for the geometry, high-drag, low-drag and rear-separation splits. Train/test, train/validation and validation/test overlap is zero in all 8 families, the nesting from super-scarce through scarce and medium to full is true, and the published test set reproduces exactly from the committed score. The splits cover 484 cases, not 500, independently confirming C7. The rear-separation split is genuinely physics-defined, taken from wake area in images rather than from an integrated coefficient.

But it is drag-confounded, and the README does not say so. Spearman between score and Cd is -0.5594, the AUC of Cd for test against train is 0.2619, and 41 of its 97 test cases are also in the low-drag test set.

A model that merely under-predicts drag will look as though it fails on separated wakes.

Use the geometry split, whose AUC is 0.538 and which is drag-neutral, as the primary out-of-distribution split; or publish the rear-separation split beside a drag-matched control. This is an easy mistake on splits published days ago.

And the fix does not transfer, which we only learned by measuring all three corpora. Re-measuring the AUC of Cd for test against train across all 24 split families: DrivAerML's geometry split reproduces at 0.5381 and is drag-neutral, but AhmedML's is 0.7164, with 41 of its 100 test cases also in the high-drag test set, and WindsorML's is 0.6046. The image-wake split is worse still, at 0.7713 and 0.6529. The geometry split is drag-neutral only on DrivAerML. A benchmark that adopts it as a universal primary split inherits a drag confound on two corpora out of three.

Separately, on all three corpora the super-scarce training set is drag-shifted high by 19 to 37 counts against its own test set, so the data-efficiency ladder confounds less data with different data. We did not find this reported.

Two further errata were found while extracting the surfaces. WindsorML runs 350-354 have no directory on the host, and run_354 appears in six test folds, so every published WindsorML test denominator is one too large (full test 36 to 35; high-drag test and low-drag test 71 to 70). And the per-run force files are variable-reference on DrivAerML but constant-reference on AhmedML, at 0.112032, and on WindsorML, at 0.112 — the same naming inversion as A1, now confirmed on the per-run files as well as the aggregates.

Supporting corrections we also owe the field

All MEASURED, from primary sources or first-hand fetches.


3. The null, stated plainly

We built an instrument to flag any published claim that exceeds its dataset's own label-noise ceiling, and ran it over every claim we had collected. The verdict is that it fires zero times. We had hoped for a positive.

claimdataseteval npublishedceiling usedverdict
DoMINO (arXiv:2507.10747)DrivAerML48Spearman 0.990.9977 random-48 / 0.9882 extreme-48UNDER
FIGConvNet (arXiv:2507.10747)DrivAerML48Spearman 0.99sameUNDER
X-MeshGraphNet (arXiv:2507.10747)DrivAerML48Spearman 0.96sameUNDER
DoMINO / FIGConvNet / X-MeshGraphNetDrivAerML48R2 0.98 / 0.97 / 0.920.9981UNDER
UniversalAGI SUV-PTDrivAerMLUNKNOWNSpearman 0.970.9989 full / 0.9977 at n = 48UNDER
PhysicsX LGM-AeroDrivAerNet++UNKNOWNSpearman 0.9350 (in-distribution)0.9475 worst case (N_eff = 1)UNDER - closest
PhysicsX LGM-AeroDrivAerNet++UNKNOWNSpearman 0.6910 (OOD)as aboveUNDER
UniversalAGI SUV-Bench-Med / -Hard; PhysicsX on Luminary SHIFT-SUVproprietaryUNKNOWN0.96 / 0.54 / 0.9732none existsNOT ASSESSABLE

The closest call is PhysicsX's in-distribution drag Spearman of 0.9350, which sits 1.25 points under 0.9475, the most pessimistic bar for that dataset. That bar in turn needs the shipped scatter column to be a standard error of the mean, which it is not (C5). On DrivAerML, DoMINO and FIGConvNet sit within 0.008 of the labels' own self-consistency, so optimising that metric further fits solver noise.


4. Our own corrections

C1. We called a published baseline "huge headroom"; it is a strawman

AirfRANS's own GraphSAGE reference reports a drag Spearman of rho = -0.303 ± 0.124 (arXiv:2212.07564, Table 5), so drag is anti-predicted. Raw XFOIL, a 1980s panel code costing a MEASURED 0.0378 s and $4.7e-7 per case, scores +0.8794; with one configuration change, moving Ncrit from 9 to 0.05, it scores +0.9949 at a MAPE of 2.3%, which is 1.30 rank-correlation units above the published baseline. We had written the strawman-baseline attack down ourselves, and then walked into it. The same change moves the median XFOIL-versus-RANS residual from +36.8 to 1.66 counts (2D convention, 1 count = 0.0001 Cd): the "physics gap" was the transition model.

C5. Our first noise ceiling misread the dataset's own column

We treated DrivAerNet++'s shipped scatter column as a standard error of the mean. arXiv:2406.09624, Appendix E.3.3, verbatim: "... forces averaged over the last 1000 iterations. We provide both the mean and standard deviation for uncertainty quantification." It is iteration scatter over a 1000-iteration SIMPLE window. A free check: the ratio of the rear-lift scatter to rear lift has a median of 0.436, and a 44% standard error on rear lift is impossible. Do not quote our earlier 0.9153 ceiling.

C6. Two statistical errors in our own calculation

First, we computed the expected Spearman between two repeat runs, which is test-retest reliability; the bar for a model-versus-label Spearman is the expected Spearman between truth and label, roughly its square root: at N_eff = 1 that is 0.9475, not 0.9153. Second, using the noisy label as truth double-counts noise into signal; empirical-Bayes shrinkage gives an R2 ceiling of 0.8986, not 0.9061.

C7. Three "missing label" counts were our own download bug, and it moved a headline by 31 points

Our first draft said these datasets ship fewer labelled runs than their papers claim, at DrivAerML 442/500, AhmedML 466/500 and WindsorML 329/355, and treated that as a defect biasing every published number. It was our defect. Our own download log records 336 + 243 + 141 exhausted-retry failures, which is our network and not HTTP 404, and the table builder wrote them as NaN, turning our failures into "missing labels". Re-probing all 403 locally absent files, 323 return HTTP 200 and were re-downloaded in 10.2 s. Corrected: DrivAerML 484 of 500, AhmedML 500 of 500, WindsorML 355 of 355 (MEASURED).

The only genuine gap is DrivAerML's 16 runs, whose design parameters are published, so clustering is testable: the mean pairwise distance among the 16 in standardised design space is 5.7739 against a 5,000-draw permutation null of 5.5971, giving p = 0.8516, not clustered, and nothing significant after Bonferroni. A caveat we do not bury: 3 of 16 parameters land at p < 0.05 where 0.8 are expected, and the probability of 3 or more under the null is 0.0429, which we do not claim at n = 16. So "the missing runs bias every published number" is NOT SUPPORTED, and two of our own numbers die instead: A6's AhmedML rho falls from 0.917 to 0.8148, and to 0.3029 under the leak-robust key, inverting our headroom ranking; and A4's DrivAerML probability that the best design is ranked first falls from 0.998 to 0.6893. The lesson: missing is MISSING, never 0 was obeyed, because the cells really were NaN, but nobody asked whether the missingness was ours. A download log is part of the dataset.

What survived an independent re-derivation: A1 reproduced from the authoritative aggregates, with DrivAerML at 24.85% flips, rho 0.6912 and top-5 agreement 0 of 5, and AhmedML at 30.34% and 0.5588, with an algebraic residual of 9.99e-08, tighter than the 3.99e-07 we first published — and it generalised into a better claim, that frontal-area spread governs the hazard. A6's DrivAerML kill held on the corrected 484 rows and under leak-robust grouping. A4's DrivAerNet++ half held.


5. Recommendations

#recommendationcloses
R1Pin the reference-area convention IN THE TASK DEFINITION and name the fileA1
R2Ship the per-geometry force file as default; or rename the defaultA6
R3Ship a per-run error estimate WITH ITS DEFINITION (SEM vs window scatter)C5+A4
R4Ship per-case force histories (~100 KB/case) so tau_int/N_eff and a correlation length are measurableA4+A8
R5Publish the EVALUATION SET's Cd spread and implied ceiling beside every SpearmanA2
R6Report drag counts and rank correlation SEPARATELY; state the count conventionA3
R7Include a TRIVIAL-BASELINE row (linear regression; frontal area; body-style label)A6
R8Report raw AND offset-corrected scores on any OOD splitoffset
R9Print the FULL force history in any grid study; run one design sweep on two meshesA7
R10Encode unscoreable as absent (never as a number)Section 4
R11Weight any field metric by the shipped variance; report the quiet/noisy ratio beside one rel-L2A8
R12Publish every OOD split beside the AUC of the obvious confounder (drag) or a drag-matched controlA9

If only three are adopted, pick R1, R5 and R7. R9 is free.


6. Limitations

ESTIMATED. DrivAerNet++'s N_eff is 3-30, central 10, inferred from scatter of 2.4% of Cd under RMS residuals below 1e-5; every DrivAerNet++ ceiling depends on it. AhmedML's sigma of 1.3 counts, range 1.0-1.8, with 1 ct = 0.001 Cd, is read off the instantaneous-Cd trace in Figure 11(a) of arXiv:2407.20801, and AhmedML uses a fixed 80-CTU budget rather than an error target. DrivAerML's sigma of 0.75 counts is MEASURED but modal, not maximal.

UNKNOWN. WindsorML uncertainty: a full-text search of arXiv:2407.19320 for "drag count", "95%", "confidence", "standard deviation" and "Meancalc" returns zero hits, so every WindsorML row here is a sensitivity analysis. The design-dependent mesh term for the four road-car datasets is UNKNOWN (A7; measured only in 2D). For Luminary SHIFT-SUV and SHIFT-Wing, the listings are public but the bytes return HTTP 401 and we obtained zero bytes, so every Luminary row in Section 3 is NOT ASSESSABLE.

Settled since the first draft: the wall-shear-stress erratum does not reach scalar Cd. PhysicsX report two wall-shear-stress defects in DrivAerNet++ — a sign convention differing from common use, on all vehicles, and an incorrect symmetry-plane reflection on the estate and notchback half-car cases — and state verbatim that they "have not ablated the effect of these corrections". We had marked scalar Cd UNVERIFIED. It is REFUTED, on four lines, with 1 count = 0.001 Cd. Scope: the reflection happens "during post-processing", while DrivAerNet++ Appendix E.4 computes forces in-solver over the last 1000 iterations, so the two pipelines do not touch. Magnitude: skin friction is 10-15% of road-car drag, or 24-36 counts (ESTIMATED, external), so a sign-flipped viscous term would put Cd 48-72 counts under the tunnel, while its own validation table shows 7. Falsification: the reflection hits the estate and the notchback, so both must carry the same offset, but OLS inside the shared WWC wheel stratum gives beta_E = +24.28 ± 0.46 counts against beta_N = +1.09 ± 0.54, a 22.8-count difference at z = 32. And there is no noise signature: both indicted styles show lower residual standard deviation, and a lower shipped scatter column, than the style that was not indicted.

Why nobody could ablate it: the file named "Updated" is bit-identical to the Average Cd column, so there is nothing to diff against. Named residual UNKNOWN: integrating pressure times normal plus wall shear stress over the shipped boundary field, for one estate, one notchback and one fastback, would settle it — 39 TB, CC BY-NC, Globus-only in practice, and not obtainable at $0. Confidence HIGH, not certain.

Scope. Everything here is scalar drag from shipped design tables, about 1.8 MB standing in for about 69 TB, plus our own 2D mesh study. No claim is made about field prediction.


Sources and artifacts

Every number above resolves to a file and a line in the published source map (scan/AUDIT_SOURCE_MAP.csv).

artifactidentifier or hash
DrivAerNet++ parametric design tableDrivAerNet_ParametricData.csv, 1.441 MB, 4,165 rows, sha256 9570826656321b9a7f468c8ef38e71709392b830917353aed97839170f26a7bd
DrivAerNet++ updated drag fileDrivAerNetPlusPlus_Cd_8k_Updated.csv, 273,749 B
Aggregate label and parameter filesforce_mom_all.csv, force_mom_{constref,varref}_all.csv, geo_parameters_all.csv, and the per-run force_mom_<n>.csv
Dataset mirrors usedHuggingFace neashton/{drivaerml,ahmedml,windsorml}. The S3 route printed in the papers, s3://caemldatasets, is dead
A7 solver imageopencfd/openfoam-run:2312
A7 per-case resultsmesh_results.csv, sha256 72fea952a7c1188525b41918da7d7ce0c7854c784a551e8ef6bc5809369ce9f3
C1 panel codeXFOIL 6.99, built from source
LicencesDrivAerML, AhmedML and WindsorML CC BY-SA 4.0; DrivAerNet++ CC BY-NC 4.0 (non-commercial); AirfRANS ODbL

The scalar-drag instrument these rules are enforced on · all research