Field-scale crop–soil models as pure JAX functions: an independent implementation from published equations, validated against DSSAT-CSM (BSD-3) and RZWQM2 outputs. The figure on the right follows your reading: first a matrix of seasons (rows) and days (columns), then measured results.
The ratios compare two programs that do different work: DSSAT-CSM writes complete daily files and runs the whole crop–soil model, Agri-JAX runs the validated day and returns ten daily series. On a whole 192-core node the first run is slower than DSSAT-CSM (0.85×); Act 3 has every row.
Scroll down and the figure changes as you read, or step through the frames with the buttons below it. Hover, focus or click an underlined number to see its conditions.
A hundred thousand seasons: DSSAT-CSM runs them one after another, Agri-JAX advances them all one day at a time
- 1.1DSSAT-CSM on one core runs the 10⁵ seasons one after another: 1426 s measured, with 30.3 GB of daily text written. The screen plays it sped up.
- 1.2On the same core Agri-JAX advances every season by one day per step. The first run, compilation included, takes 99.1 s, a warm start 56.0 s, a repeated call 48.9 s. The ten daily series stay in memory (1.5 GB); no file is written. That is 14× · 25× · 29× faster than DSSAT-CSM here (first run · warm start · repeated call).
- 1.3Why the order can change: each cell depends only on the cell to its left in the same row. A season’s today depends on its own yesterday.
- 1.4There is no arrow between rows: the water of one season never enters the equations of another.
- 1.5So the two loops swap. The outer loop runs over days and the inner loop applies the same operation to every row: one column is one instruction over a hundred thousand numbers, the shape a vector unit or a GPU is built for.
- 1.6One column is one DSSAT day: soil albedo and bucket rates, potential evapotranspiration, soil evaporation and transpiration, root water uptake, the bucket update, CERES-Maize growth, and the water ledger. Every module of the day is simulated; only what the modules that are not ported would produce (residue records, daily soil-property changes) is replayed from DSSAT.
The same matrix at the same scale; only the traversal order differs. The two sides do different work: DSSAT-CSM runs the whole crop–soil model and writes 303 kB of daily text per season, Agri-JAX runs the modules of the validated day and returns ten daily series, so the ratios are not the speed-up of one numerical kernel. Measured on one pinned core of an AMD EPYC 9654 (rorqual, 2026-09-28), one measurement per row. Cell colours are illustrative.
Existing Fortran implementations do not natively support array-style batch execution: global state and per-run file I/O bind each simulation to a single process. This is a property of the implementation, not of the language.
Sixty-five whole seasons, run free, agree with dscsm048: yield, dates, daily curves and the water ledger
- 2.1Sixty-five seasons (58 treatments of the DSSAT maize examples and 7 at the AmeriFlux corn site CA-TPA; 11,138 days) run free: soil water, evaporation and root uptake are computed by Agri-JAX, not read from DSSAT. Of the 59 seasons with grain, the worst yield differs by 1.5 % and the median by 4.0e-5; the acceptance limit is 2 %. Emergence, silking and maturity dates are equal in 65 of 65 runs. Daily LAI differs by at most 0.0061 (RMSE), profile soil water by 0.33 mm, and the daily water ledger closes to 6.6e-14 mm.
- 2.2One season day by day, UFGA8201 treatment 4 (irrigated, high nitrogen; the treatment calibrated in Act 4): LAI, biomass, grain mass and profile soil water follow the curves of DSSAT. The largest differences, LAI 0.0048, biomass 0.54 kg/ha, soil water 0.5 mm, are within the rounding of DSSAT’s own printed output.
Nitrogen off, float64 on CPU, reference dscsm048 build 486 (DSSAT-CSM 4.8.6.0). Only what no module of the day produces is replayed from DSSAT: residue records of the organic-matter module and daily soil-property changes. The example treatments agree with DSSAT at its print resolution; the CA-TPA seasons are 5 to 18 times above it in daily evapotranspiration and soil water.
Each module was also checked alone. One day at a time on DSSAT’s own inputs, XTRACT root uptake is within 1.03, potential evapotranspiration within 4.7 and soil albedo within 4.5 units of REAL*4 rounding. CERES-Maize driven by DSSAT’s soil water and transpiration: daily LAI within 0.50 %, biomass 0.63 %, yield 0.030 %, stages equal every day.
Same task, same timing boundary: clear gains on repeated calls, a mixed picture for one-off runs
- 3.1Simulating 10⁴ and 10⁵ seasons on one timing boundary. At 10⁵ seasons on one core DSSAT-CSM takes 1426 s; Agri-JAX takes 99.1 s for a first run, 56.0 s for a warm start and 48.9 s for a repeated call. On 32 cores the ratios are 2.3×, 4.1×, 22×. On a whole 192-core node DSSAT-CSM takes 17.4 s: the first run of Agri-JAX is slower (0.85×), the warm start 1.6× and the repeated call 25× faster.
- 3.2GPU: find the bottleneck, change the layout, check the model. A profile of one H100 showed about 490 short kernels per simulated day and the device idle 34–45 % of the time. Unrolling the layer loops, an execution setting chosen per backend with the process code unchanged, leaves 60–63 kernels and 2–3 % idle, and a repeated call 3.1× faster (1.31 → 0.42 s at 10⁵ seasons). The price is a longer compile (19 → 67 s). Checked: GPU against CPU, stage dates equal in 65 of 65 and yield within 1.4e-15; unrolled against looped on the GPU, yield, soil water, runoff and drainage bit for bit.
- 3.3Agri-JAX pays off with many seasons on few cores, and most of all with many calls at one array shape (calibration, sensitivity analysis, ensembles): a repeated call is 22–29× faster than DSSAT-CSM on every CPU setting measured. For up to 10³ seasons run once, plain DSSAT-CSM is all you need. The table gives the tool and the setting for each job.
Task: the 65 validation seasons replicated to B seasons (171 days on average), float64, ten daily series returned to the host. Agri-JAX perturbs 7 parameters by U(0.9, 1.1) per sample; the DSSAT replicas are unperturbed. Hardware: rorqual (Calcul Québec), 2 × AMD EPYC 9654 (192 cores), one whole NVIDIA H100 (Xeon Gold 6448Y host), JAX 0.10.2. “One core” is one pinned core, “32 cores” a 32-core allocation of such a node, not a workstation. One measurement per cell; repeated-call times vary by up to 20 % between processes on the node. CPU rows 2026-09-28, H100 rows 2026-09-29.
The two sides do different work: DSSAT-CSM writes 30.3 GB of daily text at 10⁵ seasons and runs the whole crop–soil model, Agri-JAX returns 1.5 GB of ten series. Against summary-only DSSAT-CSM the 1-core ratios roughly halve (repeated 15× instead of 29×); the node ratios do not move. DSSAT-CSM was not run on the H100 host, so the GPU has no ratio. Not measured, so no ratio is given: 10⁶ or more seasons, T4 or Colab GPUs, laptops.
One call fits a cultivar to observations and writes a row that DSSAT reads back
- 4.1
calibrate(...)reads the treatments and observations of a DSSAT maize experiment, fits the six cultivar coefficients by batched joint CMA-ES, picks the best candidate by its value as written to the.CULrow (five characters per number, so rounding can move a date by a day), writes the row and runsdscsm048with it. Example, UFGA8201: the objective falls from 1.015 to 0.0100, the held-out treatment from 1.046 to 0.0239, and DSSAT run with the written row agrees with Agri-JAX in yield within 2.9e-4. Over 22 of 22 round trips, dates are equal and yield is within 0.1 %. - 4.2Identifiability. From dates, yield, biomass and LAI alone, G2 and G3 lie along a curved valley of the loss: 24 recalibrations spread over G2 355–982 and G3 7.6–16.4. Adding grain number closes the valley around the published values (G2 831–984, G3 7.5–9.4) and cuts the held-out grain-number error (best fit of each seed): −51…−49 % → −7…−6 %. The price is the rainfed treatment 2, whose yield error grows from −5 % to −41 %: the model under-predicts kernel number under water stress.
- 4.3Why joint CMA-ES is the default. Recovering known cultivars from noise-free synthetic observations (264 fits), 225 fits reached the target, against 212 for a staged scheme that uses gradients, in 135 s against 210 s. The staged scheme needs fewer model calls per fit (443 forward and 51 gradient calls, against 1008 forward).
- 4.4Gradients, by scope: the forward model is validated; gradients of some coefficients are usable; gradients through phenology events are experimental. One gradient of the summed grain mass with respect to five coefficients takes 1.81 s for 10⁴ seasons on one H100 with the layer loops unrolled (5.02 s with loops), at the price of 316 s of compile against 64 s.
No calibration speed-up over DSSAT-CSM is claimed. For the synthetic task above, the measured time of joint CMA-ES on the 192-core node lies between 1.6× slower and 1.2× faster than an estimate for an ideally parallel DSSAT-CSM on the same node; the estimate is built from measured per-season rates, not from a calibration run in DSSAT-CSM. The loss contours are conditional slices (other coefficients fixed at the best values of a fit to the real observations), not posteriors.
calibrate works on the DSSAT v4.8.6 maize example treatments only, with nitrogen off; the tables that the free-run day reads for them are not distributed. Treatments where nitrogen changes the DSSAT yield by more than 5 % are refused.
A process is a pure function with a declared signature; swap one and the same checks run
- 5.1A process is a pure function that declares what it reads and writes. Three checks bind the author, not the numerical kernels: an AST lint at code time (AJ001–AJ012, AJ020–AJ021), a runtime check of the declared writes (
AGRI_JAX_CHECK=1), and a conformance kit at registration. Every coefficient carries its name, unit, meaning and provenance. - 5.2Rule one, pure functions: everything read is an argument and everything changed is returned. A cell written by two seasons at once would be a collision, so the state holds one row per season.
- 5.3Rule two, branch with where: on the same day the seasons are at different growth stages, so both branches are computed and each row picks its own. Both must stay finite, or the gradients turn to NaN.
- 5.4Rule three, no loops over layers, days or samples in the process code. A recurrence down the soil profile is written once; the runtime decides how it runs, as a loop on the CPU and unrolled on the GPU (Act 3).
- 5.5One line swaps the soil evaporation of the DSSAT day from Ritchie to SALUS. The replacement is checked against the declarations of the entry, and against a conformance gate, before it runs.
- 5.6What the swap changes, over 65 runs each simulated with its own method and with the other: median season soil evaporation +26 % and −14 %, median grain yield −0.07 % and −0.12 %, within −27.7 % … +16.1 %. The water ledger closes to ≤ 5.9e-14 mm in every swapped run. Each method reproduces
dscsm048when run as itself; the swapped run is a different model.
The table under “Measured” has the numbers of the swap. examples/swap_soil_evaporation.py runs the swap on a synthetic season without data; docs/swapping_a_process.md and docs/tutorial_new_process.md show the checks and a process written from scratch.