9. Reproducibility and regression testing¶
How every result of ASYNCH is checked against the reference results, and how to repeat all the checks yourself with one command.
23
C unit tests
73
Python tests
66.5 %
of the C lines run by the tests
0
reference files ever modified
Important
Rule for every change to the C code: run make check and the regression harness (with --compare-to the
previous build) before and after the change. If a result moves by more than the tolerance, the change is
scientific and must be explained in the CHANGELOG.md.
The benchmark is the set of example results shipped with the original ASYNCH repository
(examples/results/ and examples/more/*/results_benchmark/). These files are never modified or regenerated. Every
result is measured against them.
9.0 All the tests: make check¶
Run in the build folder, make check runs three sets of tests, in about a minute:
Test |
File |
What it checks |
|---|---|---|
|
|
23 C unit tests: the Runge-Kutta tables satisfy the order conditions (sum of b = 1, rows of A add up to c, …) and reach their order on y’ = y; their dense output is consistent; every built-in model has consistent sizes and all the functions the solver calls, and its equations give finite values and set every derivative; sorting and the id lookup; argument checks of |
|
|
73 tests of the Python package: runs identical to the |
|
|
every example against the reference results (9.1) |
The outcome is at the end (# PASS: 3, # FAIL: 0); the details are in tests/*.log of the build folder. If Python
or NumPy is missing, the last two are reported as SKIP. The unit tests found three bugs (B-22 to B-24, chapter 8)
when they were first written.
How much of the code the tests run (coverage)¶
Measured with tests/coverage_report.py (it needs only gcov, part of GCC):
mkdir build-cov && cd build-cov
../configure CFLAGS="-O0 -g --coverage" LDFLAGS="--coverage"
make -j4 && make check
python3 ../tests/coverage_report.py .
On 2026-09-26, make check runs 66.5 % of the lines of the C library and program (9 019 of 13 553): the solver
steps and Runge-Kutta tables 98-100 %, the C interface for other languages 89 %, the model setups 91 % and model
equations 73 %, the time loop 81 %. What stays untested needs inputs that the repository has no example for: PostgreSQL
databases (db.c, database forcings and outputs), grid-cell rain, reservoirs (steppers/forced.c), part of the dam
code, and some command-line options.
9.1 The harness¶
tests/regression/run_examples.py runs every example shipped in examples/ and
compares the produced files with those original reference (“benchmark”) files.
# after building (see 01_setup.md)
python3 tests/regression/run_examples.py # 1 MPI process
python3 tests/regression/run_examples.py --np 4 # 4 MPI processes
python3 tests/regression/run_examples.py --only model_258 # a single case
python3 tests/regression/run_examples.py --asynch /path/to/other/asynch --keep
python3 tests/regression/run_examples.py --solver 4 --tol-factor 0.01 # every example with the stiff solver
--solver and --tol-factor rewrite the global files in the temporary directory: numerical solver index, and error
tolerances multiplied by a factor. The two 2015 configurations are also run with the stiff solver (index 4, tolerances
× 0.01) as their own cases, “Rosenbrock solver”. They are compared with the 2015 references at 5·10⁻⁴: those
references were computed by Dormand–Prince, which records a peak only at the end of a step, and they differ from a run
at tolerance 10⁻⁸ by up to 4.3·10⁻⁴ (peaks, nearly all too low) and 4.2·10⁻⁴ (baseflow at the outlet), while the stiff
solver stays within 7·10⁻⁵ of that run.
The examples are copied to a temporary directory, so the repository is never modified.
--keepkeeps that directory so you can look at the outputs.Only the Python standard library is required. If
h5pyis installed (pip install h5py), the HDF5 snapshot files are compared too.The exit status is
0when everything passes, so the harness can run in CI.
Cases¶
Case |
Model |
Compared files |
Reference |
|---|---|---|---|
|
190 (constant runoff) |
|
|
|
190, tolerances from |
|
|
|
190 |
hydrographs |
|
|
254 (top layer) |
|
|
|
254 |
|
|
|
192 |
hydrograph |
|
|
196 |
idem |
idem |
|
258 |
idem |
idem, known mismatch since 1.6.0: the benchmark was produced with the evaporation error B-28 and the baseflow error B-30 |
|
259 |
idem |
idem, known mismatch (R-03, B-28, B-30) |
A known mismatch (XFAIL) is reported but does not make the run fail. A crash is
always a failure, even for those cases.
The 2015 configurations¶
The reference files in examples/results/ (.dat, .pea, .rec) were all produced in May 2015
(commit b73fc2d) with global files that are no longer in the repository: Global190.gbl (300
minutes) and Global254.gbl (6000 minutes), both starting on 2014-05-01. The input files
(topology, parameters, rain, evaporation) are byte-for-byte the same today.
examples/test_2015.gbl and examples/clearcreek_2015.gbl are those two configurations written in
today’s format; they write into examples/out_2015/. Run them like any example:
cd examples
mpirun -n 2 ../build/src/asynch test_2015.gbl # compare out_2015/test.* with results/test.*
Results: the test references are reproduced (hydrographs within 5e-7). The clearcreek references
are reproduced within the solver tolerance only when one line of model 254 is restored to its
2015 form; see issue R-02 in 08_known_issues.md.
Comparing with the original code (--compare-to)¶
Stored references can be old or incomplete, so the most direct test of a change is: does the changed code give the same answer as the code before the change, on the same inputs?
# 1. Build the unmodified original code once (default: commit 84da43a, into ~/asynch-original)
tests/regression/build_original.sh
# 2. Build your version with the SAME flags (see 01_setup.md), then:
python3 tests/regression/run_examples.py --compare-to ~/asynch-original/build/src/asynch
python3 tests/regression/run_examples.py --compare-to ~/asynch-original/build/src/asynch --np 4
Both executables run every example on identical copies of the inputs, and every output
file is compared: peak flows, hydrographs (.csv, .dat, .h5) and every snapshot
(.h5, .rec). A case line looks like
vs original: 27 of 27 output files identical
“Identical” means bit for bit. A file that differs but stays within the tolerance is
counted as “within tolerance” (use --verbose to list it). A file outside the tolerance,
or a file the original wrote but the new code did not, makes the case FAIL, even for
XFAIL cases. If the original itself crashes, its missing files are not compared.
What to expect:
processes |
same code, two runs |
what a change must achieve |
|---|---|---|
|
bit-identical |
bit-identical, unless the change is meant to alter results |
|
differ: ~1e-6 relative on 1-day runs; up to ~1e-4 absolute on the 6000-minute clearcreek run (measured: three 4-process runs of the original code differed by 8.4e-5, 9.7e-5 and 1.15e-4) |
within tolerance |
So a pure bug fix or refactoring must give “N of N output files identical” with one process. That is the strict test. Runs with several processes check that nothing breaks in parallel.
9.2 Why compare with a tolerance?¶
ASYNCH is an adaptive solver: every link chooses its own time step from an error
estimate. Anything that changes the last bits of a floating-point number (another
compiler, -O2 vs -O3, another CPU, another number of MPI processes) changes some
step sizes, and therefore the results, around the 6th significant digit.
The solver itself only guarantees accuracy up to its tolerances, which are given in
the .gbl file (typically 1e-3…1e-6 absolute and 1e-6 relative for discharge). So
differences below those tolerances carry no information. The harness accepts
and always prints the largest absolute and relative difference. A real regression (a
changed equation, a wrong unit, a parameter off by one index) produces differences of
percent or more, orders of magnitude above this threshold. The larger atol with several
processes covers the run-to-run variation measured above (up to ~1e-4). Both can be set with
--atol and --rtol.
Time of the peak. In .pea files, the time of each peak is compared with its own tolerance,
20 minutes by default (--peak-time-atol). On a flat-topped hydrograph the minute of the maximum is
ill-conditioned: two runs of the same code with 2 processes put some clearcreek peaks 15 minutes
apart while their values agree to 1e-7 m³/s. The peak value is always checked with the normal tolerance.
Files are compared by link id (.pea, .h5) or by row (.csv, .dat, .rec), so a different
order of links in the file (which happens with MPI) does not matter for .pea and .h5.
9.3 Reference baseline (commit 84da43a + example path fix)¶
Release build (-O3 -DNDEBUG), Ubuntu 24.04, GCC 13, OpenMPI 4.1, HDF5 1.10:
Case |
np=1 |
np=2 |
np=4 |
|---|---|---|---|
test (190) |
PASS |
PASS |
PASS |
clearcreek (254) |
FAIL crash (B-01) |
XFAIL R-02 |
XFAIL R-02 |
model_192 |
PASS |
PASS |
PASS |
model_196 |
PASS (bit-identical) |
PASS |
PASS |
model_258 |
PASS (bit-identical) |
PASS |
PASS |
model_259 |
XFAIL R-03 |
XFAIL R-03 |
XFAIL R-03 |
At that commit, a debug build (no -DNDEBUG) additionally failed every case with exit
code 134, because of the double fclose at shutdown (B-03, since fixed).
9.4 Useful tools¶
AddressSanitizer / UndefinedBehaviorSanitizer detect invalid memory accesses and undefined behaviour while the program runs, with the exact source line. Build a separate copy (it runs about 2–3× slower):
mkdir -p ~/build-asan && cd ~/build-asan
/path/to/asynch/configure CFLAGS="-O1 -g -fsanitize=address,undefined -fno-omit-frame-pointer" \
LDFLAGS="-fsanitize=address,undefined"
make -j
ASAN_OPTIONS=detect_leaks=0 python3 /path/to/asynch/tests/regression/run_examples.py --asynch ~/build-asan/src/asynch
Compiler warnings reveal many bugs for free:
../configure CFLAGS="-O2 -g -Wall -Wextra -Wno-unused-parameter -Wno-sign-compare"
make 2>&1 | grep warning
Bisecting across history (how R-03 was established): build an old commit in a
separate directory with git worktree and run the same example.
git worktree add /tmp/asynch-2018 cba763b
cd /tmp/asynch-2018 && autoreconf --install && mkdir b && cd b && ../configure CFLAGS="-O2 -DNDEBUG" && make -j