Verification study
These checks establish an initial working package, not independent scientific certification.
See the LuMod case study for the complete calibration-method comparison, hydrographs and sensitivity results.
LuMod example calibration
Run: python examples/real_lumod_basins.py --workers 2 --maxiter 100.
Three records distributed with LuMod 0.1.3.0, observed discharge interpreted as m³/s, upstream GR4J PET calculation, first 70% of each record used for training, first 365 daily steps excluded as warmup, remaining 30% used for validation. GR4J, differential evolution, seed 42, population multiplier 6, package search bounds and initial-state defaults. Results are local to this configuration; no observation data were invented.
LuMod example |
Record length |
Calibration NSE |
Validation NSE |
Validation KGE (2009) |
|---|---|---|---|---|
1 |
10,957 days |
0.2401 |
0.1532 |
0.4304 |
2 |
3,287 days |
0.7175 |
0.5908 |
0.7495 |
3 |
1,460 days |
0.7188 |
0.6680 |
0.7712 |
All jobs completed without errors. The optimizer reported convergence, which is not proof of a global optimum or acceptable hydrological performance. Example 1 has a poor fit under this configuration and should not be described as a successful predictive model. Further work includes basin-specific priors, initial-state sensitivity, alternative PET, multiple seeds, and independent reference validation.
Simulation microbenchmark
python benchmarks/compare_lumod.py: GR4J, LuMod example 2 (3,287 daily steps), identical forcing and parameters, parity checked, JIT warmed, 20 repetitions, macOS ARM64, Python 3.14.3.
LuMod public API median: 0.571 ms.
Prepared-array BasinForge adapter median: 0.220 ms.
Ratio: 2.59× for this particular warmed workload.
The difference is primarily wrapper overhead; equations are the same upstream kernels. This does not demonstrate universally faster calibration or linear multi-process scaling. Process startup/JIT and hardware affect results.
Automated checks
python -m pytest -q covers the Python engines, unit conversions and leap years, forcing continuity, missing-observation handling, deterministic seeds, training/validation isolation, common-vector fitting, serial/process agreement, completed-job resume fingerprints, incompatible model rejection, actual record calibration, CLI/no-overwrite behavior, multi-start selection, scenario quantiles and report fingerprint/HTML escaping. The six added hydromodel kernels are compared to independent pinned upstream checkout outputs. SMART is checked against hourly upstream execution; monthly ABCD against a literal quadratic reference and step-by-step storage conservation.
The suite provides optional Octave integration tests. Enable them with BASINFORGE_TEST_MARRMOT=1; use BASINFORGE_OCTAVE and BASINFORGE_OCTAVE_PACKAGE_LIST for a nonstandard installation.
Final development run: 145 tests passed in 44.22 seconds, with the optional Octave integration enabled. This is one macOS ARM64/Python 3.14.3/Octave 10.3.0 environment; Linux CI workflows are supplied but have not been run remotely. Cross-platform and cross-solver-version results are not certified by this local run.
Publishing checks subsequently passed 146 tests locally, including the site builder; GitHub’s Python 3.11/3.13/3.14 jobs passed. The first Linux Octave run found one cross-runtime regression difference in MARRMOT_33, up to 2.96e-5 mm on the controlled fixture, rather than an interface failure. Its stored-reference comparison now uses an explicitly documented 1e-4 mm absolute envelope; the other 46 structures retain 1e-8 mm absolute / 1e-7 relative tolerances. Public discharge validation is unchanged: materially negative flows are rejected. The stored reference is not an exact floating-point oracle across all BLAS/Octave versions.
python -m build generates a wheel/source archive including model MATLAB sources, the parameter manifest, worker wrapper and retained licenses. Version 0.1.0 is published on PyPI. No independent equivalence to airGR or the original HBV-light executable is claimed.
Optional MARRMoT verification
On this development machine, Octave 10.3.0 with optim 1.6.3, statistics 1.7.7 and struct 1.0.18 was installed in a temporary portable environment. All 47 structures are tested against separately invoked official source at revision eeb7e152d4bc194a6fa9407e3d79e2e29ba7e201. Reference generation uses benchmarks/generate_references.py with independent source checkouts, not the BasinForge adapter. Controlled forcing is explicitly artificial, not observed data. A fixed solver restart seed is applied in both reference and adapter runs.
All 47 also enter single-parameter, two-candidate calibration on a short wet fixture, including chronological holdout. These are interface checks, not effective full-model calibrations. Additional tests check full-history continuity, diagnostics, spawned multi-basin agreement and request-history-independent fallback solves. Upstream MARRMOT_33 produces roughly −1.6e-6 mm on the dry/empty-store reference fixture: raw parity is checked and the public simulation API is explicitly tested to reject that invalid discharge. It is not silently clipped. Arbitrary forcing/parameter combinations can still fail upstream numerical solvers.
Example 2 model comparison
Command: python examples/compare_real_models.py --models GR4J GR5J HYMOD_CLASSIC XAJ MARRMOT_29 --samples 12 --output results/v0.2-real-comparison.
Actual LuMod example-2 observations (3,287 daily steps), LuMod GR4J-derived PET, first 70% training, 365-step warmup, remaining 30% validation; Latin-hypercube search, 12 candidates, seed 42 and default model bounds. No observations were invented. Training loss selects the fit, not validation performance.
Implementation |
Training NSE |
Validation NSE |
Validation KGE 2009 |
Calibration elapsed |
|---|---|---|---|---|
GR4J / LuMod |
0.5947 |
0.5024 |
0.7514 |
0.003 s |
GR5J / hydromodel |
−0.0045 |
−0.2551 |
0.3441 |
0.324 s |
Classic HYMOD / hydromodel |
0.4413 |
0.4305 |
0.6638 |
0.159 s |
XAJ / hydromodel |
0.3860 |
0.4058 |
0.6754 |
1.434 s |
HYMOD structure / MARRMOT_29 |
0.5831 |
0.4899 |
0.6591 |
59.008 s |
Every search exhausted its budget; none establishes optimizer convergence. GR5J performed poorly and must not be advertised as a good predictive fit. Elapsed times exclude baseline/preflight startup, but include candidate evaluation and final train/full simulations; they are not equal-work engine speed rankings, since parameter counts/equations/solvers differ. The optional Octave backend is demonstrably slower here. Each fit has a local hydrograph/flow-duration report, parameters and simulation CSV. These results are not full field validation of all 61 implementations.