Python API reference¶
Generated from the documentation strings of the package asynch (folder python/ of the repository). The
tutorial is 10. Using ASYNCH from Python. Import names:
from asynch import Simulation, Model, GlobalConfig, io
Every method of Simulation that says collective must be called by every MPI process.
Simulation¶
- class asynch.Simulation(global_file=None, model=None, comm=None, verbose=False, load=True)[source]¶
Bases:
objectAn ASYNCH solver.
- Parameters:
global_file (str or GlobalConfig, optional) – The global file (.gbl) to read, or a
GlobalConfig.model (asynch.model.Model, optional) – A model defined in Python; replaces the model number of the global file.
comm (mpi4py.MPI.Comm, optional) – Communicator (only MPI_COMM_WORLD is fully supported by ASYNCH). Default: every process.
verbose (bool) – Print the progress messages of the C library.
load (bool) – Load everything (network, parameters, initial states, forcings, outputs) right away. Use
load=Falseto register custom outputs (set_output()) first, then callload().
Notes
The steps of
load()can also be called one by one, as insrc/asynch_cli.c.- close()[source]¶
Delete the temporary files and free the solver. The simulation cannot be used afterwards.
- property rank¶
Rank of this process (0 on a single process).
- property num_procs¶
- read_global_file(global_file)[source]¶
Read a global file (path or GlobalConfig). Checks what would otherwise stop the program.
- load(prepare_outputs=True)[source]¶
Load the network, parameters, initial states and forcings (collective), then prepare the outputs of the global file (unless prepare_outputs is False).
- prepare_outputs()[source]¶
Open the temporary files for the hydrographs and prepare the peak and hydrograph outputs (as asynch_cli.c). Checks first that every output of the global file is defined.
- advance(minutes=None, until=None, write=True)[source]¶
Integrate for minutes more, or until the time until [minutes since the start], or to the end of the simulation period (collective). Hydrographs are written to the temporary files if write.
- run(write_outputs=True)[source]¶
Integrate to the end of the period, then write the snapshot, hydrograph and peak flow outputs of the global file (collective). Same results and files as the asynch program.
- write_outputs(suffix=None)[source]¶
Write the snapshot, the hydrographs and the peak flows (collective). Raises if a file could not be written. suffix is appended to the hydrograph file name.
- snapshot(path=None)[source]¶
Write a snapshot of the current states (collective); to path if given (.rec or .h5, as the snapshot type of the global file), else to the file of the global file.
- property time¶
Current time [minutes since the start].
- property duration_total¶
Length of the simulation period [minutes].
- property begin¶
Start of the period (unix time).
- property end¶
End of the period (unix time).
- property num_links¶
Number of links in the network (all processes).
- property num_links_local¶
Number of links computed by this process.
- property link_ids¶
Link ids in location order (the order of every per-link array of this class).
- property model_uid¶
Model number of the global file.
- property global_params¶
Global parameters (a copy). Setting them after load() also recomputes the derived link parameters of every link, so the new values are used from the next step on.
- property num_link_params¶
Parameters per link (read from disk + derived).
- property num_disk_params¶
Parameters per link read from the .prm file.
- get_link_params(link_id)[source]¶
All parameters of a link (read + derived). The link must be stored on this process (always true on one process).
- set_link_params(link_id, values, update=True)[source]¶
Set the first len(values) parameters of a link (usually those read from disk) and, if update, recompute the derived parameters of every link. Ignored for links not stored on this process; call it on every process.
- update_precalculations()[source]¶
Recompute the derived parameters of every link (after several set_link_params(update=False)).
- property max_dim¶
Largest number of states of a link.
- property states¶
array [num_links, max_dim], rows in location order (see link_ids). Every process gets the whole array.
- Type:
Collective. Current states of every link
- set_states(states, time=None)[source]¶
Collective. Replace the states of every link (array [num_links, max_dim], as states) and restart the solver from them at time [minutes since the start, default: now].
- property peaks¶
Collective. (time of peak [min], peak discharge) of every link since the start, arrays in location order. Links without peak output (see the peak links of the global file) give 0.
- property num_forcings¶
- forcing_timestamps(index)[source]¶
(first, last) unix times of a forcing (database and binary forcings).
- set_forcing_state(index, t_0, first_file, last_file)[source]¶
Restart a forcing at time t_0 [min] with the files/times first_file..last_file (forecasting).
- property reservoir_forcing¶
Index of the forcing that feeds the reservoirs, or -1.
- set_initial_file(path)[source]¶
Use another initial-state file (.ini, .uini, .rec or .h5). Call before load().
- property init_timestamp¶
- set_database_connection(conninfo, index)[source]¶
Replace database connection number index (see ASYNCH_DB_LOC_* in src/constants.h).
- property snapshot_path¶
- property peaks_path¶
Peak flow file, without the .pea extension (None if there is no peak output).
- set_output(name, function, states=(), dtype=<class 'float'>)[source]¶
Define the time series output name listed in the “%Components to print” of the global file.
function(link_id, time, y) -> number, where y are the states of the link. states lists the indices of the states that function uses (they are then interpolated at the output times). dtype: float (written as double), np.float32 or int. Call after reading the global file and before load() (or before loading the initial conditions when loading step by step).
Model¶
Define a new hydrological model in Python and let the ASYNCH solver integrate it.
A model is a set of ordinary differential equations solved at every link of the river network. You describe it with names, and give the equations either
as C code (a string): it is compiled once into a small shared library and runs at the speed of the built-in models;
as Python functions: no compiler needed, handy to try an idea;
as Python functions compiled by Numba (
jit="numba"): the same functions, at the speed of C.
Measured on 5 000 links, 2 simulated hours, model 190: built-in 0.10 s, C code 0.10 s, Numba 0.13 s (plus about 1 s of compilation), plain Python functions 4.1 s. All four give identical numbers.
Example: a linear reservoir at every link, dq/dt = (inflow - q) / k, with the rain falling on the hillslope of area A_h going straight into the channel:
from asynch.model import Model
m = Model(
name="linear_reservoir",
states=["q"], # state 0: discharge [m3/s]
global_params=["k"], # residence time [min], same everywhere
params=["A_h"], # read from the .prm file [m2]
forcings=["rain"], # [mm/h]
)
m.equations = '''
double inflow = upstream_q + rain * A_h * (0.001 / 3600.0); /* mm/h * m2 -> m3/s */
d_q = (inflow - q) / k;
'''
Inside the C code, every name is a local variable: the states (q), the global parameters
(k), the link parameters (A_h) and the forcings (rain). Set the time derivative of each
state in d_<state>. upstream_<state> is the sum of that state over the upstream links (for
the states listed in dense, by default the first one). t is the time in minutes. The math
functions of C (pow, exp, fmax, …) are available.
The same model with Python functions:
import numpy as np
def equations(t, y, upstream, gp, p, forcing):
# y: states of this link; upstream: array [num_parents, max_dim] of the upstream states
inflow = upstream[:, 0].sum() + forcing[0] * p[0] * (0.001 / 3600.0)
return [(inflow - y[0]) / gp[0]]
m.equations = equations
Then run it on a network described by a global file:
from asynch import Simulation
with Simulation("my_network.gbl", model=m) as sim:
sim.run()
The model number in the global file is ignored when a model is given (it is still written in the output files).
- class asynch.Model(states, global_params=(), params=(), derived_params=(), forcings=(), dense=None, read_initial=None, nonnegative=None, param_factors=None, area=None, hillslope_area=None, areas_converted_to_m2=False, min_error_tolerances=None, name='custom', jit=None)[source]¶
Bases:
objectA model defined outside the C source of ASYNCH.
- Parameters:
states (list of str) – Names of the states, in order. State 0 should be the discharge [m3/s]: the peak flow output and the
upstream_...sums of the parents use it.global_params (list of str) – Parameters that are the same at every link, given in the global file (in this order).
params (list of str) – Parameters of each link read from the .prm file (in this order).
derived_params (list of str) – Parameters of each link computed once from the others by
precalculations.forcings (list of str) – Time-dependent inputs (rain, evaporation, …), in the order of the global file.
dense (list of str) – States that downstream links need (dense output), default: the first state.
read_initial (list of str) – States read from the initial-state file, in order; must be the first states. Default: all. The others start at 0, unless
initializesets them.nonnegative (None, "all" or "discharge") – Clip the states after every stage: “all” = every state >= 0, “discharge” = state 0 >= 1e-14 and the others >= 0 (as models 190 and 254). Ignored if
consistencyis set.param_factors (dict) – Factor applied to a parameter when it is read, e.g.
{"L": 1000.0}to convert km to m.area (str) – Link parameters holding the upstream area and the hillslope area (written in the peak flow output), and
areas_converted_to_m2if they were converted from km2 (then the peak output converts back).hillslope_area (str) – Link parameters holding the upstream area and the hillslope area (written in the peak flow output), and
areas_converted_to_m2if they were converted from km2 (then the peak output converts back).name (str) – Used for the name of the compiled library.
jit (None or "numba") – With
"numba", the model functions given as Python functions are compiled by Numba (pip install numba) into C functions: they run at the speed of C code instead of calling the Python interpreter at every evaluation. They must then use only what Numba supports (NumPy arrays, math, loops) and return a tuple, a list or an array.
Notes
Equations are set with the attributes
equations(required),precalculations,initializeandconsistency, each either C code (str) or a Python function, andsupport_codefor C helper functions.- property all_params¶
read from disk, then derived.
- Type:
Link parameters
- compile(cache_dir=None)[source]¶
Compile the C code (if any) and load it. Done automatically when the model is used. The library is cached, keyed by the source and flags, in cache_dir (default: $ASYNCH_MODEL_CACHE or ~/.cache/asynch/models). Returns its path, or None for Python equations.
Global files¶
Read, change and write ASYNCH global files (.gbl) from Python.
A global file is read by Read_Global_Data (src/config_gbl.c) one value line after the other;
lines starting with % and empty lines are comments. GlobalConfig follows that reader
field by field, so GlobalConfig.read(path).write(other) gives a file that ASYNCH reads exactly
as the original (the comments are not kept).
Example:
from asynch.config import GlobalConfig, Forcing
cfg = GlobalConfig.read("examples/test_2015.gbl")
cfg.end = "2014-05-01 10:00"
cfg.global_params[3] = 0.5 # runoff coefficient RC of model 190
cfg.hydrographs.path = "out/test_rc05.dat"
cfg.write("examples/test_rc05.gbl")
Paths inside a global file are relative to the folder ASYNCH runs in, not to the file.
- class asynch.GlobalConfig(**kw)[source]¶
Bases:
objectEvery setting of a global file. Attributes, in file order:
model (int), begin, end (
"YYYY-MM-DD HH:MM", datetime or unix time), parameters_in_filenames (0/1), outputs (list of names, e.g.["Time", "State0"]), peakflow_function ("Classic"), global_params (list of floats), steps_stored, steps_transferred, discontinuity_buffer (ints), topology, parameters, initial_state (FileRef), forcings (list ofForcing), dams (FileRef, flag 0 none, 1 .dam, 2 .qvs, 3 .dbc), reservoirs (FileRef, flag 0/1/2, extra = index of the forcing that feeds them), hydrographs (Output), peaks (PeakOutput), hydrograph_links, peak_links (Selection), snapshot (Snapshot), scratch (str), facmin, facmax, fac (floats), rkd (path of a .rkd file, or None), solver (0, 1, 2 or 4; 4 = stiff solver, see chapter 2.4 of the guide), abstol, reltol, abstol_dense, reltol_dense (lists of floats, one per state).
- class asynch.Forcing(flag=0, path=None, increment=None, file_time=None, first=None, last=None)[source]¶
Bases:
objectOne forcing (rain, evaporation, …). The flag selects the source, as in the global file:
flag
source
fields used
0
none
1
.str file (per link time series)
path
2
binary files, one per time step
path, increment, file_time, first, last
3
database
path (.dbc), increment, file_time, first, last
4
.ustr file (same rain everywhere)
path
5
irregular binary files
path, increment, file_time, first, last
6
gzipped binary files
path, increment, file_time, first, last
7
.mon file (monthly recurring values)
path, first, last
8
grid cells
path, increment, first, last
9
database, irregular time steps
path (.dbc), increment, first, last
- class asynch.Output(flag=0, interval=None, path=None, table=None)[source]¶
Bases:
objectHydrograph (time series) output.
flag: 0 none, 1 .dat, 2 .csv, 3 database (path is a .dbc, table needed), 4 .rad, 5 .h5 (packet), 6 .h5 (array). interval: minutes between written values.
- class asynch.PeakOutput(flag=0, path=None, table=None)[source]¶
Bases:
OutputPeak flow output. flag: 0 none, 1 .pea file, 2 database (path is a .dbc, table needed).
- class asynch.Snapshot(flag=0, interval=None, path=None, table=None)[source]¶
Bases:
OutputFinal state output. flag: 0 none, 1 .rec, 2 database (.dbc + table), 3 .h5, 4 .h5 every interval minutes (the path gets the time appended).
Parallel runs¶
Start parallel (MPI) runs from Python, for instance from a Jupyter notebook.
A Python program cannot become several MPI processes after it has started, so these functions start new ones with mpiexec, and wait for them:
import asynch asynch.run_parallel(“my_basin.gbl”, 4) # the asynch program on 4 processes asynch.run_script_parallel(“my_study.py”, 4) # a Python script on 4 processes (it uses Simulation)
The same as typing mpiexec -n 4 asynch my_basin.gbl or mpiexec -n 4 python3 my_study.py in a terminal. The output files are those the global file names, as with the program.
- asynch.run_parallel(global_file, processes, cwd=None, program=None, mpiexec=None, mpiexec_args=(), capture=False, timeout=None)[source]¶
Runs the asynch program on a global file with processes MPI processes, and waits for it.
- Parameters:
global_file (str) – The global file. As with the program, the file names inside it are relative to cwd.
processes (int) – Number of MPI processes.
cwd (str, optional) – Folder to run in (default: the current folder).
program (str, optional) – Paths of the asynch program and of mpiexec (default: found by find_program and find_mpiexec).
mpiexec (str, optional) – Paths of the asynch program and of mpiexec (default: found by find_program and find_mpiexec).
mpiexec_args (sequence of str) – Extra options for mpiexec, e.g. [”–hostfile”, “hosts”].
capture (bool) – Return the printed text in .stdout instead of showing it.
timeout (float, optional) – Seconds before the run is stopped.
- Returns:
Raises ParallelRunError if the run fails.
- Return type:
subprocess.CompletedProcess
- asynch.run_script_parallel(script, processes, *args, cwd=None, python=None, mpiexec=None, mpiexec_args=(), capture=False, timeout=None)[source]¶
Runs a Python script with processes MPI processes (mpiexec -n processes python script args…), and waits.
Every process runs the whole script; Simulation shares the links between them (chapter 10.9 of the guide). The other parameters are those of run_parallel; python defaults to this Python.
Input and output files¶
Read ASYNCH output files and write ASYNCH input files.
Readers (outputs of a run) return a dictionary keyed by link id:
from asynch import io
hydro = io.read_hydrographs("outputs/my_basin.dat") # .dat, .csv or .h5
t, q = hydro[2527][:, 0], hydro[2527][:, 1]
peaks = io.read_pea("examples/out_2015/test.pea") # {link: (area, time, peak)}
state = io.read_snapshot("examples/out_2015/test.rec") # .rec or .h5
The meaning of the hydrograph columns depends on the “%Components to print” block of the global file.
Writers (inputs of a run) create the text formats of docs/input_output.rst from Python data:
io.write_rvr("net.rvr", {1: [2, 3], 2: [], 3: []}) # link -> upstream links
io.write_prm("net.prm", {1: [1.2, 0.5, 0.1], 2: [...], 3: [...]}) # link -> parameters
io.write_uini("net.uini", 190, [1e-6, 0.0, 0.0])
io.write_ustr("rain.ustr", [(0, 20.0), (60, 0.0)]) # (minute, value)
Only NumPy is needed, plus h5py for the .h5 files.
- asynch.io.read_dat(path)[source]¶
Hydrographs in .dat format (global file: “1 <interval> file.dat”).
Layout: number of links, number of printed components, then for every link: “<link id> <number of rows>” followed by that many rows of components. Returns {link id: array [rows, components]}.
- asynch.io.read_csv(path)[source]¶
Hydrographs in .csv format (global file: “2 <interval> file.csv”).
Layout: a header line “Link <id> , , ,” per link, a line of output names, then one row per output time with the components of every link side by side. Returns {link id: array [rows, components]}.
- asynch.io.read_pea(path)[source]¶
Peak flows in .pea format (“Classic” peakflow function).
Layout: number of links, model id, then per link: “<id> <upstream area [km2]> <time of peak [min]> <peak discharge [m3/s]>”. Returns {link id: (area, time_of_peak, peak_value)}.
- asynch.io.read_rec(path)[source]¶
Snapshot (all states of all links at one time) in .rec format.
Layout: model id, number of links, time, then per link a line with its id and a line with its states. Returns {link id: array of states}.
- asynch.io.read_h5_snapshot(path)[source]¶
Snapshot in .h5 format (global file snapshot flag 3 or 4). Returns {link id: array of states}.
- asynch.io.read_h5_hydrographs(path)[source]¶
Hydrographs in .h5 format, packet (flag 5) or array (flag 6) layout.
Returns {link id: array [rows, components]}. For the array layout, the first column is the time (unix time) and the others are the printed components.
- asynch.io.read_hydrographs(path)[source]¶
Hydrographs from a .dat, .csv or .h5 file. Returns {link id: array [rows, components]}.
- asynch.io.read_snapshot(path)[source]¶
Snapshot from a .rec or .h5 file. Returns {link id: array of states}.
- asynch.io.write_rvr(path, parents)[source]¶
Topology (.rvr). parents: {link id: list of upstream link ids}, in the order to write.
- asynch.io.write_prm(path, params)[source]¶
Link parameters (.prm). params: {link id: sequence of the parameters read from disk}.
- asynch.io.write_uini(path, model, values, time=0.0)[source]¶
Initial state, the same at every link (.uini).
- asynch.io.write_ini(path, model, states, time=0.0)[source]¶
Initial state per link (.ini); states: {link id: values read from file}. The .rec format is the same with every state of every link (write_rec).
- asynch.io.write_rec(path, model, states, time=0.0)¶
Initial state per link (.ini); states: {link id: values read from file}. The .rec format is the same with every state of every link (write_rec).
- asynch.io.write_str(path, series)[source]¶
Forcing per link (.str); series: {link id: [(time [min], value), …]} (constant between times).
- asynch.io.write_ustr(path, points)[source]¶
Forcing uniform in space (.ustr); points: [(time [min], value), …] (constant between times).
- asynch.io.write_mon(path, values)[source]¶
Monthly recurring forcing (.mon): 12 values, January first.
- asynch.io.read_rvr(path)[source]¶
Topology (.rvr). Returns {link id: list of upstream link ids}, in file order.
- asynch.io.read_prm(path)[source]¶
Link parameters (.prm). Returns {link id: array of parameters}. The number of parameters per link is found from the file (all links have the same number).
- asynch.io.write_binary_forcing(prefix, frames, compress=False)[source]¶
Binary forcing files (global file flag 2, or 6 if compress): one file per time step, named prefix + index.
frames: {index: values}, the values of every link in the order of the topology (.rvr) file; the file with index first + k applies from minute k * (time resolution of the global file). Written as 32-bit floats in big-endian byte order (ASYNCH swaps the bytes when reading). With compress, the files are gzipped (prefix + index + “.gz”). Returns the list of files written.
- asynch.io.write_irregular_binary_forcing(prefix, frames)[source]¶
Irregular binary forcing files (global file flag 5): one file per change, named prefix + unix time.
frames: {unix time: {link id: value}}; links not listed get 0 (the format is meant for sparse fields such as rain). Written in little-endian byte order: number of links (uint32), then link id (uint32) and value (float32) for each link. Returns the list of files written.
Low level: the C functions¶
Low-level binding: loads libasynch and declares the C prototype of every function the package uses.
Nothing here knows the layout of an ASYNCH structure: the C library is only handled through an opaque
pointer (AsynchSolver*) and plain numbers and arrays. That is what makes the binding robust: a change
of the C structures cannot silently shift memory the way the old py/asynch_interface.py did.
The library is searched in this order:
the path in the environment variable
ASYNCH_LIBRARY;next to this package (a copy placed there by the user);
the build tree of this repository (
src/.libs), when the package is used from a source checkout;the system library path (
ctypes.util.find_library),$CONDA_PREFIX/lib,/usr/local/lib.
- asynch._lib.c_double_p¶
alias of
LP_c_double
- asynch._lib.c_uint_p¶
alias of
LP_c_uint
- asynch._lib.DIFFERENTIAL¶
dy/dt of one link (DifferentialFunc)
- asynch._lib.V_P¶
the same signatures with every pointer as a plain address (an int, or None for NULL): converting an address is much cheaper than building a ctypes pointer object, which matters in functions called millions of times
- asynch._lib.CHECK_CONSISTENCY¶
consistency check of one link (CheckConsistencyFunc)
- asynch._lib.PRECALCULATIONS¶
derived parameters of one link (SpecPrecalculationsFunc)
- asynch._lib.INITIALIZE¶
initial states of one link (SpecInitializeFunc)
- asynch._lib.OUTPUT_INT¶
custom time series outputs (OutputIntCallback, OutputDoubleCallback, OutputFloatCallback)
- asynch._lib.PEAKFLOW_OUTPUT¶
custom peak flow output (PeakflowOutputCallback); writes one line (< 256 bytes) into the buffer
- asynch._lib.NOT_BOUND = {'Asynch_Copy_Local_OutputUser_Data': 'see Asynch_Create_OutputUser_Data', 'Asynch_Create_OutputUser_Data': 'per-link user data for C output callbacks; Python callbacks keep their own data', 'Asynch_Custom_Model': 'takes an AsynchModel structure; use the model specification (Asynch_Install_Model)', 'Asynch_Custom_Partitioning': 'the partitioning function works on the internal Link structure', 'Asynch_Free_OutputUser_Data': 'see Asynch_Create_OutputUser_Data', 'Asynch_Get_Links': 'returns internal Link structures; use the link accessors of asynch_api.h', 'Asynch_Get_Links_Proc': 'returns internal Link structures; use the link accessors of asynch_api.h', 'Asynch_Init': 'takes an MPI_Comm, whose type depends on the MPI library; use Asynch_Init_World or Asynch_Init_Fortran_Comm', 'Asynch_Set_Size_Local_OutputUser_Data': 'see Asynch_Create_OutputUser_Data'}¶
C functions that are deliberately not bound, and why (see docs/guide/10_python.md)
- asynch._lib.mpich_libdir()[source]¶
Folder of libmpi.so.12 installed by the mpich package of PyPI (the MPI of the self-contained wheel), or None. pip puts it in the lib folder of the environment (or of the user base for pip install –user).