Large problems
Run stacks that fit neither on the GPU nor in host memory.
A single SRC sweep keeps, besides the input and output trains, one sketched
environment per site. For stacks with a large bond, such as N . V . M . U with a
bond of several thousand in M, these environments and the intermediates of the
contractions no longer fit on a GPU, and often not in host memory either.
src_method then plans the sweep to the memory at hand:
- the input cores are read one site at a time, so a train can live on disk;
- every contraction runs in batches of sketch columns (or rows), sized to the GPU budget;
- each environment is kept on the GPU, in host memory or in a scratch directory, the most recent ones on the fastest tier.
None of this changes the result beyond floating-point rounding: the random draws and the mathematics are those of a sweep that fits in memory.
Budgets
The budgets are set with Resources, passed to src, apply or compress:
from src_method import Resources, src
out = src(
N, V, M, U,
chi_out=2000,
dtype=np.complex128,
device="gpu",
resources=Resources(gpu_memory="36GB", scratch_dir="/local/scratch"),
)Every field left unset is detected when the call starts:
| Field | Default |
|---|---|
gpu_memory | free device memory, plus the free bytes of CuPy's pool, minus max(10%, 1 GiB) |
host_memory | MemAvailable from /proc/meminfo, minus 10% |
scratch_dir | tempfile.gettempdir(), which honours $TMPDIR |
Sizes are byte counts or strings: "36GB" is 36 * 10**9 bytes, "36GiB" is
36 * 2**30. On the CPU there is a single budget, host_memory. Host memory is
measured when the call starts, so inputs already held in memory are not counted
twice.
The sweep is planned to fit the GPU budget, counting the bytes its arrays use.
CuPy's memory pool also holds blocks that are split and only partly in use, so
during the sweep the pool is capped at the GPU budget plus the margin of
max(10%, 1 GiB) of the device. That margin is the one detection leaves free, so
with a detected budget the cap is the free device memory. An estimate that falls
short by more than the margin fails at once rather than when some other
allocation does; the previous limit is restored afterwards. If a single site does not fit the
budget even with batches of one column, src raises MemoryError naming the site.
Inputs on disk
A train is any sequence of per-site array-likes with shape, dtype, ndim and
np.asarray support (the SiteLike protocol, exported for type annotations):
NumPy arrays, np.memmap, zarr or HDF5 datasets. Each core
is read when the sweep reaches its site, once per pass, and a background thread
reads the next site while the current one is computed. Raw .npy files opened with
np.load(path, mmap_mode="r") are the fastest option; compressed formats may be
limited by decompression.
The scratch directory
Environments that fit in neither budget are written to a per-process directory,
<scratch_dir>/src-<pid>-<id>/, one file per site. Use node-local disk: at the
reference size of 50 sites, D_M = 4000 and chi_out = 2000 in complex128, up to
410 GB are written and read back once. The directory is removed when the call
returns, also after an error; a process killed with SIGKILL leaves it behind, and
its name identifies the process.
Reading the plan
With DEBUG logging on for the src_method logger (see
Logging), every sweep logs its plan as SRC plan: ...:
prefetch: whether the next site is read ahead;device peak,host peak,disk: the planned peaks, in bytes;scratch: the scratch directory, when anything spills to disk;tiers: where the environment of each site lives (device,hostordisk);batches: for each site, the batch of the environment, sketch and projection steps.
It then logs the time of each pass and, at the end, SRC stalls: the seconds
the sweep waited for input cores and for environments. If they are a large
fraction of the run, the disk is too slow for the compute.