src_method

Large problems

Run stacks that fit neither on the GPU nor in host memory.

A single SRC sweep keeps, besides the input and output trains, one sketched environment per site. For stacks with a large bond, such as N . V . M . U with a bond of several thousand in M, these environments and the intermediates of the contractions no longer fit on a GPU, and often not in host memory either. src_method then plans the sweep to the memory at hand:

  • the input cores are read one site at a time, so a train can live on disk;
  • every contraction runs in batches of sketch columns (or rows), sized to the GPU budget;
  • each environment is kept on the GPU, in host memory or in a scratch directory, the most recent ones on the fastest tier.

None of this changes the result beyond floating-point rounding: the random draws and the mathematics are those of a sweep that fits in memory.

Budgets

The budgets are set with Resources, passed to src, apply or compress:

from src_method import Resources, src

out = src(
    N, V, M, U,
    chi_out=2000,
    dtype=np.complex128,
    device="gpu",
    resources=Resources(gpu_memory="36GB", scratch_dir="/local/scratch"),
)

Every field left unset is detected when the call starts:

FieldDefault
gpu_memoryfree device memory, plus the free bytes of CuPy's pool, minus max(10%, 1 GiB)
host_memoryMemAvailable from /proc/meminfo, minus 10%
scratch_dirtempfile.gettempdir(), which honours $TMPDIR

Sizes are byte counts or strings: "36GB" is 36 * 10**9 bytes, "36GiB" is 36 * 2**30. On the CPU there is a single budget, host_memory. Host memory is measured when the call starts, so inputs already held in memory are not counted twice.

The sweep is planned to fit the GPU budget, counting the bytes its arrays use. CuPy's memory pool also holds blocks that are split and only partly in use, so during the sweep the pool is capped at the GPU budget plus the margin of max(10%, 1 GiB) of the device. That margin is the one detection leaves free, so with a detected budget the cap is the free device memory. An estimate that falls short by more than the margin fails at once rather than when some other allocation does; the previous limit is restored afterwards. If a single site does not fit the budget even with batches of one column, src raises MemoryError naming the site.

Inputs on disk

A train is any sequence of per-site array-likes with shape, dtype, ndim and np.asarray support (the SiteLike protocol, exported for type annotations): NumPy arrays, np.memmap, zarr or HDF5 datasets. Each core is read when the sweep reaches its site, once per pass, and a background thread reads the next site while the current one is computed. Raw .npy files opened with np.load(path, mmap_mode="r") are the fastest option; compressed formats may be limited by decompression.

The scratch directory

Environments that fit in neither budget are written to a per-process directory, <scratch_dir>/src-<pid>-<id>/, one file per site. Use node-local disk: at the reference size of 50 sites, D_M = 4000 and chi_out = 2000 in complex128, up to 410 GB are written and read back once. The directory is removed when the call returns, also after an error; a process killed with SIGKILL leaves it behind, and its name identifies the process.

Reading the plan

With DEBUG logging on for the src_method logger (see Logging), every sweep logs its plan as SRC plan: ...:

  • prefetch: whether the next site is read ahead;
  • device peak, host peak, disk: the planned peaks, in bytes;
  • scratch: the scratch directory, when anything spills to disk;
  • tiers: where the environment of each site lives (device, host or disk);
  • batches: for each site, the batch of the environment, sketch and projection steps.

It then logs the time of each pass and, at the end, SRC stalls: the seconds the sweep waited for input cores and for environments. If they are a large fraction of the run, the disk is too slow for the compute.

On this page