Sampling backends
Compile-time device selection through `.on::<B>()` and `with_backend(handle)` — the `Backend` marker, its capability subtraits, one fallback rule.
Sampling backends
Every process selects where it samples — CPU, CUDA, Metal or Apple
Accelerate — at compile time. A process is parameterised
by a backend marker B (defaulting to Cpu); the Backend trait
monomorphises sample / sample_par to that backend with no runtime branch.
This replaces the ad-hoc per-backend methods (sample_cuda,
sample_gpu, …) with one uniform selector. Those methods are now crate-private
(sample_*_impl, wired up in stochastic_rs::stochastic::device), so
.on::<B>() is not merely the recommended API — it is the only one.
Switching backend
Re-type the process by naming a backend marker, then sample as usual.
Cuda only exists when the cuda feature is compiled (an
NVIDIA CUDA 12.x toolkit and GPU), which the CI runners that build this
site do not have, so this block is illustrative rather than compiled:
use stochastic_rs::simd_rng::Unseeded;
use stochastic_rs::stochastic::device::Cuda;
use stochastic_rs::stochastic::noise::fgn::Fgn;
use stochastic_rs::traits::ProcessExt;
let cpu = Fgn::<f32, _>::new(0.7, 65_536, None, Unseeded).sample(); // default Cpu
let gpu = Fgn::<f32, _>::new(0.7, 65_536, None, Unseeded)
.on::<Cuda>()
.sample(); // same call, runs on the GPU.on::<B>() consumes the process and returns it re-typed to backend B — the
fields are moved, nothing is copied, and the choice is resolved at compile time.
The Backend marker and the FgnBackend capability
Backend itself is a bare device marker — it carries no algorithm. What a
device can compute is expressed by capability subtraits, and every process
carries the backend as its last type parameter (Name<T, S, B = Cpu>), so
.on::<B>() exists on all of them; the capability the process's bound names
says which markers it accepts:
HostBackend: Backend— the process samples on the CPU through its ownProcessExtsampler.CpuandAccelerateimplement it. Every process accepts at least these; it is the bound a new process starts with, and gaining a kernel later merely widens it, which breaks no caller. No process is bound on it alone today.FgnBackend<T>: Backend— fractional Gaussian noise on the device (the circulant-embedding FFT pipeline);Fgnis bound on it, and the fractional processes on the engine take their increments from it.SheetBackend<T>: Backend— the two-dimensional circulant embedding behind the fractional Brownian fieldFbs.EulerBackend<T>: Backend— the Euler engine's device kernels, which 129 processes step through;CpuandAccelerateimplement it as the host sampler.
All three device capabilities take the scalar, and a device implements them
only for the precision its kernels compute in: Cuda for f32 and f64,
Metal for f32, the CPU devices for both. Fgn<f64> or
Gbm<f64> on Metal is a compile error; nothing runs in f32 behind an
f64 type. Generic code spells the scalar in the bound: B: EulerBackend<T>.
Bound on Backend to say "device-parameterised"; bound on a capability when
the code actually samples through it.
| Marker | Feature | Backs onto | Hardware |
|---|---|---|---|
Cpu | — (always available, default) | ndrustfft + CPU SIMD, rayon batches | any CPU |
Cuda | cuda | cudarc + cuFFT, JIT Philox RNG | NVIDIA, CUDA 12.x |
Metal | metal | hand-written MSL via the metal crate | Apple GPU (f32) |
Accelerate | accelerate | Apple vDSP FFT (CPU), rayon batches | macOS |
Backend, Cpu, HostBackend, FgnBackend, SheetBackend and
EulerBackend are re-exported at stochastic_rs::traits; the GPU markers
live under stochastic_rs::stochastic::device and only exist when their
feature is compiled. Selecting an unavailable backend is therefore a compile
error, not a silent runtime fallback.
// docs: concepts/backends#the-backend-trait
//! Backs the always-available-marker import example on the backends
//! concept page: every marker is a [`Backend`] (the bare device marker),
//! and every marker shipped today also carries the [`FgnBackend`] fGN
//! sampling capability — for the scalar its kernels compute in: the host
//! for `f64` and `f32`, an Apple GPU for `f32` alone.
#[cfg(feature = "metal")]
use stochastic_rs::stochastic::device::Metal;
use stochastic_rs::traits::Backend;
use stochastic_rs::traits::Cpu;
use stochastic_rs::traits::FgnBackend;
#[test]
fn cpu_marker_is_a_backend_with_the_fgn_capability() {
fn assert_marker<B: Backend>() {}
fn assert_fgn_f64<B: FgnBackend<f64>>() {}
fn assert_fgn_f32<B: FgnBackend<f32>>() {}
assert_marker::<Cpu>();
assert_fgn_f64::<Cpu>();
assert_fgn_f32::<Cpu>();
#[cfg(feature = "metal")]
assert_marker::<Metal>();
#[cfg(feature = "metal")]
assert_fgn_f32::<Metal>();
}Which processes support it
The per-family inventory — which process runs on which device, at what precision, and what the Python entry point is — lives on GPU support; this section explains the mechanism.
All 131 processes carry the parameter and reach a device: 129 through the
Euler engine (EulerBackend), Fgn through its FFT pipeline (FgnBackend)
and Fbs through the sheet pipeline (SheetBackend). What a device cannot
carry is decided per configuration, not per type — see the one fallback rule
below.
One fallback rule
Some configurations exceed what the kernels carry: a grid past a per-path
array, a dimension past the four state slots, a coefficient written as a
closure, a point process in horizon mode. Every process answers
device_ready() — a method of ProcessExt, true by default — for the
configuration it holds, and a false never fails: on any backend the process
then samples on the host through its own sampler, bit-identically to the
Cpu build. try_sample is Ok on that path, and sample never panics for
a configuration — only a device that fails to open, compile or launch does
that. So gbm.on::<Cuda>() always runs on the GPU, while
wishart.on::<Cuda>() runs there for d = 2 and on the host for d = 3.
Which it is, and why, the process will say. device_fallback() returns the
reason as Some("a Wishart outside the two-dimensional cone the kernel steps") and None when a kernel carries the configuration whole;
device_ready() is the absence of a reason. The reason matters because the
fallback is silent otherwise, and silence costs an order of magnitude:
if let Some(why) = wishart.device_fallback() {
eprintln!("sampling on the host: {why}");
}Both are on ProcessExt, and both are available from Python as
process.device_fallback() and process.device_ready().
A process will also say where it is pointed. process.backend() hands back
the handle it holds — the ordinal and batch budget a type parameter does not
record — and process.probe() opens the device behind it and describes it,
or reports why it cannot be used.
Choosing the device
A backend is a value, not only a type. .on::<B>() re-types the process with
B's default handle; with_backend(handle) takes an explicit one:
use stochastic_rs::stochastic::device::Cuda;
let on_second_gpu = gbm.with_backend(Cuda::new(1).with_batch_budget(256 << 20));Cuda { ordinal, batch_budget } and Metal { ordinal, batch_budget } carry
the device ordinal (CUDA ordinal, index into Device::all()) and the bytes of
path data one launch may hold. Their
Default reads STOCHASTIC_RS_DEVICE and STOCHASTIC_RS_DEVICE_BATCH_BYTES
(0 and 1 GiB when unset), so an unchanged program picks its device from the
environment, and there is no process-wide state: two processes on two GPUs are
two handles. Cpu and Accelerate carry nothing. probe() on a handle
reports the device it opens.
Probing a device and handling its failures
A missing GPU, a runtime that will not start or a kernel that will not compile is an environmental failure, not a programming error, so it is reported as a value:
use stochastic_rs::stochastic::device::{Backend, DeviceError, Metal};
match Metal::default().probe() {
Ok(info) => println!("{} on {} ({:?})", info.backend, info.name, info.precisions),
Err(DeviceError::Unavailable(why)) => eprintln!("no device: {why}"),
Err(other) => eprintln!("{other}"),
}probe() opens the device behind the handle once and describes it
(backend, name, precisions, ordinal). After a successful probe the
plain sample() / sample_par(m) calls keep the ProcessExt signature and
panic with the device's error if something still fails at launch time (an
allocation, typically); every process also offers
try_sample() / try_sample_par(m) -> Result<_, DeviceError> on ProcessExt
(always Ok on the host devices, bit-identical to the plain calls) and
the capability traits expose try_euler_paths / try_generate_batch on the
handle for generic code. DeviceError has three kinds: Unavailable,
Compile, Launch. On the CPU devices every probe is Ok and every try_* call is
Ok with the same paths as its plain twin. From Python, a device failure is
a RuntimeError with the same message.
Precision
Apple GPUs have no f64, so Metal is f32 only: it implements the
capabilities for f32 alone, and an f64 process on it is a compile error. Cuda, Accelerate and Cpu
follow the process's T parameter in f32 and f64.
When is GPU worth it?
The discrete-GPU win shows up for large n (≥ 16 k) and batched sampling,
while for small n the CPU SIMD path wins on launch latency; on Apple Silicon
Metal leads single paths from n = 16 k and the GPU back-ends win batches by a
modest margin since the device state became reusable across calls.
See the Benchmarks page for the cross-over numbers, and the
Feature flags page for how to enable each
backend.