stochastic-rs
Concepts

Sampling backends

Compile-time device selection through `.on::<B>()` and `with_backend(handle)` — the `Backend` marker, its capability subtraits, one fallback rule.

Sampling backends

Every process selects where it samples — CPU, CUDA, Metal or Apple Accelerate — at compile time. A process is parameterised by a backend marker B (defaulting to Cpu); the Backend trait monomorphises sample / sample_par to that backend with no runtime branch.

This replaces the ad-hoc per-backend methods (sample_cuda, sample_gpu, …) with one uniform selector. Those methods are now crate-private (sample_*_impl, wired up in stochastic_rs::stochastic::device), so .on::<B>() is not merely the recommended API — it is the only one.

Switching backend

Re-type the process by naming a backend marker, then sample as usual. Cuda only exists when the cuda feature is compiled (an NVIDIA CUDA 12.x toolkit and GPU), which the CI runners that build this site do not have, so this block is illustrative rather than compiled:

use stochastic_rs::simd_rng::Unseeded;
use stochastic_rs::stochastic::device::Cuda;
use stochastic_rs::stochastic::noise::fgn::Fgn;
use stochastic_rs::traits::ProcessExt;

let cpu = Fgn::<f32, _>::new(0.7, 65_536, None, Unseeded).sample(); // default Cpu

let gpu = Fgn::<f32, _>::new(0.7, 65_536, None, Unseeded)
  .on::<Cuda>()
  .sample(); // same call, runs on the GPU

.on::<B>() consumes the process and returns it re-typed to backend B — the fields are moved, nothing is copied, and the choice is resolved at compile time.

The Backend marker and the FgnBackend capability

Backend itself is a bare device marker — it carries no algorithm. What a device can compute is expressed by capability subtraits, and every process carries the backend as its last type parameter (Name<T, S, B = Cpu>), so .on::<B>() exists on all of them; the capability the process's bound names says which markers it accepts:

  • HostBackend: Backend — the process samples on the CPU through its own ProcessExt sampler. Cpu and Accelerate implement it. Every process accepts at least these; it is the bound a new process starts with, and gaining a kernel later merely widens it, which breaks no caller. No process is bound on it alone today.
  • FgnBackend<T>: Backend — fractional Gaussian noise on the device (the circulant-embedding FFT pipeline); Fgn is bound on it, and the fractional processes on the engine take their increments from it.
  • SheetBackend<T>: Backend — the two-dimensional circulant embedding behind the fractional Brownian field Fbs.
  • EulerBackend<T>: Backend — the Euler engine's device kernels, which 129 processes step through; Cpu and Accelerate implement it as the host sampler.

All three device capabilities take the scalar, and a device implements them only for the precision its kernels compute in: Cuda for f32 and f64, Metal for f32, the CPU devices for both. Fgn<f64> or Gbm<f64> on Metal is a compile error; nothing runs in f32 behind an f64 type. Generic code spells the scalar in the bound: B: EulerBackend<T>.

Bound on Backend to say "device-parameterised"; bound on a capability when the code actually samples through it.

MarkerFeatureBacks ontoHardware
Cpu— (always available, default)ndrustfft + CPU SIMD, rayon batchesany CPU
Cudacudacudarc + cuFFT, JIT Philox RNGNVIDIA, CUDA 12.x
Metalmetalhand-written MSL via the metal crateApple GPU (f32)
AccelerateaccelerateApple vDSP FFT (CPU), rayon batchesmacOS

Backend, Cpu, HostBackend, FgnBackend, SheetBackend and EulerBackend are re-exported at stochastic_rs::traits; the GPU markers live under stochastic_rs::stochastic::device and only exist when their feature is compiled. Selecting an unavailable backend is therefore a compile error, not a silent runtime fallback.

tests/doctest_concepts_backends_import.rs
// docs: concepts/backends#the-backend-trait
//! Backs the always-available-marker import example on the backends
//! concept page: every marker is a [`Backend`] (the bare device marker),
//! and every marker shipped today also carries the [`FgnBackend`] fGN
//! sampling capability — for the scalar its kernels compute in: the host
//! for `f64` and `f32`, an Apple GPU for `f32` alone.

#[cfg(feature = "metal")]
use stochastic_rs::stochastic::device::Metal;
use stochastic_rs::traits::Backend;
use stochastic_rs::traits::Cpu;
use stochastic_rs::traits::FgnBackend;

#[test]
fn cpu_marker_is_a_backend_with_the_fgn_capability() {
  fn assert_marker<B: Backend>() {}
  fn assert_fgn_f64<B: FgnBackend<f64>>() {}
  fn assert_fgn_f32<B: FgnBackend<f32>>() {}
  assert_marker::<Cpu>();
  assert_fgn_f64::<Cpu>();
  assert_fgn_f32::<Cpu>();
  #[cfg(feature = "metal")]
  assert_marker::<Metal>();
  #[cfg(feature = "metal")]
  assert_fgn_f32::<Metal>();
}

Which processes support it

The per-family inventory — which process runs on which device, at what precision, and what the Python entry point is — lives on GPU support; this section explains the mechanism.

All 131 processes carry the parameter and reach a device: 129 through the Euler engine (EulerBackend), Fgn through its FFT pipeline (FgnBackend) and Fbs through the sheet pipeline (SheetBackend). What a device cannot carry is decided per configuration, not per type — see the one fallback rule below.

One fallback rule

Some configurations exceed what the kernels carry: a grid past a per-path array, a dimension past the four state slots, a coefficient written as a closure, a point process in horizon mode. Every process answers device_ready() — a method of ProcessExt, true by default — for the configuration it holds, and a false never fails: on any backend the process then samples on the host through its own sampler, bit-identically to the Cpu build. try_sample is Ok on that path, and sample never panics for a configuration — only a device that fails to open, compile or launch does that. So gbm.on::<Cuda>() always runs on the GPU, while wishart.on::<Cuda>() runs there for d = 2 and on the host for d = 3.

Which it is, and why, the process will say. device_fallback() returns the reason as Some("a Wishart outside the two-dimensional cone the kernel steps") and None when a kernel carries the configuration whole; device_ready() is the absence of a reason. The reason matters because the fallback is silent otherwise, and silence costs an order of magnitude:

if let Some(why) = wishart.device_fallback() {
  eprintln!("sampling on the host: {why}");
}

Both are on ProcessExt, and both are available from Python as process.device_fallback() and process.device_ready().

A process will also say where it is pointed. process.backend() hands back the handle it holds — the ordinal and batch budget a type parameter does not record — and process.probe() opens the device behind it and describes it, or reports why it cannot be used.

Choosing the device

A backend is a value, not only a type. .on::<B>() re-types the process with B's default handle; with_backend(handle) takes an explicit one:

use stochastic_rs::stochastic::device::Cuda;

let on_second_gpu = gbm.with_backend(Cuda::new(1).with_batch_budget(256 << 20));

Cuda { ordinal, batch_budget } and Metal { ordinal, batch_budget } carry the device ordinal (CUDA ordinal, index into Device::all()) and the bytes of path data one launch may hold. Their Default reads STOCHASTIC_RS_DEVICE and STOCHASTIC_RS_DEVICE_BATCH_BYTES (0 and 1 GiB when unset), so an unchanged program picks its device from the environment, and there is no process-wide state: two processes on two GPUs are two handles. Cpu and Accelerate carry nothing. probe() on a handle reports the device it opens.

Probing a device and handling its failures

A missing GPU, a runtime that will not start or a kernel that will not compile is an environmental failure, not a programming error, so it is reported as a value:

use stochastic_rs::stochastic::device::{Backend, DeviceError, Metal};

match Metal::default().probe() {
  Ok(info) => println!("{} on {} ({:?})", info.backend, info.name, info.precisions),
  Err(DeviceError::Unavailable(why)) => eprintln!("no device: {why}"),
  Err(other) => eprintln!("{other}"),
}

probe() opens the device behind the handle once and describes it (backend, name, precisions, ordinal). After a successful probe the plain sample() / sample_par(m) calls keep the ProcessExt signature and panic with the device's error if something still fails at launch time (an allocation, typically); every process also offers try_sample() / try_sample_par(m) -> Result<_, DeviceError> on ProcessExt (always Ok on the host devices, bit-identical to the plain calls) and the capability traits expose try_euler_paths / try_generate_batch on the handle for generic code. DeviceError has three kinds: Unavailable, Compile, Launch. On the CPU devices every probe is Ok and every try_* call is Ok with the same paths as its plain twin. From Python, a device failure is a RuntimeError with the same message.

Precision

Apple GPUs have no f64, so Metal is f32 only: it implements the capabilities for f32 alone, and an f64 process on it is a compile error. Cuda, Accelerate and Cpu follow the process's T parameter in f32 and f64.

When is GPU worth it?

The discrete-GPU win shows up for large n (≥ 16 k) and batched sampling, while for small n the CPU SIMD path wins on launch latency; on Apple Silicon Metal leads single paths from n = 16 k and the GPU back-ends win batches by a modest margin since the device state became reusable across calls. See the Benchmarks page for the cross-over numbers, and the Feature flags page for how to enable each backend.

On this page