stochastic-rs
Concepts

ProcessExt

The ProcessExt contract — sample, sample_map, sample_par — covering output shape, time grid, buffer reuse, and parallel determinism.

ProcessExt<T>

Every stochastic process in stochastic-rs-stochastic — diffusion, jump, volatility, interest, rough, noise — implements ProcessExt<T>. The public surface is three methods:

pub trait ProcessExt<T: FloatExt>: Send + Sync {
  type Output: Send;

  /// One sampled path.
  fn sample(&self) -> Self::Output;

  /// Map `f` over `m` independently sampled paths, in parallel.
  fn sample_map<R: Send>(&self, m: usize, f: impl Fn(&Self::Output) -> R + Sync) -> Vec<R>;

  /// `m` independently sampled paths, kept.
  fn sample_par(&self, m: usize) -> Vec<Self::Output>;
}

Under these sits a reusable per-thread sampler that owns the mutable sampling state (RNG, distribution buffers, precomputed scales). It is an implementation detail — you never name or construct it; the three methods above are the whole interface.

Output shape

Self::Output depends on the process family:

FamilyOutput
Single-factor (OU, GBM, CIR, Vasicek, …)Array1<T>
Multi-state (Heston, SABR, Bergomi, …)[Array1<T>; N]
Term structure / sheet (HJM tenors, Bgm, Fbs)Array2<T>
Variable-dimension (multivariate Hawkes)Vec<Array1<T>>
Complex-state (cfOU)Array1<Complex<T>>

The dimensional marker traits (OneDimensional, TwoDimensional, MultiDimensional<N>, CurveOutput, VariableDimensional, ComplexPathOutput) are auto-implemented from Output, so generic code can bound on the shape it expects.

Path length and time grid

Construction of any process takes (n, x0, t):

  • n: usize — number of steps (a single path is length n)
  • x0: Option<T> — initial value; None ⇒ a sensible default per process (e.g. 0 for OU, 1 for GBM)
  • t: Option<T> — horizon; None ⇒ 1.0

The time grid is uniform: Δt=t/n\Delta t = t / n. Non-uniform grids are handled by user code (sample at fine Δt\Delta t, then resample).

Choosing a method

tests/doctest_concepts_process_ext_choosing_method.rs
// docs: concepts/process-ext#choosing-a-method
//! Backs the "choosing a method" example on the ProcessExt concept page.

use stochastic_rs::simd_rng::Unseeded;
use stochastic_rs::stochastic::diffusion::gbm::Gbm;
use stochastic_rs::traits::ProcessExt;

#[test]
fn sample_sample_map_and_sample_par() {
  let gbm = Gbm::<f64, _>::new(0.05, 0.2, 64, Some(1.0), Some(1.0), Unseeded);

  // one path
  let path = gbm.sample();
  assert_eq!(path.len(), 64);

  // parallel reduction — no per-path allocation kept
  let mean = gbm
    .sample_map(10_000, |p| (p.last().unwrap() - 1.0).max(0.0))
    .iter()
    .sum::<f64>()
    / 10_000.0;
  assert!(mean >= 0.0);

  // kept paths, for plotting / path-dependent post-processing
  let paths = gbm.sample_par(1_000);
  assert_eq!(paths.len(), 1_000);
}
MethodReturnsUse when
sample()one Outputa single path, or you drive your own loop.
sample_map(m,f)Vec<R>a parallel Monte-Carlo reduction — you only need a function of each path (payoff, terminal value, a Greek). Reuses one sampler and one output buffer per worker; nothing is materialised. This is the fast path.
sample_par(m)Vec<Output>you need to keep all m paths (storage, plotting, path-dependent processing). Allocates each path fresh.

Both parallel methods build one sampler per Rayon worker, not one per path, which removes the per-path RNG-setup and (for sample_map) the per-path allocation. Reach for sample_map whenever the answer is a number, not the paths themselves.

Performance

The reuse matters most in the many-short-paths regime (large m, small n), where the fixed per-path costs are a large fraction of the work. On Apple Silicon, GBM, a parallel reduction over ≈2.6×105\approx 2.6 \times 10^5 path-elements (cargo bench --bench sampler_compare):

nold idiom (parallel sample() fold)sample_mapspeedup
64364 µs189 µs1.93×
256188 µs173 µs1.09×
1024165 µs159 µs1.04×

The dominant lever is the auto-seed counter becoming contention-free (see seeding); buffer reuse and uninitialised allocation add the rest. For longer paths (n≥256n \geq 256) the path computation itself dominates, so the speedup tapers to a few percent. Serial and single-path sampling are within a few percent of the pre-2.x figures — the win is specifically parallel and short-path.

Parallel sampling and determinism

Each Rayon worker derives its own RNG from the process's seed source. The consequences depend on the seed strategy:

  • Deterministic + sample() (serial) — bit-for-bit reproducible. Repeated calls advance the seed, so successive paths differ. A serial loop of sample() is the way to get reproducible Monte-Carlo runs.
  • Deterministic + sample_par / sample_map (parallel) — also bit-for-bit reproducible: the same seed and the same m give bit-identical output on any machine and under any Rayon thread-pool size. Each chunk's sampler is built sequentially, before any work reaches Rayon, so scheduling cannot decide which basis lands on which path. No process is an exception, full or partial, and stochastic-rs-stochastic/tests/reproducibility_all_processes.rs enumerates every concrete implementor to enforce it. The one caveat is a backend, not a process: Fgn / Fbm on the accelerate backend get thread-count-independent seed consumption but not bit-identical output, because vDSP's arithmetic is not bit-stable across Apple Silicon's P-core / E-core scheduler.
  • Unseeded — every path is independently auto-seeded in either mode.

In all modes the seeds handed out are globally unique, so paths never collide or "get stuck" on a repeated stream — including across the thread-local seed-block boundary (validated in tests/sampler_v3_rng.rs).

FGN ships a sample_pair fast path that produces two independent paths in one FFT — a 2× shortcut over calling sample() twice while maintaining exact independence.

Acceleration

Default implementation: CPU SIMD via f64x4 / f32x8 where the sampler allows. GPU backends (cuda, metal) are opt-in features and currently ship for FGN / fBM only — see the add-gpu-sampler SKILL.

Construction patterns

Every process follows the same canonical new(args…, seed) constructor — model parameters first, seed source last (see the seeding concept page for the full design). Gbm stands in for any process below; substitute the parameters of the one you want:

use stochastic_rs::simd_rng::Deterministic;
use stochastic_rs::simd_rng::Unseeded;
use stochastic_rs::stochastic::diffusion::gbm::Gbm;

// auto-seeded
let p = Gbm::<f64, _>::new(0.05, 0.2, 1_000, Some(100.0), Some(1.0), Unseeded);

// reproducible
let p = Gbm::<f64, _>::new(
  0.05,
  0.2,
  1_000,
  Some(100.0),
  Some(1.0),
  Deterministic::new(42),
);

// chain seeds: one shared source feeding several processes
let shared = Deterministic::new(42);
let p = Gbm::<f64, _>::new(0.05, 0.2, 1_000, Some(100.0), Some(1.0), shared.clone());

Use Deterministic in tests so they replay bit-exactly. Pass a shared seed source when chaining correlated processes — each seed.rng() / seed.rng_ext() call atomically advances the internal counter, so the streams diverge despite the shared root. SeedExt::reseed(u64) swaps a Deterministic source in place to sweep seeds without rebuilding the process.

On this page