Installation (Rust)
Add stochastic-rs to your Rust project — umbrella crate or per-sub-crate, with the right Cargo features and CPU / SIMD / GPU options.
Installation (Rust)
The Rust crates require Rust 1.89 or newer.
Umbrella crate (everything)
[dependencies]
stochastic-rs = "3.0.0-rc.1"Then:
// docs: getting-started/installation-rust#umbrella-crate-everything
//! Backs the umbrella-crate import example on the Rust installation page.
use stochastic_rs::prelude::*;
use stochastic_rs::quant::pricing::heston::HestonPricer;
use stochastic_rs::simd_rng::Deterministic;
use stochastic_rs::stochastic::diffusion::gbm::Gbm;
#[test]
fn umbrella_imports_resolve() {
let p = Gbm::<f64, _>::new(0.05, 0.2, 32, Some(100.0), Some(1.0), Deterministic::new(1));
let path = p.sample();
assert_eq!(path.len(), 32);
fn _type_check(_: &HestonPricer) {}
}The umbrella re-exports everything via pub use from the sub-crates, so
existing v1.x import paths keep working.
Per-sub-crate (lean)
For minimal compile time and dependency surface, depend only on the sub-crates you need:
[dependencies]
stochastic-rs-distributions = "3.0.0-rc.1" # SIMD distribution sampling
stochastic-rs-stochastic = "3.0.0-rc.1" # 131 process types
stochastic-rs-copulas = "3.0.0-rc.1" # bivariate / multivariate copulas
stochastic-rs-stats = "3.0.0-rc.1" # estimators
stochastic-rs-quant = "3.0.0-rc.1" # pricing / calibration / vol surface
stochastic-rs-ai = "3.0.0-rc.1" # neural surrogates (candle)Topology:
stochastic-rs-core (simd_rng)
└→ stochastic-rs-distributions (FloatExt, SimdFloatExt, distributions)
├→ stochastic-rs-stochastic (ProcessExt + 131 processes)
├→ stochastic-rs-copulas
└→ stochastic-rs-stats
└→ stochastic-rs-quant (ModelPricer, calibration, vol surface)
└→ stochastic-rs-aiCargo features
| Feature | Owner crate | Pulls in | Use when |
|---|---|---|---|
ai | umbrella | stochastic-rs-ai, candle-core | NN volatility surrogates |
cuda | stochastic | cudarc, cuFFT, fused Philox | Direct CUDA backend for FGN / fBM (NVIDIA, CUDA 12.x) |
metal | stochastic | metal (Apple framework) | Direct Metal backend for FGN / fBM on macOS |
accelerate | stochastic | Apple Accelerate (vDSP) | macOS-native FFT acceleration (no toolchain install) |
mimalloc / jemalloc | umbrella | mimalloc / tikv-jemallocator | Drop-in allocator for long-running MC workloads |
python | umbrella + stochastic-rs-py | pyo3, numpy | Building the Python wheel via maturin |
Default build (cargo build) is feature-light and links no GPU, no BLAS,
and no Python. Pick features explicitly for the workload at hand.
SIMD support
Numerical hot paths (FGN Davies-Harte, all Simd* distributions,
fill_slice / fill_slice_fast) use the wide
crate for portable SIMD. Lane widths in this codebase are uniformly
8-lane types:
f32x8— 8 ×f32= 256 bits (AVX2 / NEON-pair)f64x8— 8 ×f64= 512 bits (AVX-512, or 2 × AVX2 / 2 × NEON fallback)i32x8— for the integer Box-Muller / ziggurat tables
wide selects the actual SIMD instructions at build time based on
the active target features. The default x86-64 toolchain targets only
SSE4.2, which means f32x8 / f64x8 compile to scalar loops. To
unlock real SIMD, opt into a higher CPU baseline (next subsection).
| Target arch | Default ISA | What wide emits without extra flags |
|---|---|---|
x86_64-… (Linux, MSVC) | SSE4.2 (v1) | Scalar fallback (no AVX) |
x86_64-… with +avx2 | AVX2 | Full 256-bit SIMD on f32x8 |
x86_64-… with +avx512f | AVX-512 | Full 512-bit SIMD on f64x8 |
aarch64-apple-darwin | NEON | 128-bit NEON, two-pump for 256-bit ops |
Native CPU optimization
Default builds target the plain x86-64 / aarch64 baseline so the
resulting binary or wheel runs on any CPU of the same architecture.
For SIMD-heavy paths (SimdNormal::fill_slice_fast, Fgn Davies-Harte,
sample_par, …) the gap between v1 and a tuned target is large
enough to be worth raising the floor:
# Local dev / benchmarks: every feature the build host supports.
# The resulting binary only runs on this exact CPU family.
RUSTFLAGS="-C target-cpu=native" cargo build --release
RUSTFLAGS="-C target-cpu=native" cargo bench
# Higher x86-64 baselines (binary runs on any CPU meeting the level):
# v2 — SSE4.2 + POPCNT (x86_64 CPUs since ~2009)
# v3 — AVX2 + BMI2 + FMA (x86_64 CPUs since ~2013–2015)
# v4 — AVX-512 (most server CPUs only; absent from all
# AMD Zen 1–3 and Intel client 12th-gen+)
RUSTFLAGS="-C target-cpu=x86-64-v3" cargo build --releasePublic distribution (PyPI wheels, openly shared Docker images): keep
the default x86-64 baseline. pip wheel tags don't dispatch by CPU
feature level, so a v3 wheel will SIGILL on any pre-2013 hardware
(AMD Bulldozer/Piledriver, Sandy/Ivy Bridge, Atom variants). Use
v2/v3/v4 only for deployments where you've verified every target
host clears the level — typical examples are an internal Docker fleet,
a homogeneous HPC cluster, or a CI runner pinned to a known SKU. Use
target-cpu=native only for local dev and benchmarks.
RUSTFLAGS busts the build cache. Every distinct value triggers a
full workspace rebuild, and the env var fully replaces (does not
merge with) [build] rustflags = […] in any .cargo/config.toml. For
persistent local optimisation, prefer a [target.<triple>] rustflags
entry in ~/.cargo/config.toml so it composes with project configs.
GPU support
FGN and fBM ship four independent GPU / accelerator backends — pick the one that matches your hardware and toolchain. Every Rust block in this section needs hardware or an OS this site's CI runners (Linux, no GPU) do not have — an NVIDIA GPU + CUDA 12.x toolkit, a WebGPU-capable adapter, or macOS — so all of them are illustrative rather than compiled.
The .on::<B>() selector (recommended)
Select a backend at compile time with .on::<B>(), then call sample() /
sample_par() as usual — the dispatch is monomorphised, with no runtime branch:
use stochastic_rs::simd_rng::Unseeded;
use stochastic_rs::stochastic::device::Cuda;
use stochastic_rs::stochastic::noise::fgn::Fgn;
use stochastic_rs::traits::ProcessExt;
let path = Fgn::<f32, _>::new(0.7, 65_536, None, Unseeded)
.on::<Cuda>() // Cuda | Metal | Accelerate | Cpu
.sample();The markers only exist when their feature is compiled, so an unavailable backend
is a compile error rather than a runtime fallback. See the
Backends concept page. .on::<B>() is the whole
public surface — the per-backend samplers are crate-private (sample_*_impl,
wired up in stochastic_rs::stochastic::device) and are reached only through
it, as in the per-backend blocks below.
cuda — direct CUDA (NVIDIA, recommended)
Direct binding via cudarc + cuFFT
- a fused Philox RNG kernel. No
.cufiles, nonvccrequired — the kernels ship as Rust strings and JIT through cudarc.
Requires NVIDIA CUDA Toolkit 12.x and a compatible GPU.
cargo build --features cuda
cargo bench --features cuda --bench fgn_cudause stochastic_rs::simd_rng::Unseeded;
use stochastic_rs::stochastic::device::Cuda;
use stochastic_rs::stochastic::noise::fgn::Fgn;
use stochastic_rs::traits::ProcessExt;
let path = Fgn::<f32, _>::new(0.7, 65_536, None, Unseeded)
.on::<Cuda>()
.sample(); // single path on GPU
let batch = Fgn::<f32, _>::new(0.7, 65_536, None, Unseeded)
.on::<Cuda>()
.sample_par(1024); // 1024 paths in one launchmetal — direct Metal (macOS)
Direct binding via the metal crate.
Targets Apple Silicon (M1/M2/M3/M4) and Intel Macs with discrete /
integrated GPUs.
cargo build --features metal
cargo bench --features metal --bench fgn_metaluse stochastic_rs::stochastic::device::Metal;
let path = Fgn::<f32, _>::new(0.7, 65_536, None, Unseeded)
.on::<Metal>()
.sample();accelerate — Apple Accelerate (macOS, no GPU)
Routes the FFT through Apple's vDSP (part of the Accelerate framework
shipped with macOS — no extra install). Lower latency than the GPU
paths for medium-n workloads where launch overhead dominates.
cargo build --features accelerate
cargo bench --features accelerate --bench fgn_accelerateuse stochastic_rs::stochastic::device::Accelerate;
let path = Fgn::<f32, _>::new(0.7, 65_536, None, Unseeded)
.on::<Accelerate>()
.sample();Choosing a backend
| Backend | Best for | Latency floor | Throughput ceiling |
|---|---|---|---|
| CPU SIMD | Small n (≤ 4 k), single path | ≈ 8 µs | rayon × cores |
accelerate | Medium n (4 k–16 k), single path on macOS | ≈ 30 µs | one core |
metal | Large n + batches on macOS | ≈ 80 µs | full GPU |
cuda | Large n (≥ 16 k) + batches | ≈ 80 µs | full GPU |
Concrete numbers and the cross-over point are on the Benchmarks page.
Linear algebra
Cholesky, SVD, eigendecompositions and least squares run on the
pure-Rust faer and are
always compiled in — there is no BLAS/LAPACK system dependency to
install on any platform.
Verify the install
// docs: getting-started/installation-rust#verify-the-install
//! Backs the "verify the install" example on the Rust installation page.
use stochastic_rs::prelude::*;
use stochastic_rs::simd_rng::Unseeded;
use stochastic_rs::stochastic::diffusion::ou::Ou;
#[test]
fn ou_path_has_the_requested_length() {
let p = Ou::<f64, _>::new(2.0, 0.0, 1.0, 1_000, Some(0.0), Some(1.0), Unseeded);
let path = p.sample();
assert_eq!(path.len(), 1_000, "OU path of length {}", path.len());
}Paste the equivalent into your own src/main.rs to check the install by
eye:
use stochastic_rs::prelude::*;
use stochastic_rs::simd_rng::Unseeded;
use stochastic_rs::stochastic::diffusion::ou::Ou;
fn main() {
let p = Ou::<f64, _>::new(2.0, 0.0, 1.0, 1_000, Some(0.0), Some(1.0), Unseeded);
let path = p.sample();
println!("OU path of length {}", path.len());
}Run with cargo run --release. If this prints OU path of length 1000,
you are good. Continue with the
Quickstart.
stochastic-rs
Simulate 131 stochastic processes, price and calibrate against them, and move any of it to a GPU by naming one — in Rust or in Python.
Installation (Python)
Install the stochastic-rs Python bindings — pre-built wheels via pip, or build locally with maturin and Bun-equivalent uv-pip workflow.