Rust · Python · CUDA · Metal

stochastic-rs

stochastic-rs is an open-source quantitative-finance library for Rust and Python. Simulate 132 stochastic processes, price and calibrate against them, and move any of it to a GPU by naming one. The model does not change, the numbers keep their law, and the crate tells you when a device cannot take a configuration instead of quietly running it on the host.

cargo add stochastic-rs · pip install stochastic-rs

132
processes behind one trait
17×
the CPU on a batch of 200 000 paths
3 000+
tests: laws, devices, reproducibility
308
entries in the Python module

One thread, end to end

Four steps, each one real code that runs. The same four are shown in Python in the quickstart.

  1. 01

    Sample a batch

    Every process is built the same way — parameters, grid, horizon, seed — and every one gives you a single path, a parallel batch, and a mapped form over it.

    let gbm = Gbm::<f32, _>::new(
      0.05, 0.2, 1_024, Some(100.0), Some(1.0), Deterministic::new(7),
    );
    let paths = gbm.sample_par(1_000);
    
  2. 02

    Move it to a GPU

    Naming a backend is the only change, and the mapped form reads the batch where the kernel wrote it rather than copying it out — which at two hundred thousand paths is 17× the same call on the CPU.

    let terminal: Vec<f32> = gbm
      .on::<Metal>()
      .sample_map_view(200_000, |p| p[p.len() - 1]);
    
  3. 03

    Price the same model in closed form

    Simulators and pricers are separate: a pricer holds the parameters and takes the query as arguments, so a whole strike-maturity grid is one vectorised sweep.

    let model = HestonPricer::new(0.04, -0.7, 2.0, 0.04, 0.3, None);
    let call = model.price_call(100.0, 100.0, 0.05, 0.02, 0.75);
    // 7.7311
    
  4. 04

    Read a parameter back out

    The estimators take a path and return the number that generated it, with the diagnostics to judge how much to believe it.

    let fbm = Fbm::<f64, _>::new(0.7, 4_096, Some(1.0), Deterministic::new(11)).sample();
    let h = RescaledRange::default().estimate(fbm.view())?;
    // 0.61, for a true 0.7
    

What comes out

Hestonκ 2 · ξ 0.3 · ρ −0.7
60801001201400.000.250.500.751.0095%5%tS
Fractional Brownian motionone seed, three H
-3-2-100.000.250.500.751.00H 0.3H 0.5H 0.8tB
SVI surfacefour maturities
0.200.300.400.500.60-0.50-0.250.000.250.503M6M1Y2Ylog-moneynessσ

What are you here for?

To simulate

132 processes behind one trait — diffusions, jumps, stochastic volatility, short rates, fractional and rough, point processes, subordinators, conditional-variance time series. Heston and its rough relatives, Bates, SABR, CGMY, Hull-White, LMM, fBm and everything built on it.

To price or calibrate

Analytic, Fourier, ADI and Monte Carlo pricers with first- and second-order Greeks; 15 calibrators behind one trait, SVI and SSVI surfaces, curves, bonds, credit and risk.

To estimate

Eight Hurst estimators, realised measures, jump tests, cointegration, changepoints, extreme value, GARCH fitting — each carrying the paper it came from.

To draw

36 SIMD distribution samplers, most with characteristic function, pdf, cdf and moments in closed form, and 23 copulas with fitting and goodness-of-fit.

To do it from Python

308 entries, numpy in and numpy out, and the same device argument. Wheels for every platform, no feature flags to choose.

Why this one

A declaration, not a kernel

A family is written once in a small DSL — its state, its noise, its step — and the same text is rendered as CUDA C for NVRTC and as MSL for Metal. Adding a process means adding a declaration, not writing and debugging two kernels. 121 families carry 130 of the processes; the two that need pipelines of their own — fractional Gaussian noise and the fractional Brownian sheet — have them.

The speed is in the shape of the call

Ask for the grid and you get the grid; ask for a borrowed view and nothing is copied; ask for a fold and the kernel never writes it. No flags, no tuning — the same program, told what you actually want back.

Measured against mathematics

Beyond the host-versus-device suite there is one that states each model’s closed form, so a sampler wrong on both sides still fails. It earns its keep: a ziggurat tail folded inward, a circulant embedding that lost its long memory in f32, an inverse-Gaussian draw that cancelled to zero on a fine grid.

The numbers behind all three, the method that produced them, and what did not work are in Benchmarks.

Simulate a path, price an option, estimate a Hurst exponent

The quickstart does all three, in Rust and in Python side by side.