stochastic-rs
Concepts

Design philosophy

Why the library is shaped the way it is — generic over float, no statrs, paper-anchored implementations, comparison-test mandatory, plus non-goals.

What follows is the short version of the development rules at .claude/skills/dev-rules/SKILL.md. If you contribute to the library, read that SKILL once. If you only use the library, this page is enough.

1. Generic over float

Every numerical type is generic over T: FloatExt. We support f32 and f64. Reasons: GPU memory bandwidth (f32), embedded targets, and the option for higher-precision drop-ins later.

This means the library never hard-codes f64. If you contribute, your new pricer / process / estimator must also be generic. The only exception is glue code that inherently lives in one precision (e.g. numpy interop with dtype=float64).

2. Paper-anchored implementations

Every model and algorithm carries a paper citation in the source file — this is the convention itself (a module-level doc comment), not a runnable snippet:

//! Reference: Heston (1993). DOI: 10.1093/rfs/6.2.327

The implementation follows the paper exactly — no simplifications, no "I think this is equivalent" rewrites. If a formula in the paper looks numerically unstable, the fix is to cite the correction paper that addresses the instability, not to silently rewrite.

Comparison tests anchor the implementation to numerical tables in the paper.

3. No statrs for distribution math

Distribution closed forms are written from scratch in stochastic-rs-distributions, never delegated to statrs::distribution::*. The reasoning is on the DistributionExt page.

4. ndarray everywhere, never Vec<T>

All numerical arrays are ndarray::Array1<T> / Array2<T>. The library already depends on ndarray and ndrustfft. Don't re-introduce Vec<T> for numerical data — it complicates SIMD and bulk-fill paths.

5. Comparison test + criterion bench mandatory

A new module ships with both:

  • A comparison test under tests/ that validates output against the reference implementation (Python, R, MATLAB, or paper tables).
  • A criterion benchmark under benches/ that tracks performance.

This rule is enforced by review, not CI — there is no automated check that a new module shipped tests + benches.

6. Three-way error convention

One rule for every crate, from the distributions up to the calibrators:

  • Invalid parameter → panic, naming the parameter and its value. A negative volatility, a Hurst exponent outside [0, 1], a theta outside a copula's admissible interval, a grid with fewer than two points: programmer errors, not runtime states, so constructors and setters assert! them with a message of the form sigma must satisfy `sigma > 0`, got sigma = -0.2. Never a silent clamp, a silent zero or a silent None.
  • Not computable at this point → NaN, documented. A price outside a model's domain, a Greek at an expired option, an implied volatility with no root: legitimate states, returned as f64::NAN from methods whose docs say when. A function that semantically must compute a number never returns 0 in its place; unimplemented closed forms unimplemented! with the type and method named.
  • Data-dependent failure → Result::Err. The optimiser did not converge, a correlation matrix estimated from data is not positive-definite, a fit rejected its input, a vine could not be built from the supplied pairs: the caller needs a graceful path, so Calibrator::calibrate, the copula fit/new_with_corr/vine constructors and the estimators return Result. Calibrator::Error is a free associated type; the in-tree calibrators set it to anyhow::Error.
  • Environmental failure on a device → probe first, then Result or a panic that says so. No GPU behind a marker, a runtime that will not start, a kernel that will not compile, an allocation that fails: handle.probe() reports it as a DeviceError before any sampling, and the try_sample_par / try_* calls return the same error at sampling time. The plain sample* calls keep the ProcessExt signature and panic with that error's message, so a device failure is never a silent host fallback and never a bare unwrap.

7. What is stable and what may grow

Enums that catalogue an open set — copula families, day-count and holiday conventions, interpolation and finite-difference schemes, loss metrics, error kinds — are #[non_exhaustive], so a new family or scheme is not a breaking change; match them with a wildcard arm. Enums that are closed by definition (OptionType, OptionStyle, Moneyness, the payer/receiver and up/down × in/out kinds) stay exhaustive. Processes and pricers with public fields are constructed through new and the with_* setters; a struct literal is not a supported construction (a public-field process also carries its backend handle, backend: Cpu), so adding a field is not a breaking change for code that follows the constructors.

Explicit non-goals

  • Not a general-purpose Monte Carlo engine. We have variance-reduction techniques (antithetic, control variate, stratified, importance, quasi-MC, MLMC), but the design priority is quant-finance MC, not general PDE / SDE solving. (For the latter, look at nalgebra-glm, pyo3-numpy, or research-grade libraries like dolfin.)

  • Not a scipy.stats clone. We ship the distributions and estimators we need for the workflows the library is designed for. Asks like "please add a Pearson-VII distribution" without a clear quant-finance use case will be politely declined.

  • Not a closed-source-quant replacement. The vol-surface pipeline, the calibrators, and the risk metrics are competitive but specifically scoped: liquid-equity vanilla / lightly-exotic. Production deployments in fixed-income / credit / OTC commodities will likely need their own layer on top.

Edit on GitHub

Last updated on

On this page