Fair, reproducible NumSharp vs NumPy 2.4.2 performance comparison: an op × dtype × N matrix, the complete official LinearAlgebra suite under both Managed C# and OpenBLAS, targeted backend supplements for routes outside that matrix, and five complementary subsystems that fill axes the matrix cannot express.
This file is an orientation guide, not the results. For the numbers, see the canonical report
benchmark-report.md(regenerated by CI after each release) or the committed snapshot athistory/latest/benchmark-report.md. The human-facing UI is the website's Benchmarks dashboard (docs/website-src/docs/benchmarks-dashboard.md).
ratio = NumPy_ms ÷ NumSharp_ms.
>1= NumSharp faster,<1= slower,=1= parity. Higher is better.%NumPy🕐= NumSharp_ms ÷ NumPy_ms × 100 = the share of NumPy's time NumSharp uses (30% = NumSharp takes only 30% as long; <100% = faster).
Bands used by the report: ✅ ≥1.05× · 🟡 0.8–<1.05× · 🟠 0.33–<0.8× · 🔴 <0.33× · ▫ negligible
(semantic O(1)-in-N scenario, sub-µs, or >20× — excluded from every rollup/ranking) · ⚪ C# side not run.
Timing basis: every ratio compares the two sides' per-case min (best window). Min-based harnesses
run each case to a ~200 ms time budget (a >20 ms/call op runs exactly 100 times); the C# op-matrix keeps
BenchmarkDotNet's 50 iterations — never the mean, whose NumSharp-side GC/page-fault/machine-load
tails turned host state into fake regressions; per-case means stay in the report JSON
(numpy_mean_ms/numsharp_mean_ms) as tail diagnostics.
⚠️ The legacyrun-benchmarks.ps1prints the inverse (NS/NPY, lower is better). Everything else — the op-matrix report, the dashboards, every*_sheet.py— is NPY/NS. Prefer NPY/NS.
File-based dotnet run file.cs / dotnet_run compile both the script AND any #:project
NumSharp.Core in Debug, which the JIT honors even over [MethodImpl(AggressiveOptimization)].
Hand-written C# hot loops run ~2× slow; IL-emitted kernels look normal. Every timing script MUST run
dotnet run -c Release - < script.cs. The BenchmarkDotNet projects are exempt (they mandate -c Release). Full detail in CLAUDE.md.
python run_benchmark.py # interactive depth + dtype picker
python run_benchmark.py --depth measure # full official run (18 suites + backend profiles + subsystems)
python run_benchmark.py --suites arithmetic unary
python run_benchmark.py --skip-build # reuse the existing Release build
python run_benchmark.py --skip-csharp # NumPy only
python run_benchmark.py --quick # deprecated alias for --depth lightExecution depth and dtype targeting:
python run_benchmark.py --depth pass # exactly 1 BDN/NumPy call per cell; 0 warmups
python run_benchmark.py --depth light --dtypes f32,c128 # rough estimate; comma-separated dtype filter
python run_benchmark.py --depth measure # full publication profile (default with other flags)With no CLI arguments, an interactive terminal shows a depth options picker followed by a dtype
picker. Multi-dtype scenarios match a requested dtype on any side. Pass/light are validation/local
estimate profiles: they write only to results/<timestamp>/, skip complementary subsystem sheets,
and never replace canonical root, docs, or history artifacts.
run_benchmark.py is the single entry point. It builds the C# suite, runs each suite through
BenchmarkDotNet (OfficialBenchmarkConfig: InProcessEmit toolchain, 50 measured iterations, iteration
time capped at 25 ms), and reruns the same discovered LinearAlgebra classes/configuration in the
OpenBLAS executable. It sweeps NumPy across 1K / 100K / 10M in one process per suite, merges on
(op, dtype, N, scenario), records both backend profiles, selects the fastest valid timing per exact
cell, and appends the subsystem sheets. Only a successful full measure run mirrors the headline
artifacts to this folder/docs and writes the committable history/<date>_<sha>/ snapshot (+ repoints
history/latest). Subsets, targeted reruns, pass/light runs, and incomplete runs stay in their run directory.
Starting python run_benchmark.py in a terminal first checks for an unfinished run
and asks Resume this benchmark? [Y/n], showing its depth, dtypes, and last saved
progress. Enter resumes with the saved options; n opens the new-run picker.
Completed runs, saved plans, active runs, and incompatible source/environment records
are skipped. Custom run directories are remembered in results/last-run.json.
Explicit CLI arguments never trigger this prompt; non-interactive callers must pass
options such as --depth measure or --resume latest.
Run these commands from benchmark/:
python run_benchmark.py --depth measure --run-dir results/nightly --plan-only
python run_benchmark.py --resume results/nightly
python run_benchmark.py --resume latest--plan-only builds and discovers the exact selected cases, saves their identities, and stops before
timing workloads. Omit it to start measuring immediately. The total comes from live discovery,
including every selected dtype, size, and engine: NumPy, Managed C#, and the OpenBLAS profile count
separately. Do not infer the total from a scenario-title count or an older comparison report.
The parent runner keeps terminal progress and the window/tab title across suite and engine changes: completed/total matrix cases, remaining cases, failures, an estimated remaining time, and the current stage. The ETA learns from observed case durations and is approximate. Subsystems have a separate pending-stage count; their internal cells and durations are not included in the matrix ETA.
Both NumPy and BenchmarkDotNet save each completed matrix case atomically. Interrupting with Ctrl+C
keeps completed work; --resume reuses successful cases and retries failed or unfinished ones. Resume
also skips completed subsystem stages; an interrupted subsystem restarts at its stage boundary.
The saved selection is fixed. Changed source code, compiled benchmark binaries, runtimes, or recorded environment settings block
resume to avoid mixing measurements; start a new run with --rerun after making code changes.
Useful controls:
python run_benchmark.py --resume results/nightly --attempts 3 --progress-interval 10
python run_benchmark.py --resume results/nightly --process-timeout 3600 --verbose --no-title
python run_benchmark.py --depth light --operations "np.add*" --run-dir results/add-check--attempts bounds attempts per incomplete matrix suite (default 3), retaining successful cases
between attempts. --process-timeout optionally limits each child process in seconds, including
build/discovery processes; expiry stops its process tree. --verbose echoes child output as well as
logging it, and --no-title keeps progress in the terminal only.
python run_benchmark.py --rerun failed --from-results results/nightly --run-dir results/retry
python run_benchmark.py --rerun bad --from-results history/latest/benchmark-report.json --plan-only
python run_benchmark.py --rerun degraded --from-results results/current --baseline results/baseline --regression-percent 10failedselects actual recorded failures, including failed NumPy checkpoints in a run directory.badadds unmeasured/pending cases and credibleslower/much_slowercomparisons. Knownmissing_backend/not_supportedavailability states and negligible timings are excluded.degradedselects cases whose NumSharp time increased by more than 10% by default against the same operation/dtype/size/scenario and backend profile in the baseline. It compares absolute NumSharp best-window times, not changes in the NumPy/NumSharp ratio or the fastest effective backend.
Selection matches the current discovered plan exactly and includes NumPy counterparts for selected
NumSharp cells. selection.json records reasons and unmatched cases, including historical or subsystem
scenarios the official matrix cannot reproduce. Known incompatible baseline metadata is rejected;
missing historical provenance is reported. Inspect the plan before spending hours on a rerun.
Each run keeps:
| Path inside the run directory | Purpose |
|---|---|
run-config.json, plan.json |
Fixed options, source/environment provenance, exact execution plan |
build-identity.json, numpy-reference.json |
Executed binary hashes and pinned input when reusing NumPy timings |
run-state.json, events.jsonl |
Latest progress/stage state and completion events |
numpy-checkpoints/ |
Atomic NumPy case results or failure records |
csharp/, csharp-openblas/ |
Atomic BDN case reports with statistics and measurements |
logs/, selection.json |
Full child output and targeted-rerun diagnostics |
numpy-results.json, benchmark-report.* |
Rebuilt aggregate timings and comparison reports |
Older snapshots remain useful selection sources, but cannot be resumed: they omit raw BDN case
measurements and have no durable plan/checkpoints. For example, the pre-checkpoint local run
results/20260906-230207/ saved 4,996 NumPy rows and only 990 C# rows in five copied class reports,
with no final report. Other September 6 runs contain 21,720 C# rows: a fixed “5,000 tests” estimate
would miss most engine/case work. Preserve these old artifacts as evidence; use new-format runs for recovery.
| Path | What |
|---|---|
run_benchmark.py |
THE orchestrator (op matrix + backend profiles + subsystems + history snapshot). |
NumSharp.Benchmark.CSharp/ |
Core-only BenchmarkDotNet project and shared benchmark assembly — Benchmarks/<Category>/*.cs, config in Infrastructure/. |
NumSharp.Benchmark.CSharp.OpenBLAS/ |
OpenBLAS-enabled executable that discovers the shared LinearAlgebra classes and uses the same official config. |
NumSharp.Benchmark.Python/numpy_benchmark.py |
The NumPy twin — one run_<suite>_benchmarks(...) per suite. |
scripts/merge-results.py |
Joins C# + NumPy on (op, dtype, N) → benchmark-report.{md,json,csv} (NPY/NS + credibility gating). |
scripts/merge-backend-profiles.py |
Merges Managed/OpenBLAS profile JSON, preserves availability exceptions, and selects the effective timing. |
scripts/benchmark_session.py, scripts/benchmark_selection.py |
Durable run state/global progress and exact rerun selection from saved results. |
scripts/bench_common.py |
Shared build/run/parse driver for the file-based subsystems. |
scripts/snapshot_history.py |
Assembles history/<date>_<sha>/ + latest (the publish step). |
scripts/render_dashboard.py |
benchmark-report.json → benchmark-dashboard.md (ASCII sheet; run by hand — seeds the DocFX UI). |
scripts/audit_coverage.py + coverage/generated/ |
Machine-checkable API benchmark ledger (456/456 benchmarkable compatibility APIs). |
scenarios.html + scripts/generate_scenarios_html.py |
Executable 499-title × 15-dtype audit with Comparable, C# BDN, and Python coverage lenses. |
backends/ + openblas/ |
Targeted Managed/OpenBLAS runs using one JSON schema; catches MissingBackendException and NotSupportedException. |
nditer/ layout/ operand/ cast/ fusion/ |
Five appended complementary subsystems. |
history/ |
Tracked snapshots — what we commit and reference. latest → newest. |
results/ |
Gitignored raw per-run scratch. |
run-benchmarks.ps1 |
Legacy Windows-only runner (inverse NS/NPY convention). |
Regenerate the scenario/dtype audit from live BenchmarkDotNet attributes and the NumPy suite builders:
python benchmark/scripts/generate_scenarios_html.py
python benchmark/scripts/generate_scenarios_html.py --no-build --checkarithmetic · unary · reduction · broadcast · creation · manipulation · slicing · comparison · bitwise · logic · statistics · sorting · linalg · selection · fft · random · ndarray · api — each
a C# namespace filter in run_benchmark.py's SUITES map. The universal-tier contract is strict:
once an operation/dtype publishes both 1K and 100K it must also schedule 10M; the official merge
rejects asymmetric output independently for NumPy and C#. Scalar/1K-only API-dispatch cases remain
valid because they never enter the throughput matrix. For allocation-heavy families, the 10M tier
uses a symmetric 1M physical-work cap on NumPy and NumSharp while retaining N=10M as the sole
join/dashboard label; this is a workload mapping, not a fourth size tier. The experimental /
allocation classes (Dispatch, Fusion, DynamicEmission, SimdVsScalar, MultiDim, NumSharp, Allocation/)
have no NumPy twin and are not part of the official run.
Backend coverage has two layers. The official LinearAlgebra BenchmarkDotNet classes run first in the
Core-only executable and then unchanged in the OpenBLAS executable; every measured Managed cell must
have an available exact-cell OpenBLAS-profile peer or the merge fails. The targeted harness supplements
backend-only product/LAPACK routes that are not official op-matrix cells. Missing backend / unsupported
outcomes are availability states, not timings. Every profile publishes the same 1K / 100K / 10M
workload tiers. Operation-specific physical work stays bounded: LAPACK maps those tiers to matrix sides
32 / 96 / 128, for example, while the dashboard and merge use only the common tier labels. Operations
that do not consult TensorEngine.Blas still run in both processes, but their OpenBLAS-profile record
correctly says actual_backend: managed rather than crediting the native library.
Stable architecture areas: A1 NumPy twin; A2 Core-only official BDN runner; A3 OpenBLAS
official BDN runner; A4 targeted backend supplements; A5 exact-cell merge, provenance, and
coverage gates; A6 dashboard/profile consumers.
The remaining subsystem result models are appended as their own *_results.md sections:
| Subsystem | Adds |
|---|---|
nditer |
iterator machinery (construction, traversal, reductions, selection, dtypes, pathologies, dividends) × cache tier — plus the two summary cards cards/{ops,cat}.png (rendered by nditer_cards.py, committed by CI). |
layout |
reduction / copy / elementwise × 8 memory layouts × dtype (the matrix is C-contiguous only). |
operand |
1-D / scalar / mixed-operand / broadcast layouts. |
cast |
full astype src→dst × 8 layouts (no matrix coverage at all). |
fusion |
np.evaluate fused vs unfused chains. |
backends / openblas |
single-threaded targeted supplements for all 39 backend-sensitive APIs; the complete official LinearAlgebra matrix is measured by both BDN executables. |
The run leaves artifacts at several fidelities — know which is canonical:
benchmark-report.md— the canonical per-(op, dtype, N, scenario) backend-aware report + subsystem sections. Tracked; refreshed by CI. Start here for numbers.history/<date>_<sha>/— the committable snapshot (MANIFEST + report + all subsystem sheets + cards + the json/csv/numpy-results that are gitignored at the root).history/latestis the stable path docs and CI reference.benchmark-dashboard.md— a dense ASCII-bar sheet fromrender_dashboard.py. Gitignored, not wired intorun_benchmark.py/CI; run it by hand to seed the DocFX dashboard's numbers.docs/website-src/docs/benchmarks-dashboard.md— the real UI, promoted from the approved backend POC. It consumes the combined JSON directly; effective rollups and backend drill-downs are generated from measured rows, while profile JSON files remain independently inspectable.
CI (.github/workflows/benchmark.yml, post-release / manual) runs the whole harness and commits the
refreshed report + cards + profile JSON + history/ snapshot, then redeploys the docs.
- C# — add a
[Benchmark(Description = "np.foo(a)")]method to the class inBenchmarks/<Category>/that fits the suite's namespace filter.BenchmarkBasefor single-dtype (float64),TypedBenchmarkBasefor dtype-swept;[Params(...)]for size. - NumPy twin — append to the matching
run_<suite>_benchmarks(...)innumpy_benchmark.py; make.namenormalize to the C#Description(both collapse tonp.foovianormalize_op_name). - Smoke it —
dotnet build -c Release; confirm BDN discovers it (dotnet run -c Release --no-build -f net10.0 -- --list flat | grep <Class>) and the NumPy rows emit (python numpy_benchmark.py --suite <suite> --depth light). Full numbers come frompython run_benchmark.py --depth measure(or the post-release CI job).
CLAUDE.md— the exhaustive guide: architecture, every suite, infrastructure classes, the report/UI surfaces, config internals, troubleshooting, type map.nditer/README.md— the NDIter harness (section isolation, the intermittent-AV workaround, the findings ledger).NumSharp.Benchmark.CSharp/README.md— the C# project.