Visitar URL original
NumSharp/benchmark at master · SciSharp/NumSharp · GitHub
Skip to content

Latest commit

 

History

History

README.md

NumSharp Benchmarks

Fair, reproducible NumSharp vs NumPy 2.4.2 performance comparison: an op × dtype × N matrix, the complete official LinearAlgebra suite under both Managed C# and OpenBLAS, targeted backend supplements for routes outside that matrix, and five complementary subsystems that fill axes the matrix cannot express.

This file is an orientation guide, not the results. For the numbers, see the canonical report benchmark-report.md (regenerated by CI after each release) or the committed snapshot at history/latest/benchmark-report.md. The human-facing UI is the website's Benchmarks dashboard (docs/website-src/docs/benchmarks-dashboard.md).

The convention: NPY/NS (memorize this)

ratio = NumPy_ms ÷ NumSharp_ms. >1 = NumSharp faster, <1 = slower, =1 = parity. Higher is better. %NumPy🕐 = NumSharp_ms ÷ NumPy_ms × 100 = the share of NumPy's time NumSharp uses (30% = NumSharp takes only 30% as long; <100% = faster).

Bands used by the report: ✅ ≥1.05× · 🟡 0.8–<1.05× · 🟠 0.33–<0.8× · 🔴 <0.33× · ▫ negligible (semantic O(1)-in-N scenario, sub-µs, or >20× — excluded from every rollup/ranking) · ⚪ C# side not run.

Timing basis: every ratio compares the two sides' per-case min (best window). Min-based harnesses run each case to a ~200 ms time budget (a >20 ms/call op runs exactly 100 times); the C# op-matrix keeps BenchmarkDotNet's 50 iterations — never the mean, whose NumSharp-side GC/page-fault/machine-load tails turned host state into fake regressions; per-case means stay in the report JSON (numpy_mean_ms/numsharp_mean_ms) as tail diagnostics.

⚠️ The legacy run-benchmarks.ps1 prints the inverse (NS/NPY, lower is better). Everything else — the op-matrix report, the dashboards, every *_sheet.py — is NPY/NS. Prefer NPY/NS.

⚠️ Pitfall: ad-hoc dotnet run builds Debug (~2× slow)

File-based dotnet run file.cs / dotnet_run compile both the script AND any #:project NumSharp.Core in Debug, which the JIT honors even over [MethodImpl(AggressiveOptimization)]. Hand-written C# hot loops run ~2× slow; IL-emitted kernels look normal. Every timing script MUST run dotnet run -c Release - < script.cs. The BenchmarkDotNet projects are exempt (they mandate -c Release). Full detail in CLAUDE.md.

Run it

python run_benchmark.py                          # interactive depth + dtype picker
python run_benchmark.py --depth measure          # full official run (18 suites + backend profiles + subsystems)
python run_benchmark.py --suites arithmetic unary
python run_benchmark.py --skip-build             # reuse the existing Release build
python run_benchmark.py --skip-csharp            # NumPy only
python run_benchmark.py --quick                  # deprecated alias for --depth light

Execution depth and dtype targeting:

python run_benchmark.py --depth pass                         # exactly 1 BDN/NumPy call per cell; 0 warmups
python run_benchmark.py --depth light --dtypes f32,c128      # rough estimate; comma-separated dtype filter
python run_benchmark.py --depth measure                      # full publication profile (default with other flags)

With no CLI arguments, an interactive terminal shows a depth options picker followed by a dtype picker. Multi-dtype scenarios match a requested dtype on any side. Pass/light are validation/local estimate profiles: they write only to results/<timestamp>/, skip complementary subsystem sheets, and never replace canonical root, docs, or history artifacts.

run_benchmark.py is the single entry point. It builds the C# suite, runs each suite through BenchmarkDotNet (OfficialBenchmarkConfig: InProcessEmit toolchain, 50 measured iterations, iteration time capped at 25 ms), and reruns the same discovered LinearAlgebra classes/configuration in the OpenBLAS executable. It sweeps NumPy across 1K / 100K / 10M in one process per suite, merges on (op, dtype, N, scenario), records both backend profiles, selects the fastest valid timing per exact cell, and appends the subsystem sheets. Only a successful full measure run mirrors the headline artifacts to this folder/docs and writes the committable history/<date>_<sha>/ snapshot (+ repoints history/latest). Subsets, targeted reruns, pass/light runs, and incomplete runs stay in their run directory.

Long runs, progress, and recovery

Starting python run_benchmark.py in a terminal first checks for an unfinished run and asks Resume this benchmark? [Y/n], showing its depth, dtypes, and last saved progress. Enter resumes with the saved options; n opens the new-run picker. Completed runs, saved plans, active runs, and incompatible source/environment records are skipped. Custom run directories are remembered in results/last-run.json. Explicit CLI arguments never trigger this prompt; non-interactive callers must pass options such as --depth measure or --resume latest.

Run these commands from benchmark/:

python run_benchmark.py --depth measure --run-dir results/nightly --plan-only
python run_benchmark.py --resume results/nightly
python run_benchmark.py --resume latest

--plan-only builds and discovers the exact selected cases, saves their identities, and stops before timing workloads. Omit it to start measuring immediately. The total comes from live discovery, including every selected dtype, size, and engine: NumPy, Managed C#, and the OpenBLAS profile count separately. Do not infer the total from a scenario-title count or an older comparison report.

The parent runner keeps terminal progress and the window/tab title across suite and engine changes: completed/total matrix cases, remaining cases, failures, an estimated remaining time, and the current stage. The ETA learns from observed case durations and is approximate. Subsystems have a separate pending-stage count; their internal cells and durations are not included in the matrix ETA.

Both NumPy and BenchmarkDotNet save each completed matrix case atomically. Interrupting with Ctrl+C keeps completed work; --resume reuses successful cases and retries failed or unfinished ones. Resume also skips completed subsystem stages; an interrupted subsystem restarts at its stage boundary. The saved selection is fixed. Changed source code, compiled benchmark binaries, runtimes, or recorded environment settings block resume to avoid mixing measurements; start a new run with --rerun after making code changes.

Useful controls:

python run_benchmark.py --resume results/nightly --attempts 3 --progress-interval 10
python run_benchmark.py --resume results/nightly --process-timeout 3600 --verbose --no-title
python run_benchmark.py --depth light --operations "np.add*" --run-dir results/add-check

--attempts bounds attempts per incomplete matrix suite (default 3), retaining successful cases between attempts. --process-timeout optionally limits each child process in seconds, including build/discovery processes; expiry stops its process tree. --verbose echoes child output as well as logging it, and --no-title keeps progress in the terminal only.

Rerun failures, slow cases, or regressions

python run_benchmark.py --rerun failed --from-results results/nightly --run-dir results/retry
python run_benchmark.py --rerun bad --from-results history/latest/benchmark-report.json --plan-only
python run_benchmark.py --rerun degraded --from-results results/current --baseline results/baseline --regression-percent 10
  • failed selects actual recorded failures, including failed NumPy checkpoints in a run directory.
  • bad adds unmeasured/pending cases and credible slower/much_slower comparisons. Known missing_backend/not_supported availability states and negligible timings are excluded.
  • degraded selects cases whose NumSharp time increased by more than 10% by default against the same operation/dtype/size/scenario and backend profile in the baseline. It compares absolute NumSharp best-window times, not changes in the NumPy/NumSharp ratio or the fastest effective backend.

Selection matches the current discovered plan exactly and includes NumPy counterparts for selected NumSharp cells. selection.json records reasons and unmatched cases, including historical or subsystem scenarios the official matrix cannot reproduce. Known incompatible baseline metadata is rejected; missing historical provenance is reported. Inspect the plan before spending hours on a rerun.

Each run keeps:

Path inside the run directory Purpose
run-config.json, plan.json Fixed options, source/environment provenance, exact execution plan
build-identity.json, numpy-reference.json Executed binary hashes and pinned input when reusing NumPy timings
run-state.json, events.jsonl Latest progress/stage state and completion events
numpy-checkpoints/ Atomic NumPy case results or failure records
csharp/, csharp-openblas/ Atomic BDN case reports with statistics and measurements
logs/, selection.json Full child output and targeted-rerun diagnostics
numpy-results.json, benchmark-report.* Rebuilt aggregate timings and comparison reports

Older snapshots remain useful selection sources, but cannot be resumed: they omit raw BDN case measurements and have no durable plan/checkpoints. For example, the pre-checkpoint local run results/20260906-230207/ saved 4,996 NumPy rows and only 990 C# rows in five copied class reports, with no final report. Other September 6 runs contain 21,720 C# rows: a fixed “5,000 tests” estimate would miss most engine/case work. Preserve these old artifacts as evidence; use new-format runs for recovery.

What's here

Path What
run_benchmark.py THE orchestrator (op matrix + backend profiles + subsystems + history snapshot).
NumSharp.Benchmark.CSharp/ Core-only BenchmarkDotNet project and shared benchmark assembly — Benchmarks/<Category>/*.cs, config in Infrastructure/.
NumSharp.Benchmark.CSharp.OpenBLAS/ OpenBLAS-enabled executable that discovers the shared LinearAlgebra classes and uses the same official config.
NumSharp.Benchmark.Python/numpy_benchmark.py The NumPy twin — one run_<suite>_benchmarks(...) per suite.
scripts/merge-results.py Joins C# + NumPy on (op, dtype, N) → benchmark-report.{md,json,csv} (NPY/NS + credibility gating).
scripts/merge-backend-profiles.py Merges Managed/OpenBLAS profile JSON, preserves availability exceptions, and selects the effective timing.
scripts/benchmark_session.py, scripts/benchmark_selection.py Durable run state/global progress and exact rerun selection from saved results.
scripts/bench_common.py Shared build/run/parse driver for the file-based subsystems.
scripts/snapshot_history.py Assembles history/<date>_<sha>/ + latest (the publish step).
scripts/render_dashboard.py benchmark-report.json → benchmark-dashboard.md (ASCII sheet; run by hand — seeds the DocFX UI).
scripts/audit_coverage.py + coverage/generated/ Machine-checkable API benchmark ledger (456/456 benchmarkable compatibility APIs).
scenarios.html + scripts/generate_scenarios_html.py Executable 499-title × 15-dtype audit with Comparable, C# BDN, and Python coverage lenses.
backends/ + openblas/ Targeted Managed/OpenBLAS runs using one JSON schema; catches MissingBackendException and NotSupportedException.
nditer/ layout/ operand/ cast/ fusion/ Five appended complementary subsystems.
history/ Tracked snapshots — what we commit and reference. latest → newest.
results/ Gitignored raw per-run scratch.
run-benchmarks.ps1 Legacy Windows-only runner (inverse NS/NPY convention).

Regenerate the scenario/dtype audit from live BenchmarkDotNet attributes and the NumPy suite builders:

python benchmark/scripts/generate_scenarios_html.py
python benchmark/scripts/generate_scenarios_html.py --no-build --check

The op matrix (18 suites)

arithmetic · unary · reduction · broadcast · creation · manipulation · slicing · comparison · bitwise · logic · statistics · sorting · linalg · selection · fft · random · ndarray · api — each a C# namespace filter in run_benchmark.py's SUITES map. The universal-tier contract is strict: once an operation/dtype publishes both 1K and 100K it must also schedule 10M; the official merge rejects asymmetric output independently for NumPy and C#. Scalar/1K-only API-dispatch cases remain valid because they never enter the throughput matrix. For allocation-heavy families, the 10M tier uses a symmetric 1M physical-work cap on NumPy and NumSharp while retaining N=10M as the sole join/dashboard label; this is a workload mapping, not a fourth size tier. The experimental / allocation classes (Dispatch, Fusion, DynamicEmission, SimdVsScalar, MultiDim, NumSharp, Allocation/) have no NumPy twin and are not part of the official run.

Backend profiles and complementary subsystems

Backend coverage has two layers. The official LinearAlgebra BenchmarkDotNet classes run first in the Core-only executable and then unchanged in the OpenBLAS executable; every measured Managed cell must have an available exact-cell OpenBLAS-profile peer or the merge fails. The targeted harness supplements backend-only product/LAPACK routes that are not official op-matrix cells. Missing backend / unsupported outcomes are availability states, not timings. Every profile publishes the same 1K / 100K / 10M workload tiers. Operation-specific physical work stays bounded: LAPACK maps those tiers to matrix sides 32 / 96 / 128, for example, while the dashboard and merge use only the common tier labels. Operations that do not consult TensorEngine.Blas still run in both processes, but their OpenBLAS-profile record correctly says actual_backend: managed rather than crediting the native library.

Stable architecture areas: A1 NumPy twin; A2 Core-only official BDN runner; A3 OpenBLAS official BDN runner; A4 targeted backend supplements; A5 exact-cell merge, provenance, and coverage gates; A6 dashboard/profile consumers. The remaining subsystem result models are appended as their own *_results.md sections:

Subsystem Adds
nditer iterator machinery (construction, traversal, reductions, selection, dtypes, pathologies, dividends) × cache tier — plus the two summary cards cards/{ops,cat}.png (rendered by nditer_cards.py, committed by CI).
layout reduction / copy / elementwise × 8 memory layouts × dtype (the matrix is C-contiguous only).
operand 1-D / scalar / mixed-operand / broadcast layouts.
cast full astype src→dst × 8 layouts (no matrix coverage at all).
fusion np.evaluate fused vs unfused chains.
backends / openblas single-threaded targeted supplements for all 39 backend-sensitive APIs; the complete official LinearAlgebra matrix is measured by both BDN executables.

Outputs & the UI

The run leaves artifacts at several fidelities — know which is canonical:

  • benchmark-report.md — the canonical per-(op, dtype, N, scenario) backend-aware report + subsystem sections. Tracked; refreshed by CI. Start here for numbers.
  • history/<date>_<sha>/ — the committable snapshot (MANIFEST + report + all subsystem sheets + cards + the json/csv/numpy-results that are gitignored at the root). history/latest is the stable path docs and CI reference.
  • benchmark-dashboard.md — a dense ASCII-bar sheet from render_dashboard.py. Gitignored, not wired into run_benchmark.py/CI; run it by hand to seed the DocFX dashboard's numbers.
  • docs/website-src/docs/benchmarks-dashboard.md — the real UI, promoted from the approved backend POC. It consumes the combined JSON directly; effective rollups and backend drill-downs are generated from measured rows, while profile JSON files remain independently inspectable.

CI (.github/workflows/benchmark.yml, post-release / manual) runs the whole harness and commits the refreshed report + cards + profile JSON + history/ snapshot, then redeploys the docs.

Adding a benchmark for a new op

  1. C# — add a [Benchmark(Description = "np.foo(a)")] method to the class in Benchmarks/<Category>/ that fits the suite's namespace filter. BenchmarkBase for single-dtype (float64), TypedBenchmarkBase for dtype-swept; [Params(...)] for size.
  2. NumPy twin — append to the matching run_<suite>_benchmarks(...) in numpy_benchmark.py; make .name normalize to the C# Description (both collapse to np.foo via normalize_op_name).
  3. Smoke it — dotnet build -c Release; confirm BDN discovers it (dotnet run -c Release --no-build -f net10.0 -- --list flat | grep <Class>) and the NumPy rows emit (python numpy_benchmark.py --suite <suite> --depth light). Full numbers come from python run_benchmark.py --depth measure (or the post-release CI job).

Deeper docs

  • CLAUDE.md — the exhaustive guide: architecture, every suite, infrastructure classes, the report/UI surfaces, config internals, troubleshooting, type map.
  • nditer/README.md — the NDIter harness (section isolation, the intermittent-AV workaround, the findings ledger).
  • NumSharp.Benchmark.CSharp/README.md — the C# project.