Visitar URL original
Add WaveDB by Winkman4000 · Pull Request #2424 · ClickHouse/ClickBench · GitHub
Skip to content

Add WaveDB - #2424

Open
Winkman4000 wants to merge 1 commit into
ClickHouse:mainfrom
Winkman4000:add-wavedb
Open

Winkman4000 wants to merge 1 commit into
ClickHouse:mainfrom
Winkman4000:add-wavedb

Conversation

@Winkman4000

Copy link
Copy Markdown

Add WaveDB

WaveDB is a column-oriented analytical database written in Python, with its hot loops compiled by numba. It is a single developer's research engine. It runs here as a small HTTP server (wdb serve); ./query posts each statement to it and curl measures the round trip.

How it runs

  • install clones the public repo and checks out a pinned commit (aaa517b), installs pinned Python packages, and compiles every numba kernel signature the engine uses ahead of time (tools/kernel_build.py build), so no compilation happens inside the timed runs.
  • load reads the single hits.parquet with the engine's defaults. No environment variables are set anywhere.
    • --cluster-by EventTime: rows are stored ordered by EventTime (the table's sort order).
    • --cast EventDate=date_days --cast EventTime=timestamp_s: the same conversions DuckDB's load does.
    • --hash URLHash,RefererHash: a storage codec for those two columns. Each row stores its code or the distance back to the previous row with the same code. It replaces the stored codes and is only ever decoded: no index, no lookup structure.
  • start starts the server; it loads the compiled kernels and opens the database before check succeeds. Cold runs restart it and drop the page cache.

What is on disk (data-size counts all of it)

  • The column data: dictionary-encoded columns, front-coded text, block dictionaries.
  • Load statistics: per block of rows, the smallest and largest code and the non-NULL count, plus sampled entries of block dictionaries used to seek inside a column.
  • For every large text column, the character length of each dictionary entry and of each row (kept for every such column by default, no column is named).

No index, projection, materialized view or pre-aggregated table is built, by the load or by any query. No query result or intermediate result is cached; between hot runs only decoded source data (dictionaries, codes) stays in memory. COUNT(*) is the stored row count, an unfiltered COUNT(DISTINCT col) is the length of the column's own dictionary, and unfiltered MIN/MAX/COUNT can be answered from the per-block min/max and non-NULL counts (like ClickHouse's min/max-count projection). The README in the directory says the same.

Result: c6a.4xlarge, 500 GB gp2, Ubuntu 24.04, run with this directory and the repository's unmodified driver. 43/43 queries returned results, no nulls. Load 558 s, data size 9,176,178,269 bytes, concurrent QPS 4.41 with no errors. All 43 answers are checked against DuckDB in the engine's own test board.

WaveDB (https://github.com/Winkman4000/WaveDB): a column-oriented analytical database
in Python with numba-compiled kernels, served over HTTP. Default settings; install
pins engine commit aaa517b. Result: c6a.4xlarge, 2026-10-06.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018HG1WjR4535M7dRqpdWqCJ
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

This branch is waiting to be deployed

1 waiting deployment
benchmark-approval — 4d61105f Waiting Oct 6, 2026 by Winkman4000 via launch #666
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant