Repository navigation
Add WaveDB - #2424
Open
Winkman4000 wants to merge 1 commit into
Open
Add WaveDB#2424Winkman4000 wants to merge 1 commit into
Winkman4000 wants to merge 1 commit into
Conversation
WaveDB (https://github.com/Winkman4000/WaveDB): a column-oriented analytical database in Python with numba-compiled kernels, served over HTTP. Default settings; install pins engine commit aaa517b. Result: c6a.4xlarge, 2026-10-06. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018HG1WjR4535M7dRqpdWqCJ
Winkman4000
requested a deployment
to
benchmark-approval
October 6, 2026 11:32 — with
GitHub Actions
Waiting
Contributor
This branch is waiting to be deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Add WaveDB
WaveDB is a column-oriented analytical database written in Python, with its hot loops compiled by numba. It is a single developer's research engine. It runs here as a small HTTP server (
wdb serve);./queryposts each statement to it and curl measures the round trip.How it runs
installclones the public repo and checks out a pinned commit (aaa517b), installs pinned Python packages, and compiles every numba kernel signature the engine uses ahead of time (tools/kernel_build.py build), so no compilation happens inside the timed runs.loadreads the singlehits.parquetwith the engine's defaults. No environment variables are set anywhere.--cluster-by EventTime: rows are stored ordered byEventTime(the table's sort order).--cast EventDate=date_days --cast EventTime=timestamp_s: the same conversions DuckDB's load does.--hash URLHash,RefererHash: a storage codec for those two columns. Each row stores its code or the distance back to the previous row with the same code. It replaces the stored codes and is only ever decoded: no index, no lookup structure.startstarts the server; it loads the compiled kernels and opens the database beforechecksucceeds. Cold runs restart it and drop the page cache.What is on disk (
data-sizecounts all of it)No index, projection, materialized view or pre-aggregated table is built, by the load or by any query. No query result or intermediate result is cached; between hot runs only decoded source data (dictionaries, codes) stays in memory.
COUNT(*)is the stored row count, an unfilteredCOUNT(DISTINCT col)is the length of the column's own dictionary, and unfiltered MIN/MAX/COUNT can be answered from the per-block min/max and non-NULL counts (like ClickHouse's min/max-count projection). The README in the directory says the same.Result: c6a.4xlarge, 500 GB gp2, Ubuntu 24.04, run with this directory and the repository's unmodified driver. 43/43 queries returned results, no nulls. Load 558 s, data size 9,176,178,269 bytes, concurrent QPS 4.41 with no errors. All 43 answers are checked against DuckDB in the engine's own test board.