The top-level README is the quick start and the proof-it-works benchmark evidence. This folder is the reference layer underneath it: how the system is built, every command and config field, and the enterprise-facing pieces (auth, governance, observability, deployment, scaling) that don't fit in a quick start.
Written for three readers:
- Someone integrating BuffData into a pipeline or CI system → start with cli-reference.md and configuration.md.
- Someone standing up team infrastructure around it (auth, secrets, a shared LLM server, a Kubernetes deployment) → providers.md, governance.md, deployment.md.
- Someone changing BuffData's own code → architecture.md first.
| Doc | What's in it |
|---|---|
| architecture.md | The 7-stage pipeline, the provider-neutral client contract, data flow, where each guarantee (strict/local network policy, accuracy contracts) actually lives in the code |
| cli-reference.md | Every command, grouped by what it's for, with real examples |
| configuration.md | Every PipelineConfig field: type, default, what it controls, and the YAML equivalent |
| security.md | API keys in the OS keyring instead of a plaintext file (buffdata auth), and owner-only file permissions on checkpoints/reports/the audit DB |
| providers.md | Cloud providers, local LLM servers (Ollama/LM Studio/vLLM/llama.cpp), secret backends (Vault/AWS/GCP/Azure), cloud dataset storage (S3/GCS/ADLS) |
| governance.md | Access policy (RBAC), OIDC/SSO bearer-token verification, the durable audit log, Data Contracts, SBOM generation |
| observability.md | OpenTelemetry tracing and Prometheus metrics per pipeline run |
| scaling.md | Dataset sharding for orchestrator-level parallelism (Airflow/Dagster/k8s), and the Ray-based in-process alternative |
| deployment.md | Docker image, Helm chart, Terraform module -- what's verified and what isn't |
| classification-benchmark.md | Binary vs. multi-class recovery and classification accuracy across 19 real datasets with gemini-3.7-flash |
- Benchmark numbers and methodology live in the README
and
benchmarks/-- this folder links to them rather than re-copying tables that would drift out of sync. - Full flag-by-flag
--helpoutput for every command isn't reproduced verbatim here; cli-reference.md covers what each command is for and its most load-bearing flags, and points atbuffdata <command> --helpfor the complete, always-current list. - Deployment artifacts' own detailed usage notes live beside the files themselves in
deploy/; deployment.md is the map, not a copy.