This directory contains canonical schema-v2 BenchBox result bundles. Its role depends on the branch:
- On
develop, it is the curated maintainer seed and release-preview source used by the static Results Explorer build. - On
published-results, it is the complete accepted Phase 2 archive, including direct operator and community submissions.
Published-only archive bundles are expected. They do not appear in the static
Explorer until a separate reviewed promotion adds them to develop.
results-data/
bundles/
{run_id}.json # primary result bundle (schema-v2)
{run_id}.plans.json # query plans (if captured)
corpus-inventory.json # generated inventory index
The current checked-in corpus keeps bundles flat under results-data/bundles/.
The explorer pipeline and validator both scan recursively, and the seed-corpus
workflow stages PRs under results-data/bundles/{benchmark}/{platform}/sf{scale}/,
so both flat and nested layouts are supported.
Examples:
bundles/tpch_sf001_duckdb_20260403_093653_9c0925d1.jsonbundles/tpch_sf01_polars_df_20260404_191727_mcp_b897c572.jsonbundles/ssb_sf01_duckdb_sql_20260404_191819_68f79876.json
Scale factor in the path uses the raw value passed to --scale:
0.01 → sf0.01, 0.1 → sf0.1, 1.0 → sf1.0, 1 → sf1
| Label | Description |
|---|---|
maintainer-run |
Generated by BenchBox CI via the seed-corpus.yml workflow |
community-submission |
Submitted by a contributor via PR (Phase 2) |
Trust label is derived at inventory-generation time from the recorded
provenance in the submission-manifest sidecar (<stem>.manifest.json, or the
legacy directory-wide submission-manifest.json singleton):
| Sidecar state | Trust label |
|---|---|
| absent | maintainer-run (maintainer-committed) |
result_source: internal |
maintainer-run |
result_source: community |
community-submission |
result_source: vendor |
vendor-supplied |
present but no/unknown result_source |
community-submission (fails safe) |
See scripts/generate_corpus_inventory.py::_bundle_trust_label and the matching
_project/scripts/explorer_pipeline/pipeline.py::_manifest_trust_label; a test
pins the two in lockstep.
Sidecar presence used to imply community provenance on its own. That was
wrong: benchbox submit writes a sidecar unconditionally, including when a
maintainer runs it, so the 188 bundles from the maintainer's own 2026-05-02 UAT
sweep were labelled community-submission. Community submissions are not
ranking-eligible, so the public leaderboard rendered 15 of 207 results. Those
sidecars now record result_source: internal, matching the provenance this
document already described in prose.
After the 2026-08-28 trust boundary, the checked-in
corpus holds 364 maintainer-run bundles across 20 benchmarks and 58
cohorts, all at the >=3-identity validator floor. Covered families include the
local set (amplab, clickbench, coffeeshop, h2odb, joinorder, read_primitives,
ssb, tpcds, tpch, tpch_skew), the admitted datavault, flightdata, nyctaxi,
tpcdi, tpcds_obt, and tpchavoc cohorts, and 119 live cloud bundles (BigQuery,
Snowflake, Databricks; added 2026-10-02 and 2026-10-03) that also add
metadata_primitives, write_primitives, transaction_primitives and tsbs_devops.
See CORPUS_NOTES.md
for the cloud cohorts. star_schema is an alias of ssb and is not admitted
separately. See REGENERATION.md for deferral detail.
The DuckDB version-matrix cells (ClickBench / SSB / TPC-H / TPC-DS at SF 10)
use three independent power repetitions per cell and promote one median bundle
per version/benchmark cell; raw repetitions stay outside the checkout. See
CORPUS_NOTES.md for operator-run details.
Everything older than 2026-08-23 was withdrawn on 2026-08-28 as a trust
decision; see CORPUS_NOTES.md. The maintainer-run seed lane
(.github/workflows/seed-corpus.yml) runs at 07:00 UTC on the first day of
each month (0 7 1 * *) and remains callable via workflow_dispatch. Its
supported local matrix maintains TPC-H at SF 0.01 and SF 0.1 with DuckDB,
DataFusion, and Polars DataFrame; TPC-H SF 1 and TPC-DS SF 1 with DuckDB,
DataFusion, and ClickHouse Local; and SSB at SF 0.01 and SF 0.1 with DuckDB,
DataFusion, and Polars DataFrame. It does not claim to regenerate every
checked-in cohort; the authoritative cell list is the workflow file and is
summarized in SEED_CORPUS_SPEC.md.
For the up-to-date per-cohort breakdown, see
results-data/corpus-inventory.json (regenerate via
uv run -- python scripts/generate_corpus_inventory.py --write).
results-data/SEED_CORPUS_SPEC.md documents the seed-lane contract and the
validator gate.
The 325 seed-corpus bundles that claimed execution.tuning_mode == "tuned"
were removed: the 2026-07 tuning remediation (#1176 w0 discovery) found the
tuning config never reached platform adapters on the direct CLI path, so
those bundles were almost certainly physically untuned and unsound for
cross-submission comparison. The 203 bundles with no tuning_mode recorded
were left in place - they ingest honestly as "not recorded". See
results-data/REGENERATION.md for the full removed-cell checklist and the
regeneration procedure once the tuned path is fixed and verified
(post-#1176/#1180).
Community contributions use a pull request against the slim
published-results branch. That branch intentionally does not contain the full
documentation tree, so the complete runnable flow is included here:
- Install BenchBox with the extra for your platform, then run a complete suite:
uv run -- benchbox run --platform <platform> --benchmark <benchmark> --scale <sf>. - Set a stable, private
BENCHBOX_MACHINE_ID_SALTfor public submission. Store and reuse it through a secret manager or protected local environment config. - Package the latest result:
uv run -- benchbox submit --last --output ./submission. - Fork
BenchBox-dev/BenchBox, check out itspublished-resultsbranch, and copysubmission/bundle/plus the generatedsubmission/<result>.manifest.jsonintoresults-data/bundles/. - Regenerate the inventory:
uv run -- python scripts/generate_corpus_inventory.py --write. - Commit the bundle, manifest, and inventory, then open a PR against
BenchBox-dev/BenchBox:published-resultstitledresults: <benchmark> <platform> sf<scale>. - CI validates schema conformance, hashes, bundle integrity, cohort compatibility, and inventory drift before maintainer review.
The full maintained guide is published at https://benchbox.dev/docs/contributing-results.html.
Community add PRs above are the supported submission path. Separately,
maintainers handling seed or curated archive changes must follow the
published-results deletion and mirror procedure:
hand-opened maintainer PRs against published-results are deletion-only;
maintainer and seed additions go through the auto/results-mirror-* sync,
not a hand-opened add PR.
Any bundle in this directory can be re-run locally using the parameters recorded in the bundle metadata:
benchbox run \
--platform <platform> \
--benchmark <benchmark> \
--scale <scale_factor> \
--phases generate,load,power \
--compression zstd:9