Skip to content

Latest commit

 

History

History

README.md

results-data/

This directory contains canonical schema-v2 BenchBox result bundles. Its role depends on the branch:

  • On develop, it is the curated maintainer seed and release-preview source used by the static Results Explorer build.
  • On published-results, it is the complete accepted Phase 2 archive, including direct operator and community submissions.

Published-only archive bundles are expected. They do not appear in the static Explorer until a separate reviewed promotion adds them to develop.

Directory Structure

results-data/
  bundles/
    {run_id}.json            # primary result bundle (schema-v2)
    {run_id}.plans.json      # query plans (if captured)
  corpus-inventory.json      # generated inventory index

The current checked-in corpus keeps bundles flat under results-data/bundles/. The explorer pipeline and validator both scan recursively, and the seed-corpus workflow stages PRs under results-data/bundles/{benchmark}/{platform}/sf{scale}/, so both flat and nested layouts are supported.

Examples:

  • bundles/tpch_sf001_duckdb_20260403_093653_9c0925d1.json
  • bundles/tpch_sf01_polars_df_20260404_191727_mcp_b897c572.json
  • bundles/ssb_sf01_duckdb_sql_20260404_191819_68f79876.json

Scale factor in the path uses the raw value passed to --scale: 0.01 → sf0.01, 0.1 → sf0.1, 1.0 → sf1.0, 1 → sf1

Trust Labels

Label Description
maintainer-run Generated by BenchBox CI via the seed-corpus.yml workflow
community-submission Submitted by a contributor via PR (Phase 2)

Trust label is derived at inventory-generation time from the recorded provenance in the submission-manifest sidecar (<stem>.manifest.json, or the legacy directory-wide submission-manifest.json singleton):

Sidecar state Trust label
absent maintainer-run (maintainer-committed)
result_source: internal maintainer-run
result_source: community community-submission
result_source: vendor vendor-supplied
present but no/unknown result_source community-submission (fails safe)

See scripts/generate_corpus_inventory.py::_bundle_trust_label and the matching _project/scripts/explorer_pipeline/pipeline.py::_manifest_trust_label; a test pins the two in lockstep.

Sidecar presence used to imply community provenance on its own. That was wrong: benchbox submit writes a sidecar unconditionally, including when a maintainer runs it, so the 188 bundles from the maintainer's own 2026-05-02 UAT sweep were labelled community-submission. Community submissions are not ranking-eligible, so the public leaderboard rendered 15 of 207 results. Those sidecars now record result_source: internal, matching the provenance this document already described in prose.

Seed Corpus

After the 2026-08-28 trust boundary, the checked-in corpus holds 364 maintainer-run bundles across 20 benchmarks and 58 cohorts, all at the >=3-identity validator floor. Covered families include the local set (amplab, clickbench, coffeeshop, h2odb, joinorder, read_primitives, ssb, tpcds, tpch, tpch_skew), the admitted datavault, flightdata, nyctaxi, tpcdi, tpcds_obt, and tpchavoc cohorts, and 119 live cloud bundles (BigQuery, Snowflake, Databricks; added 2026-10-02 and 2026-10-03) that also add metadata_primitives, write_primitives, transaction_primitives and tsbs_devops. See CORPUS_NOTES.md for the cloud cohorts. star_schema is an alias of ssb and is not admitted separately. See REGENERATION.md for deferral detail.

The DuckDB version-matrix cells (ClickBench / SSB / TPC-H / TPC-DS at SF 10) use three independent power repetitions per cell and promote one median bundle per version/benchmark cell; raw repetitions stay outside the checkout. See CORPUS_NOTES.md for operator-run details.

Everything older than 2026-08-23 was withdrawn on 2026-08-28 as a trust decision; see CORPUS_NOTES.md. The maintainer-run seed lane (.github/workflows/seed-corpus.yml) runs at 07:00 UTC on the first day of each month (0 7 1 * *) and remains callable via workflow_dispatch. Its supported local matrix maintains TPC-H at SF 0.01 and SF 0.1 with DuckDB, DataFusion, and Polars DataFrame; TPC-H SF 1 and TPC-DS SF 1 with DuckDB, DataFusion, and ClickHouse Local; and SSB at SF 0.01 and SF 0.1 with DuckDB, DataFusion, and Polars DataFrame. It does not claim to regenerate every checked-in cohort; the authoritative cell list is the workflow file and is summarized in SEED_CORPUS_SPEC.md.

For the up-to-date per-cohort breakdown, see results-data/corpus-inventory.json (regenerate via uv run -- python scripts/generate_corpus_inventory.py --write). results-data/SEED_CORPUS_SPEC.md documents the seed-lane contract and the validator gate.

Tuned bundles dropped (2026-07-16)

The 325 seed-corpus bundles that claimed execution.tuning_mode == "tuned" were removed: the 2026-07 tuning remediation (#1176 w0 discovery) found the tuning config never reached platform adapters on the direct CLI path, so those bundles were almost certainly physically untuned and unsound for cross-submission comparison. The 203 bundles with no tuning_mode recorded were left in place - they ingest honestly as "not recorded". See results-data/REGENERATION.md for the full removed-cell checklist and the regeneration procedure once the tuned path is fixed and verified (post-#1176/#1180).

Contributing via Pull Request (Phase 2)

Community contributions use a pull request against the slim published-results branch. That branch intentionally does not contain the full documentation tree, so the complete runnable flow is included here:

  1. Install BenchBox with the extra for your platform, then run a complete suite: uv run -- benchbox run --platform <platform> --benchmark <benchmark> --scale <sf>.
  2. Set a stable, private BENCHBOX_MACHINE_ID_SALT for public submission. Store and reuse it through a secret manager or protected local environment config.
  3. Package the latest result: uv run -- benchbox submit --last --output ./submission.
  4. Fork BenchBox-dev/BenchBox, check out its published-results branch, and copy submission/bundle/ plus the generated submission/<result>.manifest.json into results-data/bundles/.
  5. Regenerate the inventory: uv run -- python scripts/generate_corpus_inventory.py --write.
  6. Commit the bundle, manifest, and inventory, then open a PR against BenchBox-dev/BenchBox:published-results titled results: <benchmark> <platform> sf<scale>.
  7. CI validates schema conformance, hashes, bundle integrity, cohort compatibility, and inventory drift before maintainer review.

The full maintained guide is published at https://benchbox.dev/docs/contributing-results.html.

Maintainer archive path (not community)

Community add PRs above are the supported submission path. Separately, maintainers handling seed or curated archive changes must follow the published-results deletion and mirror procedure: hand-opened maintainer PRs against published-results are deletion-only; maintainer and seed additions go through the auto/results-mirror-* sync, not a hand-opened add PR.

Reproducibility

Any bundle in this directory can be re-run locally using the parameters recorded in the bundle metadata:

benchbox run \
  --platform <platform> \
  --benchmark <benchmark> \
  --scale <scale_factor> \
  --phases generate,load,power \
  --compression zstd:9