This guide provides information for developers who want to contribute to BenchBox or understand its internal architecture.
- Python 3.11+
uv0.8 or newer (recommended for environment management). The committeduv.lockuses lockfilerevision = 3; an older uv silently rewrites it to revision 2 as a side effect of anyuv add/uv lockrun. A pre-commit guard (_project/scripts/check_uv_lock_revision.py, also runnable viamake uv-lock-revision-check) rejects such a downgrade — if it fires, upgrade uv (uv self update) and regenerate the lock.git
-
Clone the repository:
git clone https://github.com/BenchBox-dev/benchbox.git cd BenchBox -
Create a virtual environment:
uv venv source .venv/bin/activate -
Install dependencies:
uv pip install -e .[dev,docs]
The
[dev]extra installs testing, linting, and typing tools. Combine it with[docs]when you plan to build the Sphinx documentation locally.
BenchBox uses pytest for testing. Run tests using either make commands or direct pytest:
# Fast tests for quick feedback
make test
# or
uv run -- python -m pytest -m fast
# Full test suite
make test-all
# or
uv run -- python -m pytest
# Unit tests only
make test-unit
# or
uv run -- python -m pytest -m unit
# Integration tests
make test-integration
# or
uv run -- python -m pytest -m "integration and not live_integration"
# With coverage (fast tests only - quick feedback)
make coverage-fast
# or routine coverage (excludes stress/resource-heavy/live tests)
make coverage-all
# or full tree including opt-in stress/resource-heavy/live tests (needs services + credentials)
make coverage-opt-in-all
# or
uv run -- python -m pytest --cov=benchbox --cov-report=term-missingLinting and formatting run through Ruff:
make format
# or
uv run ruff format .
make lint
# or
uv run ruff check .Type checking is available via:
make typecheck
# or
uv run ty checkWe welcome contributions! Please see the CONTRIBUTING.md file in the root of the project for details on our development process, coding standards, and how to submit a pull request.
Maintainers and AI agents use one disposable linked worktree per task. Create it for the branch, work there until the PR merges, then remove that exact clean registration. See the disposable worktree guide for common commands and recovery scenarios. External contributors working from a fork can use the same lifecycle.
BenchBox enforces single-commit squash integration into develop with strict current-base verification:
make pr-open: Pushes the current branch, verifies it is current withorigin/develop, and creates or reuses the pull request. Auto-merge is withheld by default until the branch is marked final.make pr-ready(ormake pr-open READY=1): Runs the exact readiness transaction from caller-supplied evidence and arms auto-merge only after local/remote head, review, required-check, hold, and (for batch mode) final-tree checks pass. When a merge queue is active ondevelop, arming automatically enqueues the PR for speculative combined-tree testing.- Soundness Gate: PRs modifying soundness-critical paths (
benchbox/core/equivalence/,benchbox/core/expected_results/, etc.) cannot be auto-enqueued and require explicit maintainer review. make pr-refresh: Refreshes a stale PR branch ontoorigin/developwhen resolving merge conflicts locally. Run one branch at a time.
See the release guide for the full
maintainer workflow (version-branch flow on a single repo with
develop and main).
BenchBox exposes a CLI-independent validation workflow through
benchbox.core.validation.ValidationService. The service orchestrates preflight,
manifest, database, and platform capability checks and returns structured
ValidationResult objects that mirror the data surfaced by the CLI.
Programmatic consumers can run comprehensive validation with:
from benchbox.core.validation import ValidationService
service = ValidationService()
results = service.run_comprehensive(
benchmark_type="tpcds",
scale_factor=1.0,
output_dir=output_path,
manifest_path=manifest_path,
connection=db_connection,
platform_adapter=adapter,
)
summary = service.summarize(results)This enables benchmarks, adapters, and external automation to reuse the same validation logic without importing CLI modules.
The CLI mirrors these capabilities via the benchbox run command. Use
--enable-preflight-validation, --enable-postgen-manifest-validation, and
--enable-postload-validation to toggle each stage while the core lifecycle
runner captures the resulting validation metadata.
Benchmarks use a lightweight lazy-loading system so the core package can start
quickly and only load optional dependencies when needed. See
docs/development/import-patterns.md for guidance on adding new benchmarks to
the registry, writing tests for lazy imports, and troubleshooting missing
dependency errors.