This directory contains the BenchBox test suite. Directories group tests by purpose; pytest markers select the execution lanes.
See also: AGENTS.md for the contributor/agent guide, and docs/development/ for architecture deep-dives.
tests/
├── contracts/ # Shared contract fixtures
├── databases/ # Local test database helpers
├── docs/ # Documentation checks and test plans
├── e2e/ # End-to-end CLI workflow tests
├── examples/ # Example usage checks
├── fixtures/ # Shared test fixtures
├── integration/ # Integration tests
├── parity/ # Parity fixtures and generators
├── performance/ # Performance tests
├── system/ # Repository and CI system checks
├── uat/ # User acceptance checks and support code
├── unit/ # Unit tests
├── utilities/ # Test utilities and helpers
├── validation/ # Data and query validation checks
├── test_*.py # Root-level benchmark and runner tests
├── conftest.py # Global pytest configuration
└── README.md # This file
Pytest uses the root pytest.ini by default (fast local runs). CI-oriented
targets such as make test-ci and make coverage-fast explicitly select the
root pytest-ci.ini profile with pytest -c pytest-ci.ini.
Tests of individual components:
- benchmarks/: Core benchmark functionality
- core/: Base classes and utilities
- generators/: Data generation components
Many are isolated and fast, but the directory does not itself guarantee a
runtime or absence of external dependencies. Select the fast marker for the
curated fast lane.
End-to-end tests that validate complete CLI workflows:
- CLI option validation
- Error handling coverage
- Result schema validation
- Local platform execution
- Cloud platform dry-run workflows
- DataFrame platform workflows
Characteristics:
- Tests full CLI workflow from command to results
- Validates all platforms (local, cloud, DataFrame)
- Uses dry-run mode for cloud platforms (no credentials needed)
- Includes result file schema validation
Markers:
e2e: All E2E testse2e_quick: Quick dry-run testse2e_local: Local platform tests (full execution)e2e_cloud: Cloud platform dry-run testse2e_dataframe: DataFrame platform tests
Tests that verify component interactions:
- Database connectivity and query execution
- End-to-end benchmark workflows
- Cross-component data flow
Characteristics:
- Real database connections
- File system operations
Integration tests may use local services; live network tests use the
live_integration marker and are excluded from the default local lanes.
Tests focused on performance characteristics:
- Query execution benchmarks
- Data generation performance
- Memory usage analysis
- Scalability testing
This directory includes resource monitoring, statistical analysis, and
baseline comparisons. Use the fast, medium, slow, stress, and
resource_heavy markers to select tests by execution cost rather than by
directory name.
The standard gates are intentionally split by the risk they are meant to catch:
make test-fast: quick developer and develop-PR feedback for code-impacting changes. The matching selection in.github/workflows/pr.ymlalso collects coverage.make test-correctness-gate: bounded develop-PR real-result gate. It runs the DuckDB TPC-H matrix slice (SF=1, pinned reference qgen seed) through generate, load, and execute, then validates the emitted stream-0 results against the stored TPC-H answers with EXACT row-count checking and stored VALUE digests. The gated subset is the 18 TPC-H queries whose answer-set cardinalities are stable across dbgen builds; Q11/Q16/Q18/Q20 are excluded because their HAVING/threshold boundaries make the stored row count vary with the generated data. The subset is deliberately discriminating — it is not dominated by one-row queries and includes multiple high-cardinality answer-backed queries (e.g. Q9=175, Q2/Q21=100) — so a wrong join/filter/aggregate that still emits one row is caught. The subset shape is ratcheted intests/unit/test_standardized_test_commands.py(TestCorrectnessGateOracle). What the gate proves and what it does not:- Value + cardinality at SF=1/pinned-seed: with
BENCHBOX_EMIT_RESULT_DIGEST=1the runner emits an order-normalized digest of each stream-0 query's full result set (reusingbenchbox.core.tpchavoc.validation.calculate_checksum, the same primitive the TPC-Havoc gates use), which the gate asserts against a stored reference digest (benchbox/core/expected_results/reference_digests/tpch_value_digests_sf1.json) in addition to the row count. So a wrong-but-same-cardinality answer — a perturbed Q1 aggregate, a swapped column, a changed rounding — turns the gate RED, not just a wrong row count. Sensitivity is proven intests/unit/test_correctness_gate_value_oracle.py. - Regression snapshot, NOT an independent oracle: the reference digests were
produced by running benchbox-on-DuckDB and frozen, so the value check detects
change from that DuckDB-pinned baseline (a regression tripwire), not correctness
against an external authority. A conceptual value bug present at freeze time is
enshrined in the reference, not caught. Read a green
value+cardinalitycell as "unchanged from the frozen DuckDB answer", never "values proven correct". The reference is regenerated (never hand-copied) bymake correctness-gate-digests-regen, which reruns the same gate config and writes the file idempotently. The independence gap and the deferred cross-engine upgrade are analyzed in_project/analysis/value-digest-cross-engine-independence-decision.md. Two further fidelity properties are pinned intest_correctness_gate_value_oracle.py: the sensitivity floor is relative (~1e-6 per cell via significant-figure rounding, uniform across column magnitude), and the digest is a value+type digest (DuckDB-pinned), which is why cross-engine reuse is deferred. - Values are UNGUARDED above SF=1: stored answers and digests exist only at SF=1 (the expected-results loader raises for other scales), so the value guarantee holds at SF=1 only. There is no expected-results value or cardinality oracle above SF=1.
- Strict arming (both axes): with
BENCHBOX_STRICT_EXPECTED_RESULTS=1, every configured query must produce a non-SKIP row-count validation and, where a reference digest exists, an evaluated value digest, or the run fails. A missing digest disarms RED, never green. This is not gated on benchmark name or scale, so a future CI speedup that retargets the gate (a different benchmark, or SF<1 where no answers exist) cannot silently disarm either oracle. - No-skip guard: the Makefile target emits a JUnit report and fails unless
exactly one node ran with zero skips.
pytestexits 0 when a selected node SKIPs (e.g. duckdb unavailable, or the case dropped from the stable matrix), which would otherwise pass the gate without executing anything. - Required CI job composition: the
correctness-gatejob in.github/workflows/pr.ymlruns more than this row-count+value gate. It also runs the value-level cross-surface and TPC-Havoc equivalence gates (tpchavoc-equivalence-report,tpchavoc-dataframe-equivalence-report, and the ssb/amplab/coffeeshop/clickbench/joinorder-synthetic cross-surface reports), so the required job proves value-level equivalence across several benchmarks, not only TPC-H row counts. This composition is ratcheted intests/unit/test_standardized_test_commands.py.
- Value + cardinality at SF=1/pinned-seed: with
make test-integration: non-live, non-stress integration coverage for broader local and main/release validation.make test-local-matrix: opt-in stress matrix for the full local platform benchmark sweep.- Release canary, live cloud, Docker, and UAT evidence remain separate release signals until their cost, credential, and flake policies are suitable for blocking routine PRs.
The medium-test job in .github/workflows/pr.yml runs make test-medium,
but only when the heavy tier is needed: code-routed runs where the event is
merge_group or the change touches soundness paths or packaging
(scripts/heavy_tier_needed.py reports heavy-needed == 'true'). Ordinary
code-change PRs skip it, so do not assume medium coverage ran on a routine PR.
Its marker selection excludes slow, stress, resource-heavy, and live
integration tests. Product-critical tests that need a different selection
belong in an explicit workflow or correctness gate.
# Run the curated fast lane
make test-fast
# or
uv run -- python -m pytest -m fast
# Run specific benchmark tests
uv run -- python -m pytest tests/unit/benchmarks/test_tpch_core.py
# Run with coverage (fast tests only - quick feedback)
make coverage-fast
# or routine coverage (excludes stress/resource-heavy/live tests)
make coverage-all
# or full tree including opt-in stress/resource-heavy/live tests (needs services + credentials)
make coverage-opt-in-all
# or
uv run -- python -m pytest --cov=benchbox --cov-report=html# Quick E2E tests (dry-run mode)
make test-e2e-quick
# or
uv run -- python -m pytest -m e2e_quick
# Local platform E2E tests (full execution)
uv run -- python -m pytest -m e2e_local
# All E2E tests
uv run -- python -m pytest tests/e2e/ -v
# Specific E2E test module
uv run -- python -m pytest tests/e2e/test_cli_options.py -v# Run all tests
make test-all
# or
uv run -- python -m pytest
# Run with parallel execution
make test-parallel
# or
uv run -- python -m pytest -n auto
# Run integration tests
make test-integration
# or
uv run -- python -m pytest -m "integration and not live_integration and not stress"
# Run performance tests
uv run -- python -m pytest tests/performance/ -m performance# Run optimized development tests
uv run -- python tests/utilities/unified_test_runner.py --strategy development
# Run CI-optimized tests
uv run -- python tests/utilities/unified_test_runner.py --strategy ci
# Run specific benchmarks with parallel execution
uv run -- python tests/utilities/unified_test_runner.py --benchmark tpch tpcds --parallel --workers 4
# Run with coverage reporting
uv run -- python tests/utilities/unified_test_runner.py --coverage --report
# Run benchmark validation
uv run -- python tests/utilities/benchmark_validator.py --benchmark all --quick-check# Profile a test run
uv run -- python tests/utilities/performance_profiler.py python -m pytest tests/unit/
# Profile with detailed output
uv run -- python tests/utilities/performance_profiler.py --output performance_report.md python -m pytest tests/unit/
# Check for performance regressions
uv run -- python tests/utilities/performance_profiler.py --check-regressions python -m pytest tests/unit/# Update performance baselines
uv run -- python tests/utilities/performance_profiler.py --update-baseline python -m pytest tests/unit/
# Profile specific test categories
uv run -- python tests/utilities/performance_profiler.py python -m pytest tests/integration/ -m "integration and not slow"Tests are organized using pytest markers for selective execution:
fast: Quick tests suitable for developmentslow: Tests that take significant timememory_intensive: Tests using significant memorycpu_intensive: Tests using significant CPUio_intensive: Tests performing heavy I/O
unit: Unit testsintegration: Integration testse2e: End-to-end CLI testse2e_quick: Quick E2E tests (dry-run mode)e2e_local: Local platform E2E testse2e_cloud: Cloud platform E2E tests (dry-run)e2e_dataframe: DataFrame platform E2E testsperformance: Performance tests
sqlite: Tests requiring SQLiteduckdb: Tests using DuckDBdatabase: Tests requiring database connections
tpch: TPC-H benchmark teststpcds: TPC-DS benchmark testsssb: Star Schema Benchmark testsamplab: AMPLab Big Data Benchmark testsclickbench: ClickBench testsh2odb: H2O Database benchmark testsprimitives: Primitive operations tests
# Run only fast unit tests
make test-dev
# or
uv run -- python -m pytest -m "unit and fast"
# Run integration tests excluding slow ones
make test-integration
# or
uv run -- python -m pytest -m "integration and not slow"
# Run TPC-H related tests
make test-tpch
# or
uv run -- python -m pytest -m tpch
# Run all tests except memory intensive ones
uv run -- python -m pytest -m "not memory_intensive"tests/conftest.py acquires a single inter-process file lock at
~/.benchbox/test.lock for the duration of each parallel test session. This
prevents two concurrent pytest -n auto runs from fighting for CPU, which
would otherwise double runtime and produce flaky timing assertions.
The default is intentionally global across worktrees: ten retained worktrees
can otherwise start ten independent pytest -n auto runs and saturate a
developer workstation. For isolated debugging or sandboxed CI, set
BENCHBOX_TEST_LOCK_DIR=/path/to/dir; BenchBox will use
$BENCHBOX_TEST_LOCK_DIR/test.lock instead. make test-unlock honors the same
override.
Contract (tests/conftest.py:52-216):
- Only the xdist controller process locks - workers do not.
pytest_configurecheckshasattr(config, "workerinput")to distinguish them. - Lock is skipped when
numprocessesis 0/None (no-n auto) or whenBENCHBOX_SKIP_TEST_LOCK=1is set in the environment. - Lock path defaults to
~/.benchbox/test.lock;BENCHBOX_TEST_LOCK_DIRchanges only the directory, not the filename. - Uses
fcntl.flock(LOCK_EX | LOCK_NB)on POSIX,msvcrt.lockingon Windows. - On contention, the second local run waits up to 3600 seconds with holder
information and periodic progress messages. Set
BENCHBOX_TEST_LOCK_WAIT_SECONDS=0to fail immediately; CI does this. Ctrl-C cancels the wait without disturbing the holder. - The fd is held open for the whole session and released in
pytest_unconfigure.
When to bypass: only for intentional concurrent debug runs. Set
BENCHBOX_SKIP_TEST_LOCK=1 - but expect noisy timing results and CPU
contention. Do not bypass in CI.
On macOS, xdist workers suppress setproctitle to avoid a
launchservicesd CPU storm:
xdist calls setproctitle() twice per test (running/idle) via xdist.remote.worker_title(). At ~200 calls/second this triggers macOS launchservicesd to rebuild its process registry continuously, consuming 200%+ CPU and ~900 MB RSS - the actual root cause of the macOS beachball during parallel test runs.
The workaround walks the call stack to find the live xdist exec namespace
and replaces worker_title with a no-op. This is necessary because
import xdist.remote in an execnet __channelexec__ namespace resolves
to a different module object than the one executing - so patching by
module import alone is a no-op.
This branch only fires when sys.platform == "darwin" and the process is
an xdist worker. Linux and Windows are unaffected.
name: Test Suite
on: [push, pull_request]
jobs:
test:
runs-on: ubuntu-latest
strategy:
matrix:
python-version: ['3.11', '3.12', '3.13', '3.14']
steps:
- uses: actions/checkout@v3
- name: Set up Python ${{ matrix.python-version }}
uses: actions/setup-python@v4
with:
python-version: ${{ matrix.python-version }}
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install -e .[dev]
- name: Run fast tests
run: |
uv run -- python tests/utilities/unified_test_runner.py --strategy ci
- name: Upload coverage
uses: codecov/codecov-action@v3
with:
file: ./coverage.xmlpipeline {
agent any
stages {
stage('Setup') {
steps {
sh 'python -m pip install --upgrade pip'
sh 'pip install -e .[dev]'
}
}
stage('Fast Tests') {
steps {
sh 'uv run -- python tests/utilities/unified_test_runner.py --strategy ci --parallel --workers 4 --output junit'
}
post {
always {
junit 'test_results.xml'
}
}
}
stage('Integration Tests') {
when {
branch 'main'
}
steps {
sh 'uv run -- python tests/utilities/unified_test_runner.py --mode integration --parallel --workers 2 --coverage'
}
post {
always {
publishHTML([
allowMissing: false,
alwaysLinkToLastBuild: true,
keepAll: true,
reportDir: 'htmlcov',
reportFiles: 'index.html',
reportName: 'Coverage Report'
])
}
}
}
}
}The enhanced pytest.ini provides:
- Comprehensive marker definitions
- Parallel execution configuration
- Coverage reporting setup
- Performance optimization settings
- CI/CD friendly defaults
Global test configuration including:
- Database fixtures
- Data generation helpers
- Performance monitoring setup
- Cleanup utilities
- Use appropriate markers: Mark tests with relevant categories and characteristics
- Keep tests focused: Each test should verify one specific behavior
- Use fixtures: Leverage shared fixtures for database connections and test data
- Handle resources: Ensure proper cleanup of temporary files and connections
- Performance awareness: Use performance markers for resource-intensive tests
- Follow naming conventions: Use descriptive test names with
test_prefix - Group related tests: Organize tests by functionality and component
- Use clear assertions: Make test failures easy to understand
- Document complex tests: Add docstrings for complex test scenarios
- Establish baselines: Use the performance profiler to set baseline metrics
- Monitor regressions: Regularly check for performance regressions
- Profile selectively: Only profile tests when needed to avoid overhead
- Optimize test execution: Use caching and parallel execution for faster feedback
# Profile test execution
uv run -- python tests/utilities/performance_profiler.py python -m pytest tests/unit/ -v
# Run only fast tests
make test-fast
# or
uv run -- python -m pytest -m fast
# Use parallel execution
make test-parallel
# or
uv run -- python -m pytest -n auto# Run memory-intensive tests separately
uv run -- python -m pytest -m "memory_intensive" --maxfail=1
# Monitor memory usage
uv run -- python tests/utilities/performance_profiler.py --output memory_report.md python -m pytest tests/unit/# Run database tests with verbose output
make test-integration
# or
uv run -- python -m pytest tests/integration/ -v -s
# Test database connectivity
uv run -- python -c "import duckdb; print(duckdb.connect().execute('SELECT 1').fetchone())"# Clear pytest cache
uv run -- python -m pytest --cache-clear
# Clear custom test cache
uv run -- python tests/utilities/unified_test_runner.py --help
# Run with dry-run to see commands
uv run -- python tests/utilities/unified_test_runner.py --dry-runWhen adding new tests:
- Choose the right category: Place tests in the appropriate directory
- Add proper markers: Mark tests with relevant characteristics
- Update documentation: Add test descriptions to this README
- Consider performance: Mark resource-intensive tests appropriately
- Test your tests: Ensure new tests pass in isolation and with the full suite
- Create subdirectory in appropriate category
- Add
__init__.pyfile - Update markers in
pytest.ini - Add documentation to this README
- Update test runner configuration if needed
The test suite includes comprehensive performance monitoring:
- Execution time tracking: Monitor test duration trends
- Memory usage analysis: Track memory consumption patterns
- CPU utilization: Monitor CPU usage during test execution
- I/O monitoring: Track file system operations
- Regression detection: Automatically detect performance regressions
Performance data is stored in ~/.benchbox/test_cache/ and can be analyzed using the performance profiler utility.
The test suite includes several unified utilities for efficient testing:
utilities/unified_test_runner.py: Comprehensive test runner with multiple execution strategiesutilities/benchmark_validator.py: Unified benchmark validation utilityutilities/performance_profiler.py: Performance monitoring and profilingutilities/test_helpers.py: Common test helper functionsutilities/test_runner.py: Enhanced test runner with caching capabilities
These utilities provide a modern, efficient approach to testing and validation.
For issues with the test suite:
- Check this README for common solutions
- Review test output for specific error messages
- Use the performance profiler to identify bottlenecks
- Consult the main BenchBox documentation
- Open an issue with detailed reproduction steps