Skip to content

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

README.md

BenchBox Test Suite

This directory contains the BenchBox test suite. Directories group tests by purpose; pytest markers select the execution lanes.

See also: AGENTS.md for the contributor/agent guide, and docs/development/ for architecture deep-dives.

Test Structure

tests/
├── contracts/                # Shared contract fixtures
├── databases/                # Local test database helpers
├── docs/                     # Documentation checks and test plans
├── e2e/                      # End-to-end CLI workflow tests
├── examples/                 # Example usage checks
├── fixtures/                 # Shared test fixtures
├── integration/              # Integration tests
├── parity/                   # Parity fixtures and generators
├── performance/              # Performance tests
├── system/                   # Repository and CI system checks
├── uat/                      # User acceptance checks and support code
├── unit/                     # Unit tests
├── utilities/                # Test utilities and helpers
├── validation/               # Data and query validation checks
├── test_*.py                 # Root-level benchmark and runner tests
├── conftest.py              # Global pytest configuration
└── README.md               # This file

Pytest uses the root pytest.ini by default (fast local runs). CI-oriented targets such as make test-ci and make coverage-fast explicitly select the root pytest-ci.ini profile with pytest -c pytest-ci.ini.

Test Categories

Unit Tests (unit/)

Tests of individual components:

  • benchmarks/: Core benchmark functionality
  • core/: Base classes and utilities
  • generators/: Data generation components

Many are isolated and fast, but the directory does not itself guarantee a runtime or absence of external dependencies. Select the fast marker for the curated fast lane.

E2E Tests (e2e/)

End-to-end tests that validate complete CLI workflows:

  • CLI option validation
  • Error handling coverage
  • Result schema validation
  • Local platform execution
  • Cloud platform dry-run workflows
  • DataFrame platform workflows

Characteristics:

  • Tests full CLI workflow from command to results
  • Validates all platforms (local, cloud, DataFrame)
  • Uses dry-run mode for cloud platforms (no credentials needed)
  • Includes result file schema validation

Markers:

  • e2e: All E2E tests
  • e2e_quick: Quick dry-run tests
  • e2e_local: Local platform tests (full execution)
  • e2e_cloud: Cloud platform dry-run tests
  • e2e_dataframe: DataFrame platform tests

Integration Tests (integration/)

Tests that verify component interactions:

  • Database connectivity and query execution
  • End-to-end benchmark workflows
  • Cross-component data flow

Characteristics:

  • Real database connections
  • File system operations

Integration tests may use local services; live network tests use the live_integration marker and are excluded from the default local lanes.

Performance Tests (performance/)

Tests focused on performance characteristics:

  • Query execution benchmarks
  • Data generation performance
  • Memory usage analysis
  • Scalability testing

This directory includes resource monitoring, statistical analysis, and baseline comparisons. Use the fast, medium, slow, stress, and resource_heavy markers to select tests by execution cost rather than by directory name.

Test Execution

Automated Gate Contract

The standard gates are intentionally split by the risk they are meant to catch:

  • make test-fast: quick developer and develop-PR feedback for code-impacting changes. The matching selection in .github/workflows/pr.yml also collects coverage.
  • make test-correctness-gate: bounded develop-PR real-result gate. It runs the DuckDB TPC-H matrix slice (SF=1, pinned reference qgen seed) through generate, load, and execute, then validates the emitted stream-0 results against the stored TPC-H answers with EXACT row-count checking and stored VALUE digests. The gated subset is the 18 TPC-H queries whose answer-set cardinalities are stable across dbgen builds; Q11/Q16/Q18/Q20 are excluded because their HAVING/threshold boundaries make the stored row count vary with the generated data. The subset is deliberately discriminating — it is not dominated by one-row queries and includes multiple high-cardinality answer-backed queries (e.g. Q9=175, Q2/Q21=100) — so a wrong join/filter/aggregate that still emits one row is caught. The subset shape is ratcheted in tests/unit/test_standardized_test_commands.py (TestCorrectnessGateOracle). What the gate proves and what it does not:
    • Value + cardinality at SF=1/pinned-seed: with BENCHBOX_EMIT_RESULT_DIGEST=1 the runner emits an order-normalized digest of each stream-0 query's full result set (reusing benchbox.core.tpchavoc.validation.calculate_checksum, the same primitive the TPC-Havoc gates use), which the gate asserts against a stored reference digest (benchbox/core/expected_results/reference_digests/tpch_value_digests_sf1.json) in addition to the row count. So a wrong-but-same-cardinality answer — a perturbed Q1 aggregate, a swapped column, a changed rounding — turns the gate RED, not just a wrong row count. Sensitivity is proven in tests/unit/test_correctness_gate_value_oracle.py.
    • Regression snapshot, NOT an independent oracle: the reference digests were produced by running benchbox-on-DuckDB and frozen, so the value check detects change from that DuckDB-pinned baseline (a regression tripwire), not correctness against an external authority. A conceptual value bug present at freeze time is enshrined in the reference, not caught. Read a green value+cardinality cell as "unchanged from the frozen DuckDB answer", never "values proven correct". The reference is regenerated (never hand-copied) by make correctness-gate-digests-regen, which reruns the same gate config and writes the file idempotently. The independence gap and the deferred cross-engine upgrade are analyzed in _project/analysis/value-digest-cross-engine-independence-decision.md. Two further fidelity properties are pinned in test_correctness_gate_value_oracle.py: the sensitivity floor is relative (~1e-6 per cell via significant-figure rounding, uniform across column magnitude), and the digest is a value+type digest (DuckDB-pinned), which is why cross-engine reuse is deferred.
    • Values are UNGUARDED above SF=1: stored answers and digests exist only at SF=1 (the expected-results loader raises for other scales), so the value guarantee holds at SF=1 only. There is no expected-results value or cardinality oracle above SF=1.
    • Strict arming (both axes): with BENCHBOX_STRICT_EXPECTED_RESULTS=1, every configured query must produce a non-SKIP row-count validation and, where a reference digest exists, an evaluated value digest, or the run fails. A missing digest disarms RED, never green. This is not gated on benchmark name or scale, so a future CI speedup that retargets the gate (a different benchmark, or SF<1 where no answers exist) cannot silently disarm either oracle.
    • No-skip guard: the Makefile target emits a JUnit report and fails unless exactly one node ran with zero skips. pytest exits 0 when a selected node SKIPs (e.g. duckdb unavailable, or the case dropped from the stable matrix), which would otherwise pass the gate without executing anything.
    • Required CI job composition: the correctness-gate job in .github/workflows/pr.yml runs more than this row-count+value gate. It also runs the value-level cross-surface and TPC-Havoc equivalence gates (tpchavoc-equivalence-report, tpchavoc-dataframe-equivalence-report, and the ssb/amplab/coffeeshop/clickbench/joinorder-synthetic cross-surface reports), so the required job proves value-level equivalence across several benchmarks, not only TPC-H row counts. This composition is ratcheted in tests/unit/test_standardized_test_commands.py.
  • make test-integration: non-live, non-stress integration coverage for broader local and main/release validation.
  • make test-local-matrix: opt-in stress matrix for the full local platform benchmark sweep.
  • Release canary, live cloud, Docker, and UAT evidence remain separate release signals until their cost, credential, and flake policies are suitable for blocking routine PRs.

The medium-test job in .github/workflows/pr.yml runs make test-medium, but only when the heavy tier is needed: code-routed runs where the event is merge_group or the change touches soundness paths or packaging (scripts/heavy_tier_needed.py reports heavy-needed == 'true'). Ordinary code-change PRs skip it, so do not assume medium coverage ran on a routine PR. Its marker selection excludes slow, stress, resource-heavy, and live integration tests. Product-critical tests that need a different selection belong in an explicit workflow or correctness gate.

Quick Development Testing

# Run the curated fast lane
make test-fast
# or
uv run -- python -m pytest -m fast

# Run specific benchmark tests
uv run -- python -m pytest tests/unit/benchmarks/test_tpch_core.py

# Run with coverage (fast tests only - quick feedback)
make coverage-fast
# or routine coverage (excludes stress/resource-heavy/live tests)
make coverage-all
# or full tree including opt-in stress/resource-heavy/live tests (needs services + credentials)
make coverage-opt-in-all
# or
uv run -- python -m pytest --cov=benchbox --cov-report=html

E2E Testing

# Quick E2E tests (dry-run mode)
make test-e2e-quick
# or
uv run -- python -m pytest -m e2e_quick

# Local platform E2E tests (full execution)
uv run -- python -m pytest -m e2e_local

# All E2E tests
uv run -- python -m pytest tests/e2e/ -v

# Specific E2E test module
uv run -- python -m pytest tests/e2e/test_cli_options.py -v

Comprehensive Testing

# Run all tests
make test-all
# or
uv run -- python -m pytest

# Run with parallel execution
make test-parallel
# or
uv run -- python -m pytest -n auto

# Run integration tests
make test-integration
# or
uv run -- python -m pytest -m "integration and not live_integration and not stress"

# Run performance tests
uv run -- python -m pytest tests/performance/ -m performance

Using the Unified Test Runner

# Run optimized development tests
uv run -- python tests/utilities/unified_test_runner.py --strategy development

# Run CI-optimized tests
uv run -- python tests/utilities/unified_test_runner.py --strategy ci

# Run specific benchmarks with parallel execution
uv run -- python tests/utilities/unified_test_runner.py --benchmark tpch tpcds --parallel --workers 4

# Run with coverage reporting
uv run -- python tests/utilities/unified_test_runner.py --coverage --report

# Run benchmark validation
uv run -- python tests/utilities/benchmark_validator.py --benchmark all --quick-check

Performance Profiling

Basic Profiling

# Profile a test run
uv run -- python tests/utilities/performance_profiler.py python -m pytest tests/unit/

# Profile with detailed output
uv run -- python tests/utilities/performance_profiler.py --output performance_report.md python -m pytest tests/unit/

# Check for performance regressions
uv run -- python tests/utilities/performance_profiler.py --check-regressions python -m pytest tests/unit/

Advanced Profiling

# Update performance baselines
uv run -- python tests/utilities/performance_profiler.py --update-baseline python -m pytest tests/unit/

# Profile specific test categories
uv run -- python tests/utilities/performance_profiler.py python -m pytest tests/integration/ -m "integration and not slow"

Test Markers

Tests are organized using pytest markers for selective execution:

Execution Characteristics

  • fast: Quick tests suitable for development
  • slow: Tests that take significant time
  • memory_intensive: Tests using significant memory
  • cpu_intensive: Tests using significant CPU
  • io_intensive: Tests performing heavy I/O

Test Categories

  • unit: Unit tests
  • integration: Integration tests
  • e2e: End-to-end CLI tests
  • e2e_quick: Quick E2E tests (dry-run mode)
  • e2e_local: Local platform E2E tests
  • e2e_cloud: Cloud platform E2E tests (dry-run)
  • e2e_dataframe: DataFrame platform E2E tests
  • performance: Performance tests

Database Support

  • sqlite: Tests requiring SQLite
  • duckdb: Tests using DuckDB
  • database: Tests requiring database connections

Benchmark Types

  • tpch: TPC-H benchmark tests
  • tpcds: TPC-DS benchmark tests
  • ssb: Star Schema Benchmark tests
  • amplab: AMPLab Big Data Benchmark tests
  • clickbench: ClickBench tests
  • h2odb: H2O Database benchmark tests
  • primitives: Primitive operations tests

Example Usage

# Run only fast unit tests
make test-dev
# or
uv run -- python -m pytest -m "unit and fast"

# Run integration tests excluding slow ones
make test-integration
# or
uv run -- python -m pytest -m "integration and not slow"

# Run TPC-H related tests
make test-tpch
# or
uv run -- python -m pytest -m tpch

# Run all tests except memory intensive ones
uv run -- python -m pytest -m "not memory_intensive"

Parallel Run Mutual Exclusion (File Lock)

tests/conftest.py acquires a single inter-process file lock at ~/.benchbox/test.lock for the duration of each parallel test session. This prevents two concurrent pytest -n auto runs from fighting for CPU, which would otherwise double runtime and produce flaky timing assertions.

The default is intentionally global across worktrees: ten retained worktrees can otherwise start ten independent pytest -n auto runs and saturate a developer workstation. For isolated debugging or sandboxed CI, set BENCHBOX_TEST_LOCK_DIR=/path/to/dir; BenchBox will use $BENCHBOX_TEST_LOCK_DIR/test.lock instead. make test-unlock honors the same override.

Contract (tests/conftest.py:52-216):

  • Only the xdist controller process locks - workers do not. pytest_configure checks hasattr(config, "workerinput") to distinguish them.
  • Lock is skipped when numprocesses is 0/None (no -n auto) or when BENCHBOX_SKIP_TEST_LOCK=1 is set in the environment.
  • Lock path defaults to ~/.benchbox/test.lock; BENCHBOX_TEST_LOCK_DIR changes only the directory, not the filename.
  • Uses fcntl.flock(LOCK_EX | LOCK_NB) on POSIX, msvcrt.locking on Windows.
  • On contention, the second local run waits up to 3600 seconds with holder information and periodic progress messages. Set BENCHBOX_TEST_LOCK_WAIT_SECONDS=0 to fail immediately; CI does this. Ctrl-C cancels the wait without disturbing the holder.
  • The fd is held open for the whole session and released in pytest_unconfigure.

When to bypass: only for intentional concurrent debug runs. Set BENCHBOX_SKIP_TEST_LOCK=1 - but expect noisy timing results and CPU contention. Do not bypass in CI.

macOS-specific conftest branch (tests/conftest.py:109-122)

On macOS, xdist workers suppress setproctitle to avoid a launchservicesd CPU storm:

xdist calls setproctitle() twice per test (running/idle) via xdist.remote.worker_title(). At ~200 calls/second this triggers macOS launchservicesd to rebuild its process registry continuously, consuming 200%+ CPU and ~900 MB RSS - the actual root cause of the macOS beachball during parallel test runs.

The workaround walks the call stack to find the live xdist exec namespace and replaces worker_title with a no-op. This is necessary because import xdist.remote in an execnet __channelexec__ namespace resolves to a different module object than the one executing - so patching by module import alone is a no-op.

This branch only fires when sys.platform == "darwin" and the process is an xdist worker. Linux and Windows are unaffected.

CI/CD Integration

GitHub Actions Example

name: Test Suite
on: [push, pull_request]

jobs:
  test:
    runs-on: ubuntu-latest
    strategy:
      matrix:
        python-version: ['3.11', '3.12', '3.13', '3.14']

    steps:
    - uses: actions/checkout@v3
    - name: Set up Python ${{ matrix.python-version }}
      uses: actions/setup-python@v4
      with:
        python-version: ${{ matrix.python-version }}

    - name: Install dependencies
      run: |
        python -m pip install --upgrade pip
        pip install -e .[dev]

    - name: Run fast tests
      run: |
        uv run -- python tests/utilities/unified_test_runner.py --strategy ci

    - name: Upload coverage
      uses: codecov/codecov-action@v3
      with:
        file: ./coverage.xml

Jenkins Pipeline Example

pipeline {
    agent any

    stages {
        stage('Setup') {
            steps {
                sh 'python -m pip install --upgrade pip'
                sh 'pip install -e .[dev]'
            }
        }

        stage('Fast Tests') {
            steps {
                sh 'uv run -- python tests/utilities/unified_test_runner.py --strategy ci --parallel --workers 4 --output junit'
            }
            post {
                always {
                    junit 'test_results.xml'
                }
            }
        }

        stage('Integration Tests') {
            when {
                branch 'main'
            }
            steps {
                sh 'uv run -- python tests/utilities/unified_test_runner.py --mode integration --parallel --workers 2 --coverage'
            }
            post {
                always {
                    publishHTML([
                        allowMissing: false,
                        alwaysLinkToLastBuild: true,
                        keepAll: true,
                        reportDir: 'htmlcov',
                        reportFiles: 'index.html',
                        reportName: 'Coverage Report'
                    ])
                }
            }
        }
    }
}

Test Configuration

pytest.ini

The enhanced pytest.ini provides:

  • Comprehensive marker definitions
  • Parallel execution configuration
  • Coverage reporting setup
  • Performance optimization settings
  • CI/CD friendly defaults

conftest.py

Global test configuration including:

  • Database fixtures
  • Data generation helpers
  • Performance monitoring setup
  • Cleanup utilities

Best Practices

Writing Tests

  1. Use appropriate markers: Mark tests with relevant categories and characteristics
  2. Keep tests focused: Each test should verify one specific behavior
  3. Use fixtures: Leverage shared fixtures for database connections and test data
  4. Handle resources: Ensure proper cleanup of temporary files and connections
  5. Performance awareness: Use performance markers for resource-intensive tests

Test Organization

  1. Follow naming conventions: Use descriptive test names with test_ prefix
  2. Group related tests: Organize tests by functionality and component
  3. Use clear assertions: Make test failures easy to understand
  4. Document complex tests: Add docstrings for complex test scenarios

Performance Testing

  1. Establish baselines: Use the performance profiler to set baseline metrics
  2. Monitor regressions: Regularly check for performance regressions
  3. Profile selectively: Only profile tests when needed to avoid overhead
  4. Optimize test execution: Use caching and parallel execution for faster feedback

Troubleshooting

Common Issues

Slow Test Execution

# Profile test execution
uv run -- python tests/utilities/performance_profiler.py python -m pytest tests/unit/ -v

# Run only fast tests
make test-fast
# or
uv run -- python -m pytest -m fast

# Use parallel execution
make test-parallel
# or
uv run -- python -m pytest -n auto

Memory Issues

# Run memory-intensive tests separately
uv run -- python -m pytest -m "memory_intensive" --maxfail=1

# Monitor memory usage
uv run -- python tests/utilities/performance_profiler.py --output memory_report.md python -m pytest tests/unit/

Database Connection Issues

# Run database tests with verbose output
make test-integration
# or
uv run -- python -m pytest tests/integration/ -v -s

# Test database connectivity
uv run -- python -c "import duckdb; print(duckdb.connect().execute('SELECT 1').fetchone())"

Test Cache Management

# Clear pytest cache
uv run -- python -m pytest --cache-clear

# Clear custom test cache
uv run -- python tests/utilities/unified_test_runner.py --help

# Run with dry-run to see commands
uv run -- python tests/utilities/unified_test_runner.py --dry-run

Contributing

When adding new tests:

  1. Choose the right category: Place tests in the appropriate directory
  2. Add proper markers: Mark tests with relevant characteristics
  3. Update documentation: Add test descriptions to this README
  4. Consider performance: Mark resource-intensive tests appropriately
  5. Test your tests: Ensure new tests pass in isolation and with the full suite

Adding New Test Categories

  1. Create subdirectory in appropriate category
  2. Add __init__.py file
  3. Update markers in pytest.ini
  4. Add documentation to this README
  5. Update test runner configuration if needed

Performance Monitoring

The test suite includes comprehensive performance monitoring:

  • Execution time tracking: Monitor test duration trends
  • Memory usage analysis: Track memory consumption patterns
  • CPU utilization: Monitor CPU usage during test execution
  • I/O monitoring: Track file system operations
  • Regression detection: Automatically detect performance regressions

Performance data is stored in ~/.benchbox/test_cache/ and can be analyzed using the performance profiler utility.

Test Utilities

The test suite includes several unified utilities for efficient testing:

  • utilities/unified_test_runner.py: Comprehensive test runner with multiple execution strategies
  • utilities/benchmark_validator.py: Unified benchmark validation utility
  • utilities/performance_profiler.py: Performance monitoring and profiling
  • utilities/test_helpers.py: Common test helper functions
  • utilities/test_runner.py: Enhanced test runner with caching capabilities

These utilities provide a modern, efficient approach to testing and validation.

Support

For issues with the test suite:

  1. Check this README for common solutions
  2. Review test output for specific error messages
  3. Use the performance profiler to identify bottlenecks
  4. Consult the main BenchBox documentation
  5. Open an issue with detailed reproduction steps