Skip to content

Latest commit

 

History

History
249 lines (194 loc) · 9.25 KB

File metadata and controls

249 lines (194 loc) · 9.25 KB

Platform Documentation

Documentation for database platforms supported by BenchBox.

Platform Guides

DataFrame Platforms (Native API)

BenchBox supports benchmarking DataFrame libraries using their native APIs instead of SQL. This enables direct performance comparison between SQL and DataFrame paradigms on identical workloads.

Available DataFrame Platforms

Platform CLI Name Family Status Documentation
Polars polars-df Expression Production-ready Polars DataFrame
Pandas pandas-df Pandas Production-ready Pandas DataFrame
Dask dask-df Pandas Production-ready Dask DataFrame
cuDF cudf-df Pandas Production-ready cuDF DataFrame
PySpark pyspark-df Expression Production-ready PySpark DataFrame
LakeSail lakesail-df Expression Production-ready LakeSail DataFrame
DataFusion datafusion-df Expression Production-ready DataFusion DataFrame
Databricks databricks-df Expression Production-ready Databricks DataFrame
# Quick start with DataFrame platforms
benchbox run --platform polars-df --benchmark tpch --scale 0.1    # Recommended - fast
benchbox run --platform pandas-df --benchmark tpch --scale 0.1    # Familiar API
benchbox run --platform dask-df --benchmark tpch --scale 0.1      # Distributed
benchbox run --platform cudf-df --benchmark tpch --scale 0.1      # GPU (Linux only)
benchbox run --platform pyspark-df --benchmark tpch --scale 0.1   # Spark ecosystem
benchbox run --platform lakesail-df --benchmark tpch --scale 0.1  # Sail (fast Spark)
benchbox run --platform datafusion-df --benchmark tpch --scale 0.1
benchbox run --platform databricks-df --benchmark tpch --scale 0.1  # Databricks Connect

# Compare SQL vs DataFrame on same workload
benchbox run --platform polars --benchmark tpch --scale 0.1       # SQL mode
benchbox run --platform polars-df --benchmark tpch --scale 0.1    # DataFrame mode

SQL Platforms

Core Local Databases

These platforms are included in the base BenchBox installation with no additional dependencies:

  • DuckDB - Embedded analytical database (default local platform)
  • SQLite - Embedded row-store database for lightweight testing

Local/Embedded Analytics Engines

Traditional Relational Databases

  • PostgreSQL - Open-source relational database with TimescaleDB support
  • CedarDB - High-performance HTAP engine over PostgreSQL wire protocol (formerly Umbra)

PostgreSQL Extensions

  • pg_duckdb - DuckDB-powered Postgres extension for vectorized OLAP
  • pg_mooncake - Columnar storage for Postgres with Iceberg and Delta Lake support
  • paradedb - Hybrid search and analytics Postgres extension (BM25 via pg_analytics)
  • citus - Distributed Postgres extension with sharded tables

Self-Hosted OLAP Databases

  • StarRocks - MPP columnar OLAP database (MySQL protocol + Stream Load)
  • Apache Doris - MPP real-time analytical database (MySQL protocol + Stream Load)
  • SingleStore - Distributed SQL with columnstore analytics (MySQL protocol + LOAD DATA)
  • QuestDB - Time-series database (PostgreSQL wire protocol + REST API)

Distributed SQL Engines

  • PrestoDB - Distributed SQL query engine (Facebook fork)
  • Trino - Distributed SQL query engine (community fork, formerly PrestoSQL)
  • Apache Spark - Unified analytics engine for large-scale data processing
  • LakeSail Sail - High-performance Spark-compatible engine (Spark Connect)
  • Apache Gluten + Velox - Spark plugin that offloads physical operators to a vectorized C++ engine (Linux-only, Docker on macOS/Windows)

Cloud Data Warehouses

GPU-Accelerated Platforms

  • CUDF - NVIDIA RAPIDS GPU-accelerated DataFrames

Platform Categories

By Installation Complexity

Zero Config (included in base install):

  • DuckDB, SQLite

Single Extra (one pip install command):

  • DataFusion, Polars, PostgreSQL, ClickHouse

Cloud SDK Required (authentication setup needed):

  • Databricks SQL, Databend, BigQuery, Redshift, Snowflake, Amazon Athena, Firebolt, Azure Synapse Analytics

Infrastructure Required (external cluster needed):

  • Trino, Presto, Spark, LakeSail, Gluten+Velox (Linux/Docker), StarRocks, Doris, QuestDB, ClickHouse (server mode)

By Use Case

Local Development & Testing:

  • DuckDB (recommended), SQLite, DataFusion, Polars

Production Benchmarking:

  • Databricks, Snowflake, BigQuery, Redshift

Self-Hosted Analytics:

  • ClickHouse, StarRocks, Doris, QuestDB, PostgreSQL, Trino, Presto, Spark, LakeSail, Gluten+Velox

GPU Workloads:

  • CUDF (NVIDIA GPUs required)

Quick Start by Platform

Local Platforms (No Setup Required)

# DuckDB - Default, included in base install
benchbox run --platform duckdb --benchmark tpch --scale 0.01

# SQLite - Included in base install
benchbox run --platform sqlite --benchmark tpch --scale 0.01

Cloud Platforms (Credentials Required)

# Databricks - Requires DATABRICKS_TOKEN and DATABRICKS_HOST
benchbox run --platform databricks --benchmark tpch --scale 1.0

# BigQuery - Requires GOOGLE_APPLICATION_CREDENTIALS
benchbox run --platform bigquery --benchmark tpch --scale 1.0

# Snowflake - Requires SNOWFLAKE_USER, SNOWFLAKE_PASSWORD, SNOWFLAKE_ACCOUNT
benchbox run --platform snowflake --benchmark tpch --scale 1.0

Future Platforms

Related Documentation

:maxdepth: 1

platform-selection-guide
quick-reference
comparison-matrix
support-status
deployment-modes
dataframe
duckdb
sqlite
polars
pandas-dataframe
modin-dataframe
dask-dataframe
cudf
pyspark-dataframe
datafusion-dataframe
databricks-dataframe
spark
lakesail
velox
velox_jar_setup
starrocks
doris
singlestore
questdb
databend
cedardb
clickhouse-local-mode
clickhouse-server
clickhouse-cloud
clickhouse-migration
workaround-index
postgresql
pg_duckdb
pg_mooncake
paradedb
citus
presto
trino
snowflake
databricks
bigquery
redshift
motherduck
ducklake
starburst
quanton
athena
athena-spark
firebolt
microsoft-fabric
fabric-lakehouse
azure-platforms
datafusion
influxdb
timescaledb
aws-glue
gcp-dataproc
dataproc-serverless
emr-serverless
fabric-spark
synapse-spark
snowpark-connect