Overview
chDB is described as a fast, reliable and scalable in-process database that brings the ClickHouse engine into an application process. It is an open-source library, so users can customise and extend it, and it is maintained alongside a community that publishes updates and security practices. ClickHouse positions chDB within its wider ecosystem for local development, in-process analytics embedded in an application, and production workloads.
- Open-source in-process database library
- Blazing fast SQL engine
- Supports 80+ data formats
- Zero-copy data exchange with DataFrames
- Native Pandas-like Pythonic interface
Python integration
chDB is available as an in-process OLAP SQL engine for Python, letting developers use the ClickHouse engine directly in Python code. It supports the Python DB API 2.0, so it integrates with familiar Python libraries and tools. Native objects in the chosen programming language can be queried directly, which reduces latency and simplifies data processing. Beyond Python, chDB provides bindings for a number of other programming languages.
- In-process OLAP SQL engine for Python
- Python DB API 2.0 support
- Direct querying of native language objects
- Bindings for multiple programming languages
chDB 4.0 and the DataStore API
chDB 4.0 introduces a DataStore API that allows Pandas code to execute on the ClickHouse engine. Operations are recorded rather than executed immediately; the full pipeline is compiled into optimised ClickHouse SQL at execution time and run on the vectorised, multi-threaded engine. The compilation step applies filter pushdown, column pruning and limit propagation, and the API provides seamless engine fallback and a one-import migration path. chDB 4.0 can be run in Hex notebooks with no setup, with an extended free trial offered through ClickHouse.
- Lazy execution pipeline compiled to optimised SQL
- Filter pushdown, column pruning and limit propagation
- Seamless engine fallback
- One import migration
- Available in Hex notebooks
Embedded deployment and data exchange
chDB runs embedded, with no requirement to install or run ClickHouse services, which streamlines deployment and reduces system complexity and resource usage. This makes it suited to lightweight and embedded applications. DataFrames are read and written directly with no serialisation cost, and numeric and fixed-width columns share memory buffers with Pandas via the buffer protocol. Input and output cover Parquet, CSV, JSON, Arrow, ORC and more than 80 further formats.
- No ClickHouse service installation required
- Zero-copy DataFrame read and write
- Shared memory buffers with Pandas via the buffer protocol
- Parquet, CSV, JSON, Arrow, ORC and 80+ formats
Use in machine learning and data science
ClickHouse references chDB on its machine learning and GenAI use-case page as the way to access the full power of ClickHouse within Python code. The page presents chDB alongside ClickHouse capabilities for data preparation aggregations, vector search and inference through user-defined functions. chDB also features in the ClickPy demo, which analyses Python package download data from PyPI.
- Listed on the ClickHouse machine learning and GenAI use-case page
- Referenced in the ClickPy Python package analytics demo
Sources
-
chDB - a fast, reliable, and scalable in-process database | ClickHouse https://clickhouse.com/chdb Verified 02 Oct 2026
-
Machine learning and GenAI with ClickHouse | ClickHouse for ML and data science | ClickHouse https://clickhouse.com/use-cases/machine-learning-and-data-science Verified 17 Sep 2026
Last verified 17 Sep 2026. This entry is compiled from the public web pages listed above. Nothing here is stated that those pages do not, and each of them was read on the date shown.