Back to blog

// OSSeva Blog

Migration

Apache Druid vs ClickHouse vs Apache Pinot: Real-Time Analytics Architectures Compared

Matt Reynolds8 min read

The short answer

Choose ClickHouse when you want a general-purpose analytical SQL database: wide SQL support including joins, a simpler server layout, and one system for ad hoc analysis, logs and dashboards. Choose Druid or Pinot when the workload is event data streamed from Kafka and queried by many users at once through an application, and you are willing to run a multi-service cluster for it. All three are columnar, all three are Apache 2.0, and all three are fast on the workloads they were designed for. The differences that matter in production are architecture, update handling and how much coordination infrastructure you have to operate.

Druid vs ClickHouse vs Pinot at a glance

Apache DruidClickHouseApache Pinot
Described asReal-time analytics database for slice-and-dice OLAPColumn-oriented SQL DBMS for OLAPDistributed OLAP database for real-time, user-facing analytics
LicenceApache 2.0Apache 2.0, plus a commercial cloud offeringApache 2.0
ServicesCoordinator, Overlord, Broker, Router, Historical, Middle Manager and Peon or IndexerClickHouse servers, plus Keeper for replicated tablesController, Broker, Server, Minion
StorageSegments in deep storage (for example S3 or HDFS), cached on HistoricalsMergeTree tables on each server, parts merged in the backgroundSegments assigned to servers, managed by Apache Helix
CoordinationZooKeeper by default, plus a metadata storeClickHouse Keeper or ZooKeeper, for replication onlyApache Helix on ZooKeeper
UpdatesStreaming inserts; updates through batch jobsAppend-first; update engines and UPDATE statementsImmutable segments; upserts for streaming entity data
QueryDruid SQL and native queries; joins and lookupsFull SQL with many join typesSQL

Apache Druid

Druid's documentation describes it as a real-time analytics database for fast slice-and-dice analytics on large data sets, and says it works best with event-oriented data. Its architecture splits work across services: Coordinators manage where segments live, Overlords assign ingestion tasks, Brokers fan queries out and merge results, Routers front the API and web console, Historicals serve committed data, and Middle Managers with Peons, or the Indexer, ingest new data. The docs suggest grouping them onto Master, Query and Data servers.

Data is stored as columnar segments, partitioned first by time, with Roaring or CONCISE bitmap indexes for filtering, approximate algorithms for count-distinct and quantiles, and optional rollup that pre-aggregates rows at ingest. Segments live in deep storage, which Druid does not provide itself, so you bring S3, HDFS or similar. The documentation is direct about the limits: Druid supports streaming inserts but not streaming updates, and it is a poor fit if you need low-latency updates of existing records by primary key. ZooKeeper is still the default for cluster state and leader election.

ClickHouse

ClickHouse calls itself a column-oriented SQL database management system for online analytical processing, available as open source and as a cloud service. Its core is the MergeTree family of table engines, built for high ingest rates and large volumes: each insert creates a part, and background merges combine parts, with the primary key setting sort order. Replicated tables keep their metadata in ClickHouse Keeper, or in ZooKeeper 3.4.5 or newer, though Keeper is recommended; unreplicated servers need neither.

ClickHouse is optimised for read-heavy analytics and append-only writes. Its guidance is to model updates and deletes as inserts where possible, using engines such as ReplacingMergeTree, and to use UPDATE or ALTER TABLE ... UPDATE when you must change rows in place. Its SQL is broad, with inner, outer, semi, anti and ASOF joins, which makes it easier to use for exploratory analysis than systems built around one large table per query. The project supports only its latest monthly releases and two LTS lines at a time, so clusters fall out of support quickly; our ClickHouse LTS upgrade post explains the cycle.

Apache Pinot

Pinot describes itself as a distributed OLAP database for real-time, user-facing use cases: many concurrent queries from an application, on data that is fresh from a stream. It runs controller, broker, server and minion nodes, uses Apache Helix for cluster management, and Helix keeps its state in ZooKeeper. Pinot assumes stored data is immutable, which simplifies storage and replication, but supports upserts on streaming entity data and background purges for privacy requests. Brokers and servers scale separately, so query volume and data size can grow independently.

How to choose

  • Choose ClickHouse for analysts and engineers running varied SQL, log and observability analytics, and dashboards over large tables, when you want the smallest number of moving parts and joins that behave like a relational database.
  • Choose Druid for time-series event data streamed from Kafka or loaded from object storage, with many dimensions, high-cardinality counting and rollup, served to analytical applications, when you already operate ZooKeeper and deep storage.
  • Choose Pinot for analytics embedded in a product, where end users run queries at high concurrency on fresh streaming data and you need upserts on entity records.

Operational cost belongs in the decision. Druid and Pinot clusters are several services plus ZooKeeper; ClickHouse is one server type plus Keeper for replication. All three need someone who understands segment or part management when things go wrong. If you are comparing ClickHouse with a transactional database instead, see ClickHouse vs PostgreSQL.

Druid vs ClickHouse performance

We do not publish speed figures, because every public comparison is tuned for one side. The architectures point to where each is strong. Druid's time partitioning, bitmap indexes and rollup favour filtered aggregations over time ranges. ClickHouse's sorted MergeTree parts favour wide aggregations and joins over large tables. Pinot is designed for low-latency queries at high concurrency. Test with your own data, query mix, concurrency and freshness requirements, and include ingestion and compaction load, not just query time.

The ZooKeeper underneath

Each of these systems can depend on ZooKeeper, and the ensemble is often older than the analytics engine above it. Druid uses it by default and older Druid releases bundle ZooKeeper 3.5.9, from a line that reached end of life on 1 June 2022. ClickHouse accepts any ZooKeeper from 3.4.5 for replicated tables, and Pinot depends on it through Helix. Our ClickHouse Keeper vs ZooKeeper post covers the move to Keeper.

Where OSSeva fits

OSSeva for Apache Druid ships patched, signed builds for archived Druid releases with the bundled ZooKeeper patched in the same build, and OSSeva Assure adds an upgrade plan across every major version between your release and the current one. OSSeva for ClickHouse ships patched builds for releases outside the community support window, patches the ZooKeeper ensemble under replicated tables, and plans the move to ClickHouse Keeper. For Pinot, see Apache Pinot support. Each is priced per cluster and sits under one contract with Kafka, ZooKeeper and PostgreSQL. See Apache Druid support options and ClickHouse support providers, or book a discovery call for a quote.

Frequently asked questions

Druid vs ClickHouse: which should I use?

ClickHouse for general analytical SQL, joins and a simpler deployment. Druid for streamed event data with time-based queries, rollup and high-cardinality counting served to applications. If your team writes ad hoc SQL all day, ClickHouse is usually the easier fit.

Druid vs ClickHouse performance: which is faster?

Each is fast on the workload it was built for: Druid on filtered aggregations over time with rollup and bitmap indexes, ClickHouse on large scans, aggregations and joins over sorted parts. No public benchmark settles it for your data, so test with your own queries and concurrency.

Druid vs Pinot: what is the difference?

Both are real-time OLAP systems with segment-based storage and several node types on ZooKeeper. Pinot is aimed at user-facing analytics with very high query concurrency and supports upserts on streaming data. Druid adds rollup at ingest and deep storage as the system of record, and supports streaming inserts but not streaming updates.

Can ClickHouse update rows?

Yes, but it is built for appends. ClickHouse recommends modelling updates as inserts with engines such as ReplacingMergeTree, and offers UPDATE and ALTER TABLE ... UPDATE when rows must change in place.

Does Druid still need ZooKeeper?

By default, yes, for cluster state and Coordinator and Overlord leader election. Running without it needs the Kubernetes extension, which Druid's documentation marks experimental.

Tags

Apache DruidClickHouseApache PinotAnalyticsComparison

Ready to get your open source under control?

Talk to an OSSeva engineer about CVE coverage, compliance, and migration support for your stack.