CDC using gRPC protocol EARLY ACCESS

Asynchronous replication of data changes (inserts, updates, and deletes) to external databases or applications

Change data capture (CDC) in YugabyteDB provides technology to ensure that any changes in data due to operations such as inserts, updates, and deletions are identified, captured, and made available for consumption by applications and other tools.

Architecture

Every YB-TServer has a CDC service that is stateless. The main APIs provided by the CDC service are the following:

  • createCDCSDKStream API for creating the stream on the database.
  • getChangesCDCSDK API that can be used by the client to get the latest set of changes.

Stateless CDC Service

See Change data capture for more details and limitations.

CDC streams

YugabyteDB automatically splits user tables into multiple shards (also called tablets) using either a hash- or range-based strategy. The primary key for each row in the table uniquely identifies the location of the tablet in the row.

Each tablet has its own WAL file. WAL is NOT in-memory, but it is disk persisted. Each WAL preserves the order in which transactions (or changes) happened. Hybrid TS, Operation ID, and additional metadata about the transaction is also preserved.

How does CDC work

YugabyteDB normally purges WAL segments after some period of time. This means that the connector does not have the complete history of all changes that have been made to the database. Therefore, when the connector first connects to a particular YugabyteDB database, it starts by performing a consistent snapshot of each of the database schemas.

The YugabyteDB Debezium connector captures row-level changes in the schemas of a YugabyteDB database. The first time it connects to a YugabyteDB cluster, the connector takes a consistent snapshot of all schemas. After that snapshot is complete, the connector continuously captures row-level changes that insert, update, and delete database content, and that were committed to a YugabyteDB database.

How does CDC work

The core primitive of CDC is the stream. Streams can be enabled and disabled on databases. You can specify which tables to include or exclude. Every change to a watched database table is emitted as a record in a configurable format to a configurable sink. Streams scale to any YugabyteDB cluster independent of its size and are designed to impact production traffic as little as possible.

EA In v2026.1.1.0 and later, you can use standard PostgreSQL replication slot commands with the yb_grpc output plugin to create gRPC CDC streams.

Creating a slot returns a user-chosen slot_name and a stream UUID (yb_stream_id in pg_replication_slots) for the connector.

You can also use the legacy yb-admin create_change_data_stream command to create streams.

Note

For v2026.1.1.0 and later, PostgreSQL replication slot syntax is recommended.

See Create a gRPC CDC stream.

You configure the maximum batch size in YugabyteDB, while the polling frequency is configured on the connector side.

Stream classification and metadata

A CDCSDK stream is treated as a gRPC stream when its replication slot plugin name is absent, empty, or yb_grpc; otherwise it is a logical replication stream.

gRPC streams carry different metadata depending on how (and when) they were created:

  • PostgreSQL syntax (yb_grpc): User-provided slot name, yb_grpc plugin, a replica_identity_map, and a cdc_state slot entry. Before-image format comes from per-table replica identity (no stream-level record_type).
  • yb-admin create_change_data_stream: Auto-generated slot name (grpc_<stream_id>), yb_grpc plugin, and a cdc_state slot entry. No replica_identity_map; before-image format still comes from the stream-level record_type passed to yb-admin.
  • Pre-existing gRPC streams (created in versions earlier than v2026.1.1.0): On master leader bringup after upgrading to v2026.1.1.0 or later, these streams are automatically backfilled to match the yb-admin shape above: auto-generated slot name (grpc_<stream_id>), yb_grpc plugin, and a cdc_state slot entry. They keep using record_type and do not receive a replica_identity_map.
Creation method Replication slot name Plugin replica_identity_map Slot entry in cdc_state Before-image source
PostgreSQL syntax (yb_grpc) User-provided yb_grpc Yes Yes Per-table replica identity
yb-admin create_change_data_stream (after finalization) grpc_<stream_id> yb_grpc No Yes Stream-level record_type
Pre-existing streams after backfill grpc_<stream_id> yb_grpc No Yes Stream-level record_type

Record-format discriminator: Logical replication streams and gRPC streams created via PostgreSQL syntax carry a replica_identity_map and no record_type option. gRPC streams created via yb-admin (and backfilled pre-existing streams) carry a record_type option and no replica_identity_map.

Connector tasks can consume changes from multiple tablets. At least once delivery is guaranteed. In turn, connector tasks write to the Kafka cluster, and tasks don't need to match Kafka partitions. Tasks can be independently scaled up or down.

The connector produces a change event for every row-level insert, update, and delete operation that was captured, and sends change event records for each table in a separate Kafka topic. Client applications read the Kafka topics that correspond to the database tables of interest, and can react to every row-level event they receive from those topics. For each table, the default behavior is that the connector streams all generated events to a separate Kafka topic for that table. Applications and services consume data change event records from that topic. All changes for a row (or rows in the same tablet) are received in the order in which they happened. A checkpoint per stream ID and tablet is updated in a state table after a successful write to Kafka brokers.

CDC guarantees

CDC in YugabyteDB provides technology to ensure that any changes in data due to operations (such as inserts, updates, and deletions) are identified, captured, and automatically applied to another data repository instance, or made available for consumption by applications and other tools. CDC provides the following guarantees.

Per-tablet ordered delivery

All data changes for one row or multiple rows in the same tablet are received in the order in which they occur. Due to the distributed nature of the problem, however, gRPC replication does not guarantee order across tablets.

Consider the following scenario:

  • Two rows are being updated concurrently.
  • These two rows belong to different tablets.
  • The first row row #1 was updated at time t1, and the second row row #2 was updated at time t2.

In this case, it is possible for CDC to push the later update corresponding to row #2 change to Kafka before pushing the earlier update, corresponding to row #1.

At-least-once delivery

Updates for rows are pushed at least once. With the at-least-once delivery, you never lose a message, however the message might be delivered to a CDC consumer more than once. This can happen in case of a tablet leader change, where the old leader already pushed changes to Kafka, but the latest pushed op id was not updated in the CDC metadata.

For example, a CDC client has received changes for a row at times t1 and t3. It is possible for the client to receive those updates again.

No gaps in change stream

When you have received a change for a row for timestamp t, you do not receive a previously unseen change for that row from an earlier timestamp. This guarantees that receiving any change implies that all earlier changes have been received for a row.