Data systemsUpstream2026
Checkpoint safety for change-data-capture in Apache SeaTunnel
Pinned down what a CDC checkpoint is allowed to contain when a fetcher fails mid-handoff.
Source
Split fetcher
Reads change events
Handoff
Record queue
Failure injected here
Emit
Emitter
Delivers records downstream
Recover
Checkpoint and restore
Offsets verified on replay
Context
Apache SeaTunnel moves data between a large catalog of databases and services, and its change-data-capture readers are what keep downstream copies consistent with source databases. A checkpoint that advances past records that were never delivered is silent data loss.
The problem
A reported MySQL CDC incident raised the question of whether a checkpoint could record an offset for records that had not yet been handed off when the fetcher failed. The shared reader had no direct test for that failure window, so nobody could answer with confidence.
My role
Investigated the failure window, designed the test harness, and worked the change through maintainer review.
Approach
- Built a reusable harness around the real reader, split fetcher, emitter, checkpoint lock, and serialized split state rather than mocks.
- Injected failures at exact boundaries: before handoff, during partial delivery, and during snapshot creation.
- Exercised real checkpoint-lock contention instead of timing sleeps, which is what usually makes concurrency tests lie.
- Restored state into a fresh reader to prove the remaining records replay after recovery.
Outcome
- The checkpoint ownership contract is now enforced by tests in the shared CDC base that SeaTunnel's CDC connectors build on.
- Verified that queued positions are not checkpointed as emitted, without changing runtime code or checkpoint formats.
- Merged after review by two maintainers, with the full cross-platform build matrix passing.