Apache Software Foundation / SeaTunnel
Merged upstreamReliabilityMerged Oct 2, 2026

Always stop E2E engine containers during teardown

Centralized resilient E2E teardown so every engine container is stopped and the host mount cleanup is attempted even when in-container cleanup or an earlier stop fails.

apache/seatunnel · #12481

Test-infrastructure reliability

Failed or partially started Flink, Spark, and SeaTunnel containers can no longer abort teardown before sibling containers and host resources are cleaned up.

Problem

Flink, Spark, and SeaTunnel teardown executed an in-container rm before stopping engine containers. If a container never started or had already died, that command threw and skipped later stops and host cleanup. A leaked Flink JobManager could retain the shared jobmanager alias, attract the next test's TaskManager, and leave the new job waiting until timeout.

Approach

Adds a shared stopContainersAndDeleteVolume helper that treats in-container cleanup as best effort, stops each container independently, always attempts host-directory deletion, preserves the first real stop or delete failure with later failures suppressed, and restores interrupt status when required.

Impact and scope

  • Prevents one failed engine container from leaking its sibling containers into subsequent E2E cases.
  • Removes a shared-network alias collision that could turn one startup failure into a later, difficult-to-diagnose workflow hang.
  • Applies the same teardown guarantees to Flink, Spark, and the SeaTunnel Zeta engine through one reusable helper.
  • Keeps cleanup failures observable while distinguishing best-effort in-container removal from failures to stop containers or delete host resources.
  • Changes test infrastructure only; no production engine, connector, checkpoint, protocol, or user configuration behavior was modified.

Validation

  • All 51 seatunnel-e2e-common unit tests passed on JDK 8 and JDK 11, including not-running containers, cleanup exceptions, stop failures, suppressed failures, and interrupt restoration.
  • Real-container trials recovered from killed Spark, Zeta, and Flink containers and successfully ran the next job; the prior Flink path leaked its first JobManager in 11 of 11 runs, while the repaired revision completed teardown in 12 of 12 runs.
  • An Apache SeaTunnel member approved the final authored change, and all four current hosted checks pass. Linux host-directory ownership remains an observable follow-up rather than a claimed solved condition.