Work

Research2026

EventFlowSentry: reproducible fault testing for streaming

Turns event-time failures in streaming pipelines into reproducible, explainable experiments.

  1. Model

    • Fault definition

      Late, reordered, replayed events

  2. Run

    • Pipeline under test

  3. Check

    • Metamorphic oracle

      Compare related runs

  4. Explain

    • Replayable trace

How an experiment runs.

Context

Streaming systems can pass ordinary tests and still produce different results under late data, reordering, replay, or recovery. Those are exactly the conditions production traffic creates.

The problem

Random chaos testing finds failures but rarely explains them, and a failure you cannot replay is hard to fix or trust. The research question is how to generate timing faults whose expected behavior stays checkable.

My role

Author. Designed the fault model, the metamorphic oracle, and the evaluation.

Approach

  • Represents event-time perturbations as explicit, repeatable experimental inputs.
  • Uses metamorphic relations to compare related runs when no single expected output exists.
  • Keeps the full execution conditions so every observed failure can be replayed and inspected.

Outcome

  • Published as a public preprint with a DOI.
  • Connects directly to the CDC and checkpointing work I do upstream.
Next storyA consumer-facing AI assistant on Amazon Bedrock AgentCore