Three Streaming Engines, Three Answers: Spark, Flink, and Kafka Streams
The choice is rarely about features. It is about the latency you actually need, how much state you will carry, and whether you want a cluster to operate.
The choice is rarely about features. It is about the latency you actually need, how much state you will carry, and whether you want a cluster to operate.
Checkpoints belong to the engine and exist for failure. A savepoint belongs to you and exists for the planned changes that failure recovery cannot help with.
Keeping state on the heap is fast until it is too big. Keeping it on disk is unbounded until every lookup costs a read. The choice is made once and felt daily.