When interviewers ask about batch versus stream processing, they want to see that you understand the core difference, can articulate the trade‑offs, and can map the concepts to real‑world systems.
One‑Sentence Definitions
- Batch processing: Running a job on a finite, collected data set at a scheduled time.
- Stream processing: Continuously handling an unbounded flow of events as they arrive.
How Each Mechanism Works
Batch
- Ingestion – Data is collected (e.g., nightly logs, daily sales CSVs).
- Storage – Files land in a data lake or a relational database.
- Job Execution – A scheduler (cron, Airflow) launches a job that reads the whole set, transforms it, and writes the result.
- Output – Results are stored for downstream reporting or analytics.
Typical tools: Hadoop MapReduce, Spark in batch mode, Snowflake, traditional ETL pipelines.
Stream
- Ingestion – Events are emitted by producers (web clicks, sensor readings) to a broker.
- Processing – A streaming engine pulls events, applies operators (filter, window, join) in near‑real‑time.
- State Management – Windows or keyed state keep track of aggregates.
- Output – Processed events are written to dashboards, alerts, or downstream services.
Typical tools: Apache Kafka + Flink, Kinesis + Lambda, Spark Structured Streaming, Pulsar.
Trade‑offs
| Aspect | Batch | Stream |
|---|---|---|
| Latency | Minutes‑to‑hours; acceptable for reports that don’t need instant freshness. | Milliseconds‑to‑seconds; needed for real‑time alerts or personalization. |
| Complexity | Simpler pipelines, easier debugging; jobs run in isolation. | Requires handling out‑of‑order events, state checkpoints, and exactly‑once semantics. |
| Throughput | Very high; can process terabytes in a single job because it can batch I/O. | High but bounded by per‑event processing time; back‑pressure handling is crucial. |
| Fault tolerance | Retry whole job; easier to recover from failures. | Requires checkpointing, replay, and idempotent operators. |
| Use case fit | End‑of‑day financial statements, nightly data warehouse loads. | Fraud detection, live recommendation, monitoring dashboards. |
Concrete Example
Imagine an e‑commerce platform:
- Batch: Every night, a Spark job reads the day’s order CSVs from S3, aggregates total sales per product, and writes the summary to a Redshift table for the business intelligence team.
- Stream: As each order is placed, a Kafka event is produced. A Flink job updates a running count of items sold per product and pushes the latest numbers to a real‑time dashboard that drives dynamic pricing.
Both pipelines may use the same raw data source, but the batch job gives a clean, audited snapshot, while the stream job provides instant visibility.
Typical Interview Questions
- When would you choose batch over stream, and vice‑versa?
- Highlight latency requirements, data volume, and operational overhead.
- How do you handle late‑arriving data in a stream?
- Explain event‑time windows, watermarks, and allowed lateness.
- What guarantees does your streaming platform provide?
- Discuss at‑least‑once vs exactly‑once semantics and checkpointing.
- How do you ensure consistency between batch and stream results?
- Mention the “lambda architecture” pattern or reconciliation jobs.
- What are the cost implications of each approach?
- Batch often uses spot or low‑priority compute; stream may need always‑on resources.
60‑Second Spoken Answer
"Batch processing is about running a job on a bounded data set at a scheduled interval—think of nightly sales reports that read a day’s CSV files, aggregate them, and store the results. Stream processing, on the other hand, continuously consumes an unbounded flow of events, applying transformations in near‑real time—for example, updating a live dashboard as each order arrives.
The trade‑off is latency versus complexity. Batch can handle huge volumes with simple, fault‑tolerant jobs, but the results are delayed. Stream gives you sub‑second insights, but you need to manage state, out‑of‑order events, and exactly‑once guarantees. In practice, you pick batch when the business can wait for a nightly snapshot, and stream when you need immediate reactions, like fraud detection."
How to Practice This
- Write the answer on paper – Keep it under 90 seconds; count words to gauge length.
- Record yourself – Use Call Assistant to capture the audio, then replay to check pacing and clarity.
- Tie it to your résumé – Pick a project where you built either a batch pipeline or a streaming job, and rehearse linking the concept to that experience.
FAQ
- Q: Can a system be both batch and streaming? A: Yes. Many architectures combine both, using batch for historical analysis and streaming for real‑time alerts, often called a hybrid or lambda architecture.
- Q: What is the main challenge of stream processing? A: Managing state and handling out‑of‑order events while guaranteeing correct results under failure conditions.
- Q: Do modern tools blur the line between batch and stream? A: Tools like Spark Structured Streaming let you write code that can run in either mode, but the underlying semantics—bounded vs unbounded data—still dictate latency and complexity.
- Q: How do you test a streaming job? A: Use deterministic test data, simulate late events, and verify that checkpoints restore the correct state after a failure.
Frequently asked questions
Can a system be both batch and streaming?
Yes. Many architectures combine both, using batch for historical analysis and streaming for real‑time alerts, often called a hybrid or lambda architecture.
What is the main challenge of stream processing?
Managing state and handling out‑of‑order events while guaranteeing correct results under failure conditions.
Do modern tools blur the line between batch and stream?
Tools like Spark Structured Streaming let you write code that can run in either mode, but the underlying semantics—bounded vs unbounded data—still dictate latency and complexity.
How do you test a streaming job?
Use deterministic test data, simulate late events, and verify that checkpoints restore the correct state after a failure.
#concept#batch vs stream processing#interview#data engineering#real-time