Kafka is a distributed, partitioned, replicated commit log. In plain terms, it stores streams of immutable records that producers append and consumers pull, preserving order within each partition. This simple model underpins many real‑time data pipelines, event‑driven architectures, and micro‑service communication.

Core Mechanics

Topics and Partitions

  • Topic: a logical name for a stream (e.g., order-events).
  • Partition: each topic is split into ordered logs; partitions enable parallelism and scaling.
  • Offset: a monotonically increasing number that identifies a record’s position within a partition.

Producers and Consumers

  • Producers write records to a topic, choosing a partition via a key or round‑robin.
  • Consumers pull records by offset. A consumer group ensures each partition is read by only one member, providing load‑balancing and fault tolerance.

Replication and Fault Tolerance

  • Each partition has a leader and one or more followers. Followers replicate the leader’s log.
  • If the leader fails, a follower is elected, guaranteeing no data loss beyond the last committed offset.

Retention Policies

  • Logs are retained by size or time (e.g., 7 days or 500 GB). After retention, older segments are deleted, freeing space while preserving recent data for replay.

Trade‑offs

AspectBenefitCost
ScalabilityHorizontal scaling by adding partitions.More partitions increase coordination overhead and can lead to uneven load if keys are skewed.
DurabilityReplication ensures data survives broker failures.Requires extra disk and network resources; latency grows with replication factor.
OrderingGuarantees order per partition.Global ordering across partitions is not provided; you must design keys carefully.
LatencyLow latency for high‑throughput streams.For low‑volume topics, latency can be higher due to batching and pull‑based consumption.
ComplexityRich ecosystem (Kafka Streams, Connect, kSQL).Operational complexity: monitoring broker health, managing offsets, handling rebalances.

Concrete Example

Imagine an e‑commerce site that needs to process orders in real time. When a user clicks "Buy", the front‑end sends an order‑created event to a Kafka topic orders. The topic has three partitions keyed by customer_id. A set of micro‑services—inventory, payment, shipping—each belong to a consumer group that reads from orders. Because each service reads the same offsets, they can replay events for debugging or reprocess after a bug fix. Replication across three brokers ensures the order data survives a broker crash.

Typical Interview Questions

  1. What is a Kafka topic and why is it partitioned?
    • Explain the logical stream, parallelism, and ordering guarantees per partition.
  2. How does Kafka achieve durability?
    • Discuss leader‑follower replication, ISR (in‑sync replicas), and acknowledgment settings (acks=all).
  3. What is the difference between at‑least‑once and exactly‑once semantics?
    • Clarify that Kafka provides at‑least‑once by default; exactly‑once requires idempotent producers and transactional APIs.
  4. How would you handle a consumer that falls behind?
    • Mention offset management, increasing fetch.max.bytes, adding more partitions, or using a separate consumer group for replay.
  5. When would you choose Kafka over a traditional message queue?
    • Highlight durability, replayability, and high throughput for stream processing.

60‑Second Spoken Answer

"Kafka is a distributed commit log that lets you publish streams of immutable records and have multiple consumers read them in order. Each topic is split into partitions, which are replicated across brokers for durability. Producers write to a partition based on a key, and consumers pull records by offset, forming consumer groups that balance load and provide fault tolerance. The trade‑offs are that you only get ordering per partition, latency can rise for low‑volume streams, and you need to manage broker health and rebalances. A common use case is an e‑commerce platform where order events are streamed to Kafka, and downstream services—inventory, payment, shipping—consume the same stream to stay in sync."

How to Practice This

  1. Write the answer on paper – keep it under 90 seconds, focus on definition, mechanism, trade‑offs, and a concrete example.
  2. Record yourself – play back the recording and note filler words or rambling sections.
  3. Run a mock interview with Call Assistant – let it listen, surface follow‑up prompts, and keep your story anchored to your resume.

FAQ

  • Q: Do I need to know the exact replication factor for every Kafka deployment? A: No. Explain the concept of replication for durability and mention that typical setups use three replicas, but the factor depends on reliability requirements.
  • Q: How does Kafka differ from RabbitMQ? A: Kafka stores a persistent log and enables replay, while RabbitMQ focuses on message routing and typically discards messages after acknowledgment.
  • Q: What does "consumer lag" mean? A: It’s the difference between the latest offset in a partition and the offset a consumer group has processed, indicating how far behind the consumer is.
  • Q: Can Kafka guarantee exactly‑once processing? A: Yes, but only when you use idempotent producers and the transactional API; otherwise, the default is at‑least‑once.

Frequently asked questions

Do I need to know the exact replication factor for every Kafka deployment?

No. Explain the concept of replication for durability and mention that typical setups use three replicas, but the factor depends on reliability requirements.

How does Kafka differ from RabbitMQ?

Kafka stores a persistent log and enables replay, while RabbitMQ focuses on message routing and typically discards messages after acknowledgment.

What does "consumer lag" mean?

It’s the difference between the latest offset in a partition and the offset a consumer group has processed, indicating how far behind the consumer is.

Can Kafka guarantee exactly‑once processing?

Yes, but only when you use idempotent producers and the transactional API; otherwise, the default is at‑least‑once.

#concept#Kafka#interview#streaming#distributed-systems