When an interview asks you to design a fraud detection pipeline, the conversation is less about memorising a diagram and more about showing how you reason about data, latency, and risk. You’ll need to walk the interviewer through requirements, sketch a high‑level architecture, dive into the hardest pieces, and be ready to discuss trade‑offs. Below is a practical walkthrough you can adapt on the fly.
1. Clarify Functional Requirements
Start by confirming what the system must do. Typical functional specs include:
- Real‑time scoring: Incoming transactions must be evaluated within a few hundred milliseconds.
- Batch enrichment: Periodic jobs that recompute risk scores using more expensive models.
- Alerting & action: The pipeline should emit a decision (allow, review, block) and optionally trigger downstream workflows.
- Feedback loop: Outcomes (e.g., chargeback, manual review) feed back into model training.
- Auditability: Every decision must be traceable to the input data and model version.
Ask clarifying questions: expected transaction volume, tolerance for false positives, compliance constraints, and whether the system will serve multiple product lines.
2. Define Non‑Functional Requirements
Non‑functional goals shape the design more than the functional ones:
| Requirement | Typical Target | Why It Matters |
|---|---|---|
| Latency | < 200 ms for real‑time scoring | Users expect instant approval/decline. |
| Throughput | Tens of thousands of TPS (adjust per volume) | System must handle peak load without queuing. |
| Availability | 99.9 %+ (or higher for financial services) | Downtime directly impacts revenue. |
| Scalability | Horizontal scaling across regions | Fraud patterns differ globally. |
| Data durability | 24‑48 h retention for audit logs | Regulators require traceability. |
| Model freshness | Daily or hourly retraining | New fraud tactics emerge quickly. |
Mention that exact numbers can be tuned later; the interview is about the reasoning.
3. Identify Core Entities & Data Flow
A fraud detection pipeline typically revolves around these entities:
- Transaction – raw event (amount, merchant, device, timestamps).
- Feature vector – enriched representation (user history, device risk, geo‑velocity).
- Model version – the binary or parameters used for scoring.
- Score – numeric risk output (e.g., 0‑1).
- Decision – business rule applied to the score (allow/review/block).
- Feedback – label from downstream (chargeback, manual review outcome).
The flow is:
[Ingestion] → [Feature Store] → [Scoring Service] → [Decision Engine] → [Alert/Action]
↑ ↓
[Feedback Loop] ←───────────────────────────────
4. High‑Level Architecture (Text Diagram)
+----------------+ +----------------+ +-------------------+
| Ingestion | ---> | Feature Store | ---> | Scoring Service |
| (Kafka/REST) | | (Redis/BigTable) | | (TensorFlow/ONNX) |
+----------------+ +----------------+ +-------------------+
|
v
+-------------------+
| Decision Engine |
| (Rule + Model) |
+-------------------+
|
v
+-------------------+
| Alert / Action |
| (Pub/Sub, API) |
+-------------------+
|
v
+-------------------+
| Feedback Service |
| (Batch jobs) |
+-------------------+
Key Components
- Ingestion: Use a durable log (Kafka, Pulsar) to guarantee at‑least‑once delivery. A thin API layer can accept HTTP/HTTPS calls for legacy systems.
- Feature Store: A low‑latency key‑value store (e.g., Redis) holds recent user/device aggregates; a cold store (e.g., BigQuery) keeps historic data for batch enrichment.
- Scoring Service: Deploy models as stateless containers behind a load balancer. Keep model versions in a central registry so you can roll back.
- Decision Engine: Combine the model score with business rules (thresholds, velocity checks). This layer is where you enforce compliance limits.
- Alert / Action: Publish decisions to downstream systems – fraud ops dashboards, payment gateways, or a webhook for third‑party risk platforms.
- Feedback Service: Periodic jobs ingest chargeback or manual review outcomes, update feature aggregates, and trigger model retraining.
5. Deep Dive: Real‑Time Scoring
Challenges
- Cold‑start: New users have no history. Mitigate with generic device risk scores or fallback heuristics.
- Model latency: Complex neural nets can exceed the 200 ms budget. Use model quantization, compiled inference (e.g., ONNX Runtime), or a two‑stage approach: a lightweight rule filter first, then a heavy model for borderline cases.
- Feature freshness: Features must reflect the last few minutes of activity. Cache recent aggregates in an in‑memory store and invalidate on new events.
Design Choices
- Stateless containers: Allows horizontal scaling; each request carries the transaction ID and looks up needed features.
- Batch vs. streaming: For ultra‑low latency, stream features directly from the ingestion topic; for richer context, pull from the feature store asynchronously.
- Model versioning: Store model binaries in an object store (S3‑compatible) and reference them via a hash. The scoring service reads the hash at startup and can hot‑swap on a signal.
6. Deep Dive: Feedback Loop & Model Retraining
Why It Matters
Without a feedback loop, the system drifts as fraudsters adapt. The loop closes the circle: labeled outcomes → feature updates → model retraining → deployment.
Implementation Sketch
- Collect labels: Export decisions and outcomes nightly to a data lake.
- Feature engineering: Join raw transactions with historic aggregates to build a training set.
- Training: Use a distributed ML platform (Spark, Flink) to train a gradient‑boosted tree or deep model.
- Evaluation: Compare ROC‑AUC against the current production model; only promote if improvement exceeds a threshold.
- Deployment: Store the new model version; trigger a rolling rollout of the scoring service.
7. Trade‑offs & Follow‑Up Questions
Interviewers love to probe the edges. Be ready to discuss:
- Consistency vs. latency: Strong consistency (e.g., reading from a relational store) adds latency. Explain why eventual consistency is acceptable for risk scores.
- Cold‑storage cost: Storing every transaction indefinitely is expensive. Propose tiered storage (hot for recent, cold for compliance‑required retention).
- Model explainability: Some regulators require a rationale for a block. Show how you can surface feature importance from a tree‑based model.
- Fail‑over handling: If the scoring service crashes, the pipeline should fall back to a rule‑only mode to maintain availability.
- Multi‑region deployment: Discuss data locality, cross‑region replication, and latency impact for global merchants.
8. Sample Answer (45‑90 seconds)
"Sure, let me outline a fraud detection pipeline. First, we ingest every transaction into a durable log like Kafka. A feature store, backed by Redis for hot data and a data warehouse for historic aggregates, enriches each event with user‑level risk signals. The enriched vector is sent to a stateless scoring service that loads the latest model from an object store; we keep latency under 200 ms by using a compiled runtime and caching recent features. The score feeds a decision engine that applies business thresholds and compliance rules, then publishes an allow, review, or block decision via Pub/Sub to downstream systems. Finally, we close the loop: outcomes such as chargebacks are batched nightly, used to retrain the model, and the new version rolls out gradually. This architecture balances real‑time latency, horizontal scalability, and auditability while allowing us to evolve the model as fraud tactics change."
9. How to Practice This
- Sketch the flow on a whiteboard: Start with ingestion and add one component at a time. Explain each addition aloud.
- Run a mock interview: Use Call Assistant to record your answer, then listen back to ensure you stay within the time budget and hit the key points.
- Flip the script: After presenting, ask yourself follow‑up questions about latency, model drift, and data consistency. Write short bullet‑point answers for each.
FAQ
- What if the interviewer asks for exact latency numbers?
- Explain that you’d target sub‑200 ms for real‑time scoring, citing typical industry expectations, and note that the exact figure depends on transaction volume and model complexity.
- How do you handle a sudden spike in traffic?
- Mention auto‑scaling of the scoring containers, back‑pressure on the ingestion topic, and a fallback rule‑only mode to keep the system responsive.
- Can you use a single database for both features and logs?
- Generally not advisable; separating hot feature lookups (key‑value store) from immutable logs (append‑only log) reduces contention and improves latency.
- What if the model produces too many false positives?
- Discuss adjusting thresholds, adding a secondary review queue, and incorporating more nuanced features to improve precision without sacrificing recall.
Frequently asked questions
What if the interview asks for exact latency numbers?
State that you’d aim for sub‑200 ms end‑to‑end scoring, which is typical for payment‑oriented fraud systems, and explain that the exact target depends on transaction volume and model complexity.
How do you handle a sudden traffic spike?
Describe auto‑scaling of stateless scoring containers, back‑pressure on the ingestion queue, and a rule‑only fallback path that keeps decisions fast even if the model layer is overloaded.
Can one database serve both feature storage and transaction logs?
Usually not; a key‑value store for hot feature lookups and an append‑only log for immutable transaction records keep read/write patterns separate and preserve low latency.
What if the model generates too many false positives?
Adjust decision thresholds, introduce a secondary manual review queue, and enrich the feature set to improve precision while maintaining recall.
#system design#fraud detection#pipeline#architecture#interview#a fraud detection pipeline