When an interviewer asks you to design a web crawler, they’re testing how you turn a vague problem into a concrete, scalable system. The conversation usually starts with a high‑level description, then dives into details like politeness policies, duplicate detection, and fault tolerance. Below is a walkthrough you can use in a real interview, from gathering requirements to handling the toughest edge cases.

1. Clarify the scope first

Functional requirements

  • Accept a seed URL list and discover new URLs by following links.
  • Store fetched pages (HTML, metadata, outbound links) for later analysis.
  • Provide an API for querying the crawl status and retrieving stored content.

Non‑functional requirements

  • Throughput: The system should fetch many pages per second (the exact number depends on the target domain, but you can ask the interviewer for a ball‑park).
  • Politeness: Respect robots.txt and per‑host rate limits.
  • Scalability: Handle growth from a few thousand pages to billions.
  • Fault tolerance: Recover from node failures without losing progress.
  • Freshness: Optionally re‑crawl pages after a configurable interval.

Ask the interviewer which of these are most important for the scenario. A typical interview might focus on scalability and politeness, while freshness can be a follow‑up discussion.

2. Define the core API

A clean API helps you anchor the design and shows you think about the consumer.

POST /crawl/start
Body: { "seedUrls": ["https://example.com"], "maxDepth": 3 }

GET /crawl/status/{crawlId}
Response: { "queued": 1200, "inProgress": 30, "completed": 470 }

GET /page/{urlHash}
Response: { "url": "https://example.com", "html": "<html>…", "outLinks": ["https://example.com/about"] }

You can mention a simple authentication token if security is relevant, but keep the focus on the system internals.

3. High‑level component diagram

+-------------------+        +-------------------+        +-------------------+
|   Front‑End API   | <----> |   Crawl Manager   | <----> |   Scheduler Queue |
+-------------------+        +-------------------+        +-------------------+
                                 |       ^
                                 v       |
                        +-------------------+   |
                        |   Fetch Workers   |---+
                        +-------------------+   |
                                 |               |
                                 v               v
                        +-------------------+   +-------------------+
                        |   Parser Workers  |   |   Storage Service |
                        +-------------------+   +-------------------+
  • Crawl Manager: validates input, creates a crawl ID, and writes the seed URLs to the scheduler.
  • Scheduler Queue: holds URLs to be fetched, ordered by priority (e.g., breadth‑first, domain‑aware).
  • Fetch Workers: retrieve raw HTTP responses, obeying per‑host rate limits.
  • Parser Workers: extract links, compute a fingerprint (e.g., SHA‑256 of the URL) to detect duplicates, and push new URLs back to the scheduler.
  • Storage Service: persists page content, metadata, and link graph (often in a blob store + relational DB).

4. Deep dive: The hard parts

4.1 Politeness and rate limiting

Websites expect crawlers to limit requests per host. Implement a token bucket per domain in the scheduler. When a fetch worker asks for a URL, the scheduler checks the bucket; if empty, the URL is re‑queued with a delay. This design keeps the rate‑limit logic centralized and avoids pounding a single host.

4.2 Duplicate detection

Storing every URL ever seen quickly becomes expensive. Use a Bloom filter for a fast, memory‑efficient “maybe seen” check, backed by a persistent set for false positives. When a parser discovers a link, it first checks the Bloom filter; if the filter says “maybe new,” the URL’s hash is written to a durable key‑value store (e.g., a distributed hash table) before being enqueued.

4.3 Fault tolerance

  • Worker crashes: Use a message queue (e.g., Kafka) with at‑least‑once delivery semantics. If a fetch worker dies after pulling a URL, the message remains uncommitted and will be retried.
  • Scheduler node failure: Keep the queue replicated across multiple nodes; leader election (via Raft or etcd) ensures continuity.
  • Storage outages: Write ahead logs and periodic snapshots let you restore data without losing crawled pages.

4.4 Scaling the fetch layer

Fetching is I/O‑bound, so you can horizontally scale fetch workers. A common pattern is to run a pool of lightweight async processes (e.g., using asyncio or Go’s goroutines) that each maintain a few concurrent connections per host. The scheduler can assign a host affinity token so that workers reuse TCP connections, reducing latency.

4.5 Managing crawl depth and frontier explosion

Depth‑first can quickly blow up the frontier if the site has many outbound links. Enforce a maxDepth parameter: each URL carries a depth counter, and the parser discards links that exceed the limit. Additionally, you can prune low‑value URLs by scoring them (e.g., based on PageRank or content type) before adding them to the queue.

5. Trade‑offs to discuss

AspectSimple approachMore complex approach
SchedulerSingle FIFO queuePriority queue with per‑host buckets
Duplicate detectionStore every URL in DBBloom filter + persistent set
FetchingOne thread per workerAsync pool with connection reuse
StorageSingle relational DBBlob store + graph DB
  • Complexity vs. performance: A priority queue adds latency but gives better control over politeness and freshness.
  • Consistency: At‑least‑once delivery simplifies recovery but can cause duplicate fetches; you can de‑duplicate later using content hashes.
  • Cost: Using a distributed blob store reduces DB load but introduces eventual consistency concerns.

Bring up these trade‑offs when the interviewer asks “What would you change if the crawl grew to billions of pages?” and show that you can pivot the design without rewriting everything.

6. Typical follow‑up questions

  1. How would you handle robots.txt updates while a crawl is in progress?
    • Cache the parsed rules per domain and refresh the cache on a timer; workers consult the cache before each request.
  2. What if the crawl must respect a global bandwidth limit?
    • Introduce a global token bucket that fetch workers acquire before opening a connection.
  3. How do you ensure fresh content for high‑traffic sites?
    • Add a re‑crawl scheduler that periodically enqueues URLs whose last‑fetch timestamp exceeds a threshold.
  4. Can you support incremental crawls?
    • Store a fingerprint of page content (e.g., SHA‑256) and compare it on subsequent fetches; only store if the fingerprint changed.
  5. What monitoring metrics would you expose?
    • Queue length, fetch latency, error rate, per‑host request rate, and storage growth.

Answering these shows you think beyond the initial diagram.

7. Sample answer snippet (45‑90 seconds)

"The crawler starts with a seed list that the Crawl Manager writes to a scheduler queue. Workers pull URLs, respect per‑host rate limits enforced by token buckets, and fetch the page. The raw response goes to a Parser Worker, which extracts outbound links, checks a Bloom filter to avoid duplicates, and pushes new URLs back to the scheduler with an incremented depth counter. All pages and metadata are persisted in a blob store, while a relational table tracks the link graph. The design is horizontally scalable: you can add more fetch workers for throughput, and the scheduler is replicated for fault tolerance. Politeness, duplicate detection, and re‑crawl policies are handled centrally, making the system easy to reason about and extend."

Practicing this short narrative aloud helps you stay concise. Tools like Call Assistant can record your rehearsals and keep the follow‑up questions on topic, ensuring the story stays grounded in your own experience.

8. How to practice this

  1. Sketch the diagram on paper – spend 5 minutes drawing the components and their data flow without looking at any notes.
  2. Record a 60‑second pitch – use a phone or Call Assistant to capture yourself explaining the design; listen back for filler words and clarity.
  3. Answer three follow‑up questions – pick common variations (e.g., handling robots.txt, bandwidth limits, incremental crawls) and write a concise response for each, then rehearse them aloud.

By following this structure, you’ll demonstrate that you can turn an abstract crawling problem into a robust, production‑ready architecture while keeping the conversation focused and efficient.

Frequently asked questions

What is the difference between a crawler and a scraper?

A crawler discovers and fetches pages across the web, building a link graph. A scraper extracts specific data from a known page or set of pages. Crawlers handle breadth, while scrapers focus on depth.

How do token buckets enforce politeness?

Each domain gets a bucket that refills at a configured rate (e.g., one token per second). A fetch request consumes a token; if none are available, the request is delayed, ensuring the crawler does not exceed the allowed request rate.

Why use a Bloom filter for duplicate detection?

A Bloom filter provides a fast, memory‑efficient way to test if a URL has possibly been seen before. It reduces database lookups, and false positives are harmless because the URL will be stored only once after a secondary check.

Can a web crawler be fully real‑time?

Fully real‑time crawling is rare because politeness and network latency impose limits. However, you can approach near‑real‑time for a small set of high‑priority sites by tightening rate limits and using aggressive parallelism.

#system design#web crawler#architecture#interview#scalability#a web crawler