Feature flags have become a staple of modern continuous‑delivery pipelines. In a system‑design interview you’ll be expected to discuss both the business purpose and the technical choices that make a flag service reliable at scale.

1. Clarify Functional Requirements

  • Create / Update flag – Define a flag name, default value, and optional targeting rules (e.g., user segment, environment).
  • Read flag – Client libraries query the service to decide which code path to execute.
  • Rollout control – Ability to enable a flag for a percentage of users or for specific cohorts.
  • Versioning / Auditing – Keep a history of changes for compliance and rollback.
  • Admin UI – Dashboard for product owners to toggle flags and view metrics.

Non‑functional Requirements

RequirementWhy it mattersTypical target
Low latency readsClient code runs on the request path; extra milliseconds hurt response time.< 10 ms per read
Strong consistency for writesA flag change must be visible to all clients quickly to avoid split‑brain bugs.Immediate visibility
High availabilityFeature flags often control safety‑critical features.99.9 %+ uptime
Horizontal scalabilityLarge SaaS products may have millions of flag checks per second.Scale‑out across nodes
Security & RBACOnly authorized engineers should modify flags.Role‑based policies

2. Core Data Model

CREATE TABLE flags (
    flag_id   UUID PRIMARY KEY,
    name      TEXT UNIQUE NOT NULL,
    description TEXT,
    created_at TIMESTAMP NOT NULL DEFAULT now()
);

CREATE TABLE flag_versions (
    version_id UUID PRIMARY KEY,
    flag_id    UUID REFERENCES flags(flag_id),
    value      BOOLEAN NOT NULL,
    rollout    JSONB,            -- e.g., {"percentage":20,"segment":"beta"}
    created_by TEXT NOT NULL,
    created_at TIMESTAMP NOT NULL DEFAULT now()
);

flags stores the identity; flag_versions records each change. The JSON column lets you store arbitrary targeting rules without a rigid schema.

3. Public API (REST + gRPC)

MethodPathPurpose
POST /flagsCreate a new flag
PATCH /flags/{id}Update value or rollout rules
GET /flags/{name}Retrieve the latest version
GET /flags/{name}/historyAudit trail
POST /flags/batchBulk read for client SDKs

A lightweight gRPC endpoint (GetFlag) is often used by SDKs to reduce overhead.

4. High‑Level Architecture

+----------------+      +-------------------+      +-------------------+
|  Client SDKs   |<---->|  API Gateway/    |<---->|  Write‑through    |
| (Node, Java,   |      |  Load Balancer    |      |  Cache (Redis)    |
|  Go, Python)   |      +-------------------+      +-------------------+
+----------------+                |                         |
                                 |                         |
                                 v                         v
                        +-------------------+   +-------------------+
                        |  Relational DB    |   |  Pub/Sub (Kafka)  |
                        |  (Postgres)       |   |  for invalidation |
                        +-------------------+   +-------------------+
  • Write path – API writes to Postgres, then pushes an invalidation event to Kafka. The cache updates synchronously (write‑through) and also listens to the topic for eventual consistency across nodes.
  • Read path – SDK first checks Redis; a miss falls back to the DB. Cache TTL is short (seconds) because flags change frequently.
  • Admin UI – Communicates through the same API, adding audit logging.

5. Deep Dive: Consistency & Propagation

Problem

When a flag is toggled, every running instance must see the new value quickly, otherwise you risk a mixed state where some users see the new feature and others do not.

Solution Options

  1. Push‑based invalidation – After a write, publish a message to a topic. All cache nodes subscribe and evict the key. Latency is typically a few hundred milliseconds.
  2. Polling – SDKs poll the API at a configurable interval (e.g., 30 s). Simpler but slower to converge.
  3. Hybrid – Use push for critical flags and polling for less‑sensitive ones.

Trade‑off Discussion

Push gives faster convergence but adds operational complexity (message ordering, retries). Polling is robust but may cause brief inconsistencies. In an interview, argue that a push‑based approach is preferred for safety‑critical flags, while a fallback poll mitigates missed messages.

6. Deep Dive: Targeting Rules Engine

Flags often need to evaluate expressions like:

(user.country == "US" && user.plan == "premium") || (user.id % 100 < 5)

A naïve approach stores the rule as a string and evaluates it at runtime, which can be slow and insecure.

Design pattern:

  • Parse the rule into an abstract syntax tree (AST) when the flag is saved.
  • Serialize the AST into a compact binary format (e.g., Protocol Buffers).
  • At read time, the SDK loads the pre‑compiled AST and runs a fast interpreter.

This reduces per‑request CPU and allows you to enforce a whitelist of operators, preventing injection attacks.

7. Scaling Considerations

  • Sharding – Partition flags by name hash across multiple DB instances if the flag count grows into the hundreds of thousands.
  • Read‑heavy workload – Deploy read‑replicas and configure the cache to prefer them for cache‑miss fallback.
  • Multi‑region – Replicate the cache and DB per region; use a global Kafka cluster to propagate changes.
  • Rate limiting – Protect the read endpoint with token buckets to avoid denial‑of‑service from runaway SDKs.

8. Common Follow‑Up Questions

  1. How would you handle versioned rollouts (e.g., canary)? Answer: Store rollout percentages in the rollout JSON field and let the SDK evaluate the user’s bucket. The service can expose a GET /flags/{name}/evaluate?userId=… endpoint for server‑side evaluation.
  2. What if a client is offline for a long period? Answer: SDKs cache the last known flag state locally and mark it stale after a TTL. When connectivity returns, they refresh via the batch API.
  3. How do you ensure auditability? Answer: Write every PATCH to an immutable flag_audit table and expose it via the history endpoint. Optionally stream audit events to a log aggregation system.
  4. Can you support A/B test metrics? Answer: Integrate with an analytics pipeline; the SDK can emit an event when a flag is evaluated, tagging the experiment ID.
  5. What if a flag change causes a cascade of failures? Answer: Implement a “kill‑switch” that forces all flags to a safe default, stored in a separate fast‑path table that bypasses normal routing.

9. Sample Answer Template (45‑90 seconds)

"The core of a feature‑flag service is a low‑latency read path backed by a write‑through cache. I’d store flags in a relational table with a versioned history, and expose a simple REST API for create, update, and read. Writes go to Postgres, then publish an invalidation event on Kafka so every cache node evicts the stale entry. Clients read from Redis first; a miss falls back to the DB. For targeting rules I’d parse the expression into an AST at save‑time, serialize it, and evaluate it in the SDK to keep per‑request cost low. The main trade‑off is push‑based invalidation versus polling – I’d choose push for safety‑critical flags and fall back to a short poll interval for resilience. Scaling is handled by sharding flags by name hash, adding read replicas, and replicating the cache per region. Auditing is provided by an immutable audit table and a history endpoint."

10. How to practice this

  1. Sketch the diagram on paper – Start from the API gateway and add cache, DB, and pub/sub layers. Explain each arrow in a sentence.
  2. Run a mock interview – Use Call Assistant to record your answer, then replay it and check that you stay within the 45‑90 second window.
  3. Implement a tiny prototype – Build a REST endpoint that writes to SQLite and publishes to a local Redis channel; measure read latency and iterate on the cache strategy.

FAQ

  • What is the difference between a feature flag and a toggle? A feature flag is a broader concept that can include targeting rules, rollouts, and versioning; a toggle is a simple on/off switch without granularity.
  • Do I need a separate service for flags, or can I embed them in my application? Embedding works for very small teams, but a dedicated service provides central governance, auditability, and the ability to change flags without redeploying each service.
  • How do I guarantee that a flag change is visible to all instances instantly? By using a push‑based invalidation mechanism (e.g., Kafka) combined with a write‑through cache, you can achieve sub‑second propagation in most environments.
  • What security considerations are important for a flag service? Enforce role‑based access control on the API, encrypt data in transit, and log every change for compliance. Additionally, validate targeting expressions to prevent injection attacks.

Frequently asked questions

What is the difference between a feature flag and a toggle?

A feature flag can carry targeting rules, rollout percentages, and version history, while a toggle is a simple binary switch without granularity.

Do I need a separate service for flags, or can I embed them in my application?

Embedding works for tiny projects, but a dedicated service gives central governance, audit trails, and the ability to change flags without redeploying each microservice.

How do I guarantee that a flag change is visible to all instances instantly?

Use a push‑based invalidation channel (e.g., Kafka) that notifies every cache node after a write, achieving sub‑second propagation in most setups.

What security considerations are important for a flag service?

Enforce role‑based access control on the API, encrypt traffic, log every modification, and validate targeting expressions to avoid injection attacks.

#system design#feature flag#architecture#scalability#interview#a feature flag service