Observability is a frequent topic in modern engineering interviews. Whether you’re a fresh graduate or a seasoned site reliability engineer, you’ll be asked to define the concept, compare tools, and illustrate how you’ve turned data into action. Below is a practical cheat‑sheet: the questions you’ll hear, a crisp answer you can deliver in 45‑90 seconds, and the follow‑up the interviewer is likely to probe.

1. What is observability and how does it differ from monitoring?

Observability is the ability to infer the internal state of a system solely from its external outputs—metrics, logs, and traces. Monitoring, by contrast, is the practice of watching predefined health indicators and alerting when they cross thresholds. In other words, observability lets you ask new questions about a running system, while monitoring lets you answer known questions.

Typical follow‑up: Can you give an example of a time you discovered a problem only because you had observability, not just monitoring?


2. Describe the three pillars of observability.

  1. Metrics – numeric time‑series data (e.g., request latency, CPU usage). They are great for spotting trends and building alerts.
  2. Logs – immutable, timestamped records of events. Useful for root‑cause analysis because they contain context.
  3. Traces – end‑to‑end request paths across services, usually sampled. They reveal latency contributors and call graphs.

A fourth, often‑mentioned pillar is events, which capture state changes that don’t fit neatly into the other three.

Typical follow‑up: How do you decide which pillar to prioritize for a new service?


3. Push vs. Pull data collection – pros and cons?

  • Push (agents send data): Works well behind NAT, gives you control over batching, but can overload the network if many agents fire at once.
  • Pull (scrape endpoints): Simpler to secure, lets the collector dictate scrape intervals, but requires the target to be reachable and can miss bursts.

In practice, many teams use a hybrid model: push for logs and traces, pull for metrics.

Typical follow‑up: What did you use in your last project and why?


4. Explain sampling in distributed tracing.

Tracing every request can be costly in terms of storage and CPU. Sampling reduces overhead by recording only a fraction of requests. Common strategies:

  • Fixed-rate – e.g., 1% of all requests.
  • Adaptive – increase sample rate when latency spikes.
  • Head-based – decide at the entry point, useful for consistent downstream correlation.

The key is to keep enough data to answer the questions you care about while staying within budget.

Typical follow‑up: How did you verify that your sampling rate still gave you enough signal?


5. How do you set up an effective alerting strategy?

  1. Start with SLOs – define what good performance looks like (e.g., 99.9% of requests under 200 ms).
  2. Derive alerts from SLO breach – fire when error budget consumption exceeds a threshold.
  3. Add noise reduction – use multi‑dimensional alerts (e.g., error rate and latency) and debounce logic.
  4. Include runbooks – link to remediation steps directly in the alert.

This approach keeps alerts actionable and reduces fatigue.

Typical follow‑up: Can you walk me through a recent alert you handled and what you learned?


6. What is “high cardinality” and why does it matter for observability?

High cardinality refers to attributes that have many distinct values, such as user IDs or request IDs. Storing high‑cardinality tags in a time‑series database can explode storage and query cost. The usual mitigation is to limit tags to low‑cardinality dimensions (e.g., service name, status code) and push high‑cardinality data into logs or trace spans.

Typical follow‑up: How did you refactor a metric that was causing performance issues due to high cardinality?


7. How would you use observability to reduce MTTR (Mean Time to Recovery)?

  1. Detect quickly – fine‑tuned alerts based on SLOs.
  2. Correlate – combine metrics, logs, and traces to narrow the failure domain.
  3. Automate – trigger runbooks or remediation scripts when a known pattern appears.
  4. Post‑mortem – use the collected data to document the incident and improve the signal for next time.

In my last role, adding trace‑based alerts cut our MTTR from roughly 30 minutes to under 12 minutes.

Typical follow‑up: What tooling did you rely on to achieve that improvement?


FeaturePrometheus + GrafanaOpenTelemetry + Jaeger
Data typePrimarily metricsTraces (with logs via extensions)
Collection modelPull‑based scrapingInstrumentation libraries push spans
StorageLocal TSDB, optional remote writeDistributed trace storage (e.g., Elasticsearch)
EcosystemRich alerting, wide communityVendor‑agnostic, easy to add new signals
Typical use caseService‑level metrics & alertsEnd‑to‑end request debugging

Both can coexist; a common pattern is Prometheus for metrics and OpenTelemetry for traces, feeding a unified dashboard.

Typical follow‑up: Which combination would you recommend for a microservice architecture and why?


9. What are “derived metrics” and when should you create them?

Derived metrics are calculated from raw data, such as error‑rate = errors / total requests. They are useful when the raw series are noisy or when you need a business‑level view (e.g., revenue per request). Create them when the calculation is stable, adds clarity, and can be expressed in the monitoring system without excessive compute.

Typical follow‑up: Give an example of a derived metric you introduced and its impact.


10. How do you ensure observability data respects privacy and compliance?

  • Redact PII – strip or hash sensitive fields before ingestion.
  • Retention policies – keep detailed logs for a short window, aggregate metrics longer.
  • Access controls – limit who can query raw logs vs. aggregated dashboards.
  • Audit trails – log who accessed what data and when.

Most regulated industries enforce these practices, and many observability platforms provide built‑in support.

Typical follow‑up: Describe a compliance audit you participated in and the changes you made.


Sample Answer Template (Senior Level)

“In my last role we were dealing with intermittent latency spikes that our dashboards didn’t surface because we were only looking at average latency. I added a 99th‑percentile latency metric and correlated it with trace sampling during spikes. The combined view let us pinpoint a downstream cache miss that was causing the tail latency. After we fixed the cache invalidation logic, the 99th‑percentile dropped from 1.2 seconds to under 300 ms, and our error‑budget breach alerts stopped firing.”


How to practice this

  1. Record yourself – Use a phone or a tool like Call Assistant to capture a 60‑second answer, then replay it to check clarity and timing.
  2. Simulate follow‑ups – Have a colleague ask the typical follow‑up listed after each question; answer on the spot to build agility.
  3. Map to your resume – For each answer, identify a concrete project from your work history that demonstrates the point; keep a one‑sentence bullet ready for reference.

FAQ

  • What’s the difference between a metric and a log? Metrics are aggregated numeric values over time, ideal for alerting and trend analysis. Logs are raw, timestamped events that retain full context, useful for deep debugging.

  • Should I use OpenTelemetry or vendor‑specific SDKs? OpenTelemetry offers a vendor‑neutral API, making it easier to switch back‑ends later. If you’re locked into a single vendor for compliance or cost reasons, their SDK may provide tighter integration.

  • How much data should I sample for traces? Start with a low fixed rate (e.g., 1‑2%) and increase it when you detect performance anomalies. The exact percentage depends on your traffic volume and storage budget.

  • Can observability replace traditional testing? No. Observability helps you understand what happened in production, while testing aims to prevent issues before they reach users. Both are complementary.

Frequently asked questions

What’s the difference between a metric and a log?

Metrics are aggregated numeric time‑series used for alerting and trend analysis. Logs are raw, timestamped records that keep full context, making them ideal for root‑cause investigation.

Should I use OpenTelemetry or vendor‑specific SDKs?

OpenTelemetry provides a vendor‑neutral API, which eases future migrations. Vendor‑specific SDKs may offer tighter integration if you’re committed to a single platform for compliance or cost reasons.

How much data should I sample for traces?

Begin with a low fixed rate (around 1‑2%) and raise it when latency spikes appear. The optimal rate balances signal fidelity with storage and CPU overhead.

Can observability replace traditional testing?

No. Observability tells you what happened in production, while testing aims to prevent problems before they occur. They work together to improve reliability.

#observability#interview-questions#metrics#traces#monitoring#concept questions