Design Real-time Analytics System

Near real-time analytics that ingests billions of events, aggregates them into time-windowed metrics, and powers live dashboards with sub-minute freshness.

Functional requirements

  • Ingest high-throughput event streams via authenticated API.
  • Durably buffer events and support at-least-once delivery to processors.
  • Aggregate metrics in tumbling/sliding windows with late-event handling.
  • Power live dashboards and APIs with sub-minute freshness.
  • Support ad-hoc drill-down on raw events for debugging.
  • Backfill and reprocess history when aggregation logic changes.
  • Detect anomalies and trigger alerts on threshold/ML rules.
  • Multitenant isolation with per-tenant quotas and cost controls.

Non-functional requirements

  • Availability: 99.95%+; single AZ failure should not halt ingest or queries.
  • Latency: ingest API p99 < 150 ms; dashboard API p95 < 300 ms; end-to-end freshness < 60 s.
  • Durability: raw events persisted with replication and backups; immutability for audit.
  • Scalability: independently scale ingest, queues, processors, and query tiers across multi-AZ deployments.
  • Consistency: at-least-once across stream; idempotent processors; read-after-write for configs.
  • Security: TLS 1.2+ in transit, encryption at rest, per-tenant API keys/OAuth, and RBAC for dashboards.
  • Resilience: timeouts, circuit breakers, retries/backoff, DLQs, and backpressure.
  • Observability: traces/logs/metrics for queue depth, lag, watermark, window lateness, and query p99.

How the design evolves

Stage 1: Monolith MVP

Single service handles ingest, simple rollups, and queries backed by one database.

What was missing: Edge protection, buffering, scalable processors, and OLAP tier.

Why that's risky: SPOF and DB contention; limited throughput and freshness.

What gets added: Nothing yet (MVP).

Trade-offs: Fast to build; not production-ready.

Stage 2: Edge, Throttling, and Durable Buffer

Protect ingress and decouple ingest from processing with a queue.

What was missing: Ingress protection and decoupling from processors.

Why that's risky: Overload propagates to compute and storage.

What gets added: Edge + rate limits + durable queue with optional raw landing.

Trade-offs: Additional infra and ops complexity.

Stage 3: Stream Processing and Query Tier

Add stream processor, rollup store, and query API with cache.

What was missing: Scalable processing and fast query path.

Why that's risky: Synchronous compute and DB contention.

What gets added: Stream processor, rollup store, query API, and cache with a dedicated read path from edge.

Trade-offs: Eventual consistency of aggregates.

Stage 4: DLQ, Late Events, and Backfills

Harden pipelines with DLQ, late event handling, and backfill workers.

What was missing: Failure isolation and historical correction.

Why that's risky: Poison events and late data skew aggregates.

What gets added: DLQ isolation, replayer workers, and controlled backfill workers for late/corrective processing.

Trade-offs: Higher operational complexity (DLQ workflows, replay safety, and backfill controls).

Stage 5: Multi-Region Read Scale and OLAP

Add OLAP store/read replicas and region-friendly query performance.

What was missing: Read scale and flexible ad-hoc querying.

Why that's risky: Query hotspots and slow ad-hoc scans.

What gets added: Global CDN, rollup read replica, and OLAP tier for ad-hoc analytics without impacting writes.

Trade-offs: Replica lag and ETL freshness trade-offs must be managed with SLOs.

Stage 6: Observability and Anomaly Alerts

Add monitoring, lag SLOs, and anomaly detection with alerting.

What was missing: Operational visibility and automated incident response.

Why that's risky: Lag/latency regressions unnoticed; slow manual response.

What gets added: End-to-end monitoring/tracing, anomaly detection, and automated alert notifications.

Trade-offs: Additional cost and potential noisy alerts.

Frequently asked questions

Exactly-once vs at-least-once: what should I pick?

Prefer at-least-once with idempotent processors and de-duplication keys. Exactly-once often adds complexity without clear benefits; track offsets and use transactional writes only where strictly required.

How do you handle late and out-of-order events?

Use event-time processing with watermarks and allowed lateness. Emit corrections to the same aggregate keys; maintain compactable state and a backfill worker to reconcile windows that exceed lateness thresholds.

How do you keep dashboards fast at scale?

Front the query API with a cache for hot dashboards, denormalize aggregates for the most common queries, and use replicas for read scaling. Limit query shapes, paginate, and pre-warm caches around known spikes.

What do you alert on for stream health?

Alert on consumer lag, watermark stall, processing error rate, queue depth, and end-to-end freshness. Create SLOs and burn-rate alerts to react before users notice.

How do you run backfills safely in production?

Throttle backfills, write to side tables or temp partitions, and validate results before swapping. Use idempotent jobs and pause conflicting online tasks. Monitor impact on primary queues and stores.

PRISM
System Design Interview
Round 1 of 4 · Architecture Design · 60:00 remaining
PRISM logo
AI Interview
Interview Prep
Interview Challenges
Design Your Own System NEW
System Architectures
Interactive Roadmap System Design Guides
Notifications
  • No new notifications
Feedback
Signed in
Phase
Design Your Own System
Phase 01: Thinking in Systems
Upcoming

Components

User
CDN
Load Balancer
Server
Cache
Database
Blob Storage
Search Index
Queue
Worker
Rate Limiter
Service
API Gateway
Reverse Proxy
WebSocket Server
Third-party API

Inspector

Notes
Use clear names so your design intent is easy to understand.
Good
Name by business meaning
"Order API", "Restaurant Service"
Avoid
Generic names = zero signal
"Server 1", "API", "Queue"
A short description for each component makes feedback much better.
100%
Start by identifying:
  • Users & entry points
  • APIs & services
  • Databases & storage
  • Traffic flow & scale

Round 1 of 4 Architecture Design

Run a simulation to see results.

Time Remaining
60:00
System Design Interview

What are the core functional requirements?
What are the key non-functional constraints?

Questions

Start Evaluation to unlock questions.

Components Added
No components listed
Click "+ Add" to document components introduced in this stage.
Design Decisions

Click Simulate to run your design and see results here.

Internal notes — not shown to learners.

EVALUATE MODE

Test yourself like it's the real thing.

A structured 4-module evaluation that mirrors how top companies assess system design candidates.

Architecture Design
Draw your system on the canvas. Define components, connections, and data flow.
FR & NFR Requirements
Answer functional and non-functional requirement questions about your design.
MCQ Round
Multiple choice questions testing your depth on the chosen system.
Tradeoff Analysis
Justify your design decisions and defend your architectural tradeoffs.
AI Report Generated
A R S
Used by engineers preparing for FAANG & top-tier companies
Choose a Problem
No problem selected
  • 30 min
  • 45 min
  • 60 min
Round 2 of 4
MCQ Round
Answer multiple-choice questions based on your design.

Exit Interview?

You're in the middle of an interview session. Leaving now will end your current attempt.

Your progress will be saved.

Open a saved design

Select a design to load into the canvas.

My Evaluations

Your past evaluation sessions

Here’s a simple request flow that follows the expected layer order.

External User
→
Edge CDN → API Gateway → Load Balancer
→
Compute App Servers / Services
→
DataAccess Cache
→
Storage Database / Search Index
→
Async Queue → Worker

Tip: keep arrows moving forward through layers (Edge → Compute → Storage). Avoid sending storage back to compute.

Evaluation Instructions

Read the rules carefully before starting. The test auto-submits on refresh.

Before you start

  • Build your architecture on the canvas. The timer starts when you click Start Evaluation.
  • Don't forget to answer Functional Requirement and Non Functional Requirements.
  • When satisfied with your design, click Next to lock it and view the questions.
  • Please answer final step questions to complete the evaluation.

Dos

  • Do read each question carefully before answering.
  • Do include required components to maximize component coverage.
  • Do save a copy of your design if you want to keep it before submission.

Don'ts

  • Don't refresh or close the tab during an active evaluation — this will auto-submit your answers.
  • Don't switch app modes or open another tab while the evaluation is running.
  • Don't attempt to edit the design after clicking Next; the workspace will be locked.

All the best!!

Confirm

Input

Notice

Evaluation Report:

Evaluation Complete

Generating Your Report

Hang tight — our AI is evaluating your design…

Did you know?

Loading…

Share feedback

Tell us what worked well and what we can improve.

Let's personalize this

Answer a couple of quick questions so we can tailor your journey and missions.

You can change this anytime from your Profile.

Your personalized missions are ready

We tailored these first steps based on your answers.

    PRISM Welcome Gift

    This is a personal welcome gift from PRISM.

    Congratulations.

    You explored PRISM.

    You earned Apprentice.

    As a welcome gift, unlock Full PRISM Access for the configured trial duration.

    This starts Trial. Trial timer begins only after you activate this gift.

    Welcome to PRISM

    We've prepared a personalized Apprentice Journey based on your goals and experience.

    This journey introduces you to the capabilities of PRISM that are most relevant to you.

    Complete all 8 missions to earn your Apprentice title. 8 MISSIONS

    PRISM Surprise Offer

    Complete your Apprentice Journey to unlock a special gift from PRISM.

    • No payment required
    • No credit card required
    • Just complete the journey
    MISSION CONTROL
    0 / 8 missions complete
    NEXT UP Continue your missions
    View full roadmap →
    Mission Complete 0 / 8 Completed Next: Keep going
    SYSTEM BRIEF

    ⬤ System Constraints

    What the system must do — every item is a user-facing behaviour your architecture must support.

      ↑ Engineering Constraints

      These are the failure modes you must design against — latency SLAs, durability targets, traffic ceilings.

        ⇆ Architecture Constraints

        ◈ Core Concepts to Master

        Your Journey
        PHASE – –
        0 / 0 0%
        0
        Mock Interview Checklist

        Pick a topic to start

        Explore concept overviews, real-system examples, key tradeoffs, and interview talking points for each roadmap section.

        Topic-Wise Progress
        Experience Points 0 XP
        Read subtopics & solve challenges to earn XP
        Theory Read +0 XP
        Solved +0 XP
        Streak Bonus +0 XP
        Theory Read 0%
        — Mastered — Solved
        Weekly Streak 0 day streak
        Mon
        Tue
        Wed
        Thu
        Fri
        Sat
        Sun
        Keep going — log in daily to build your streak!
        0 0%
        Skill Profile
        Recommended Next
        🎯 Your Focus

        You haven't explored enough yet.

        Focus on
        → Understanding System Design
        → Estimating Scale
        Next Action
        Continue → Next: –
        Mock Interview Checklist
        Architecture DNA
        Engineering Profile
        Phase Mastered

        You've conquered this phase. These are the skills you now own:

          +500 XP

          Engineering Profile

          Company Interview Paths

          Progress Summary