Advanced System design concept · Distributed Systems & Data · 45 mins read

Service Communication Patterns

How services talk to each other — request/response or events — and how to deliver and publish messages reliably.

Synchronous vs Asynchronous Communication

Request/response versus messaging between services: coupling, failure propagation and how to choose.

Intuition

Checkout calls payment, which calls fraud, which calls a risk-scoring service. If risk scoring slows down, every service up the chain waits, threads fill up and checkout fails — even though only one small service had a problem.

How services communicate determines how failures spread. Most availability incidents in microservice systems involve long synchronous chains.

Mental Model

A synchronous call requires both services to be up at the same time and the caller to wait; availability multiplies along the chain (three services at 99.9% give about 99.7%). An asynchronous message is written to a broker and processed later; the sender only needs the broker to be up, and the receiver can be down, slow or scaled independently. Keep calls for questions that need an answer now; use messages for commands and notifications that can complete later. Think of it like: A phone call needs both people available at once. A text message can be read when the other person is free.

Building Blocks

  • Request/response: REST or gRPC calls where the caller waits for the result.
  • Commands: Asynchronous messages asking one specific service to do something, such as 'send this email'.
  • Events: Facts that something happened, such as 'OrderPlaced', which any number of services can react to.
  • Timeouts and circuit breakers: Limits that stop a slow dependency from tying up the caller.

Definitions

Temporal coupling
Two services must both be available at the same time for an operation to succeed.
  • Synchronous calls create it; queues remove it.
Cascading failure
A failure in one service spreading to its callers as they wait or retry.
  • Prevented with timeouts, circuit breakers and bulkheads.
Request–reply over messaging
An asynchronous request with a correlation ID, where the response arrives on a reply queue.
  • Useful for long-running work.

Patterns

  • Synchronous read, asynchronous side effects — Most user-facing write flows.
  • Accept and process later (202 Accepted) — Work that takes seconds or minutes.
  • Local copy of another service's data — When a service would otherwise call another on every request.

Strategies

  • Limit synchronous depth When: When designing call graphs. How: Keep user-facing paths to one or two synchronous hops; push everything else onto messages. Example: The product page reads from one aggregated read model instead of calling six services.
  • Protect every synchronous call When: Whenever a call remains synchronous. How: Set a timeout below the caller's own deadline, retry only idempotent calls with backoff, and add a circuit breaker and a fallback. Example: Recommendations time out after 150 ms and the page shows best-sellers instead.

Choosing between commands, events and calls

Ask what the sender needs. If it needs data or a decision before it can continue — 'is this card valid?' — use a synchronous call. If it needs a specific service to do something but not right now — 'generate this invoice' — send a command to that service's queue. If it is simply announcing a fact, without caring who reacts — 'order placed' — publish an event.

Events give the loosest coupling: the publisher does not know its consumers, so new consumers can be added without changing it. The price is that the overall flow is no longer written down in one place, and the data each service sees is eventually consistent. Good designs mix all three, choosing per interaction rather than declaring the whole system 'event-driven'.

Tradeoffs

DecisionUpsideDownside
SynchronousSimple, immediate result, easy to trace and debug.Temporal coupling; latency and failures add up along the chain.
AsynchronousServices fail and scale independently; the broker absorbs spikes.Eventual consistency, duplicate handling, harder tracing and more infrastructure.

Real World

SystemHow it's used
AmazonOrder placement confirms quickly while fulfilment, notifications and recommendations run as asynchronous downstream processes.
NetflixWraps synchronous dependencies with timeouts, fallbacks and circuit breakers so one failing service does not break the page.

Interview

Questions interviewers ask

  • Which parts of checkout should be synchronous?
  • What is the difference between a command and an event?
  • How do you stop one slow service from taking down others?

What a strong answer covers

Classify each interaction as call, command or event, explain temporal coupling and availability math, and protect remaining calls.

Common traps

  • Making everything synchronous because it is simpler.
  • Making the user wait on work that could be done later.

Quiz

Three services in a synchronous chain each have 99.9% availability. Roughly what is the chain's availability?
  1. 99.9%
  2. 99.7%
  3. 99.99%
  4. 97%

0.999³ ≈ 0.997, because all three must be up for a request to succeed.

What is temporal coupling?
  1. Shared clocks
  2. Both services must be available at the same time
  3. Services deployed together
  4. Using timestamps

A synchronous call fails if the receiver is down at that moment.

Sending a welcome email after sign-up is best done…
  1. Synchronously before responding
  2. Asynchronously via a message
  3. In the database trigger
  4. By the client

The user does not need to wait for it, and an email outage should not block sign-up.

How does an event differ from a command?
  1. Events are faster
  2. An event states a fact for anyone; a command asks a specific service to act
  3. Commands are always synchronous
  4. No difference

Events are broadcast facts; commands are directed requests.

Which protects a caller from a slow synchronous dependency?
  1. More retries without limit
  2. A timeout plus a circuit breaker and fallback
  3. Bigger thread pools only
  4. Synchronous replication

Timeouts bound waiting, and the breaker stops calling a failing service.

Event Streaming

Durable, partitioned, replayable logs such as Kafka: how they differ from queues, and how ordering and consumer groups work.

This section is part of the full PRISM roadmap, with worked examples, trade-off tables, interview questions and a quiz.

Unlock the full lesson

Delivery Semantics

At-most-once, at-least-once and exactly-once processing: what each guarantees and how idempotent consumers make duplicates harmless.

This section is part of the full PRISM roadmap, with worked examples, trade-off tables, interview questions and a quiz.

Unlock the full lesson

Transactional Outbox & CDC

Publish events reliably with database changes using an outbox table, and stream changes out with change data capture.

This section is part of the full PRISM roadmap, with worked examples, trade-off tables, interview questions and a quiz.

Unlock the full lesson

Practice service communication patterns in PRISM

Concepts stick when you watch them fail. Build an architecture that depends on service communication patterns, push traffic through it in the PRISM simulator, and see the latency and error rates change as you adjust the design.