Microservices Behind an API Gateway: A Real Case Study
Five services, one gateway, and the honest reason we split a Django app that was working fine. Service boundaries that survived, the two I drew wrong and pulled back, what belongs in a gateway and what must never, and the failure modes that only exist once a call can cross a network.
The first version of the library system was one Django app and it worked fine. It is still the thing I would build again for the same requirements. What changed was not the load — it was that four different teams needed to ship on four different schedules, and every deploy meant coordinating all of them.
That is the honest reason we split it, and it is worth saying plainly because "we needed to scale" is the reason people give and it is almost never true. This post is the architecture we ended up with: five services behind one gateway, what each boundary cost, the two boundaries I drew wrong, and the failure modes that only appear once a call can cross a network.
The system
Five services and why each one exists
A public library: catalogue, borrowing, members, notifications, search. Roughly 40,000 members across a dozen branches — not a scale story.
| Service | Owns | Why it is separate |
|---|---|---|
| Catalogue | Titles, physical copies, branch stock | Changes on a cataloguing schedule — weekly bulk imports, nothing else touches it |
| Lending | Loans, holds, renewals, fines | The transactional core, and the only service with genuinely hard consistency rules |
| Members | Identity, cards, contact details | Different data-protection rules and a different audit requirement from everything else |
| Search | A read-only projection of catalogue + availability | Completely different scaling and latency profile; rebuilt from events, owns no truth |
| Notifications | Email and SMS delivery | Talks to flaky third parties and must never be able to fail a checkout |
The mistakes
Two boundaries I drew wrong
Both had the same shape and it took an incident each to see it.
Fines as its own service. It felt clean — money is its own domain.
But a return can settle a fine, and a fine can block a loan. Every checkout became a two-service transaction with no transaction. We wrote compensating logic, then compensating logic for the compensating logic.
Fines inside Lending. One database, one transaction, one place where "can this member borrow" is answered.
Fines is now a module with its own folder and its own tests. It has all the isolation it ever needed and none of the network.
The second was Holds, split out for exactly the same reason and pulled back for exactly the same one. The rule that came out of both:
The gateway
What belongs in it, and what must never
The gateway is the most tempting place to put things and the worst place to put most of them. Everything routes through it, so anything you add there becomes a dependency of the whole system.
TLS termination · authentication (verify the token, attach the claims) · rate limiting · routing · request IDs · CORS · the timeout budget
Every one of these is true for all traffic and has nothing to do with any domain.
Authorisation ("may this member borrow") · response aggregation across services · any database · any business rule · any retry that changes state
Each of these makes the gateway know something a service should own, and a shared deploy for a domain change.
Contracts
The interface is the product
Inside a monolith a bad function signature is a refactor. Across a network it is a coordinated release. Three rules kept that manageable:
-
Additive changes only, until a major version
New optional fields are free. Removing a field, renaming one, or narrowing a type is a breaking change and needs a version. Consumers must ignore fields they do not recognise — write that down, because a strict parser on the consumer side turns every additive change into a breaking one.
-
The contract lives in the repo that serves it
An OpenAPI spec next to the code, generated in CI, published as an artifact. A spec maintained separately from the implementation is documentation, and documentation drifts.
-
Consumer-driven contract tests in CI
Lending's test suite asserts the shape it needs from Members. That test runs in Members' pipeline. This is the single highest-value thing on the list — it is the difference between finding a break at merge and finding it in production.
Failure
The modes that only exist once there is a network
This is the part the diagrams never show and the part that decides whether the split was worth it.
| Failure | What it looks like | What we do |
|---|---|---|
| Cascading timeout | Members is slow; Lending waits; the gateway waits; every request in the system is now slow | A timeout budget set at the gateway and divided going down, never a per-call default |
| Retry storm | A struggling service gets three times its normal traffic from well-meaning clients | Retries only on idempotent reads, jittered backoff, and a circuit breaker that gives up |
| Partial write | The loan is recorded; the notification never sends | Notifications is event-driven and at-least-once, so it is a delayed email, not a lost loan |
| Stale projection | Search says a copy is available; Lending says it is out | Lending is authoritative at checkout; the UI never promises what only Search knows |
| The whole bus is down | Events queue up; search goes stale; email stops | Nothing synchronous fails. Borrowing a book still works, which is the point of the shape |
The timeout budget is the load-bearing idea
Every call in a chain having its own 30-second default is how one slow service takes down a system. The budget is set once at the edge and spent going down:
# The gateway gives the whole request 3 seconds and says so in a header.
# Each hop spends from what is left and passes the remainder on.
def forward(request, service):
remaining = float(request.headers.get("X-Deadline-Ms", 3000))
if remaining <= 50:
raise DeadlineExceeded() # do not start work you cannot finish
started = time.monotonic()
response = http.post(
service.url,
timeout=remaining / 1000,
headers={"X-Deadline-Ms": str(remaining - 50)}, # reserve overhead
)
spent = (time.monotonic() - started) * 1000
metrics.observe("hop_ms", spent, service=service.name)
return response
The remaining <= 50 check matters more than it looks. Without it, a service at the end of a chain starts a two-second database query on behalf of a caller that gave up a second ago — burning capacity for a response nobody will read, at exactly the moment the system is already struggling.
Honesty
What this cost, and what should have stayed together
- Local development got much worse. Five services, a bus and a gateway is a compose file and a lot of RAM. New developers lose a day on it. This is a real, recurring tax that no diagram shows.
- Every cross-service change is two pull requests in an order that matters. Additive-first discipline makes it survivable rather than pleasant.
- Debugging needs distributed tracing from day one. Without a request ID threaded end to end, "it was slow" is unanswerable. We added it late and regretted the gap.
- Notifications did not need to be a service. It needed to be a queue. It is separate now largely because it was separate then, which is not a reason.
- If the team were still four people, one Django app would win. The split bought independent deploys for independent teams. That is the only thing it bought, and it is worth having only when you have the second thing.
The checklist
- Split for release cadence and team ownership, not for load you do not have.
- Never split a transaction — things consistent within one user action live together.
- The gateway does authentication, not authorisation, and holds no database.
- Contracts live with the code that serves them, generated in CI.
- Consumer-driven contract tests run in the provider's pipeline.
- A timeout budget from the edge, divided down the chain, checked before starting work.
- Retry only idempotent reads, with jitter and a circuit breaker.
- Anything that can be async, is — search projections and notifications never block a user.
- A request ID threaded through everything from the first day.
- Write down what each service owns, in one sentence. If you cannot, the boundary is wrong.
Where to start on Monday
Do not start by drawing services. Start by writing the one sentence each candidate service owns, then look for a user action that crosses two of them and has to be consistent. Every one of those is a boundary you are about to draw in the wrong place.
- service boundaries
- API gateway
- timeout budget
- circuit breaker
- contract tests
- event bus
- read projections
- request IDs
The real-time layer sits awkwardly on top of all this, which is its own problem and its own post: adding WebSockets to a microservice architecture. If you are running this on your own hardware rather than a managed platform, the CI/CD guide covers getting five services deployed without a cloud platform underneath you.
Comments (0)
No comments yet
Be the first to share a thought on this article.
Join the conversation