Skip to content
A AhsanLab.Tech
Architecture 16 min read · April 28, 2026

Microservices Behind an API Gateway: A Real Case Study

Five services, one gateway, and the honest reason we split a Django app that was working fine. Service boundaries that survived, the two I drew wrong and pulled back, what belongs in a gateway and what must never, and the failure modes that only exist once a call can cross a network.

A Ahsan Habib Save

The first version of the library system was one Django app and it worked fine. It is still the thing I would build again for the same requirements. What changed was not the load — it was that four different teams needed to ship on four different schedules, and every deploy meant coordinating all of them.

That is the honest reason we split it, and it is worth saying plainly because "we needed to scale" is the reason people give and it is almost never true. This post is the architecture we ended up with: five services behind one gateway, what each boundary cost, the two boundaries I drew wrong, and the failure modes that only appear once a call can cross a network.

5servicesfrom one Django app
2boundaries redrawnboth split a transaction
1gatewayand it stays thin

The system

Five services and why each one exists

A public library: catalogue, borrowing, members, notifications, search. Roughly 40,000 members across a dozen branches — not a scale story.

WEB + MOBILE one base URL BRANCH KIOSK self-checkout API GATEWAY authn · routing rate limit · TLS CATALOGUE titles, copies, branches LENDING loans, holds, fines MEMBERS identity, cards SEARCH read-only projection NOTIFICATIONS email, SMS · async only EVENT BUS loan.created, etc. search and notifications never sit on the critical path
FIGURE 1 — THE SHAPE · Solid arrows are synchronous calls a user waits on. Dashed arrows are events nobody waits on. The ratio between those two is the architecture.
ServiceOwnsWhy it is separate
CatalogueTitles, physical copies, branch stockChanges on a cataloguing schedule — weekly bulk imports, nothing else touches it
LendingLoans, holds, renewals, finesThe transactional core, and the only service with genuinely hard consistency rules
MembersIdentity, cards, contact detailsDifferent data-protection rules and a different audit requirement from everything else
SearchA read-only projection of catalogue + availabilityCompletely different scaling and latency profile; rebuilt from events, owns no truth
NotificationsEmail and SMS deliveryTalks to flaky third parties and must never be able to fail a checkout
The test that actually decides a boundary Not "is this a different noun" — every table is a different noun. The question is whether the two halves change for different reasons, on different schedules, under different failure tolerances. Catalogue and Search are the same nouns and belong apart. Loans and Fines are different nouns and belong together.

The mistakes

Two boundaries I drew wrong

Both had the same shape and it took an incident each to see it.

Wrong

Fines as its own service. It felt clean — money is its own domain.

But a return can settle a fine, and a fine can block a loan. Every checkout became a two-service transaction with no transaction. We wrote compensating logic, then compensating logic for the compensating logic.

Right

Fines inside Lending. One database, one transaction, one place where "can this member borrow" is answered.

Fines is now a module with its own folder and its own tests. It has all the isolation it ever needed and none of the network.

The second was Holds, split out for exactly the same reason and pulled back for exactly the same one. The rule that came out of both:

If two things must be consistent within one user action, they belong in one service. A distributed transaction is not an architecture, it is a bill you pay every day for a diagram that looked tidy once.

The gateway

What belongs in it, and what must never

The gateway is the most tempting place to put things and the worst place to put most of them. Everything routes through it, so anything you add there becomes a dependency of the whole system.

Belongs

TLS termination · authentication (verify the token, attach the claims) · rate limiting · routing · request IDs · CORS · the timeout budget

Every one of these is true for all traffic and has nothing to do with any domain.

Does not belong

Authorisation ("may this member borrow") · response aggregation across services · any database · any business rule · any retry that changes state

Each of these makes the gateway know something a service should own, and a shared deploy for a domain change.

Authentication yes, authorisation no The gateway proves who you are and passes the claims down. Whether you may do the thing is a question only the service holding the data can answer, because the answer depends on state the gateway does not have. Put it in the gateway and you either duplicate rules or make the gateway call services to decide whether to call services.

Contracts

The interface is the product

Inside a monolith a bad function signature is a refactor. Across a network it is a coordinated release. Three rules kept that manageable:

  1. Additive changes only, until a major version

    New optional fields are free. Removing a field, renaming one, or narrowing a type is a breaking change and needs a version. Consumers must ignore fields they do not recognise — write that down, because a strict parser on the consumer side turns every additive change into a breaking one.

  2. The contract lives in the repo that serves it

    An OpenAPI spec next to the code, generated in CI, published as an artifact. A spec maintained separately from the implementation is documentation, and documentation drifts.

  3. Consumer-driven contract tests in CI

    Lending's test suite asserts the shape it needs from Members. That test runs in Members' pipeline. This is the single highest-value thing on the list — it is the difference between finding a break at merge and finding it in production.

Failure

The modes that only exist once there is a network

This is the part the diagrams never show and the part that decides whether the split was worth it.

FailureWhat it looks likeWhat we do
Cascading timeout Members is slow; Lending waits; the gateway waits; every request in the system is now slow A timeout budget set at the gateway and divided going down, never a per-call default
Retry storm A struggling service gets three times its normal traffic from well-meaning clients Retries only on idempotent reads, jittered backoff, and a circuit breaker that gives up
Partial write The loan is recorded; the notification never sends Notifications is event-driven and at-least-once, so it is a delayed email, not a lost loan
Stale projection Search says a copy is available; Lending says it is out Lending is authoritative at checkout; the UI never promises what only Search knows
The whole bus is down Events queue up; search goes stale; email stops Nothing synchronous fails. Borrowing a book still works, which is the point of the shape

The timeout budget is the load-bearing idea

Every call in a chain having its own 30-second default is how one slow service takes down a system. The budget is set once at the edge and spent going down:

gateway middleware — the budget travels with the request
# The gateway gives the whole request 3 seconds and says so in a header.
# Each hop spends from what is left and passes the remainder on.

def forward(request, service):
    remaining = float(request.headers.get("X-Deadline-Ms", 3000))
    if remaining <= 50:
        raise DeadlineExceeded()          # do not start work you cannot finish

    started = time.monotonic()
    response = http.post(
        service.url,
        timeout=remaining / 1000,
        headers={"X-Deadline-Ms": str(remaining - 50)},   # reserve overhead
    )
    spent = (time.monotonic() - started) * 1000
    metrics.observe("hop_ms", spent, service=service.name)
    return response

The remaining <= 50 check matters more than it looks. Without it, a service at the end of a chain starts a two-second database query on behalf of a caller that gave up a second ago — burning capacity for a response nobody will read, at exactly the moment the system is already struggling.

Honesty

What this cost, and what should have stayed together

  • Local development got much worse. Five services, a bus and a gateway is a compose file and a lot of RAM. New developers lose a day on it. This is a real, recurring tax that no diagram shows.
  • Every cross-service change is two pull requests in an order that matters. Additive-first discipline makes it survivable rather than pleasant.
  • Debugging needs distributed tracing from day one. Without a request ID threaded end to end, "it was slow" is unanswerable. We added it late and regretted the gap.
  • Notifications did not need to be a service. It needed to be a queue. It is separate now largely because it was separate then, which is not a reason.
  • If the team were still four people, one Django app would win. The split bought independent deploys for independent teams. That is the only thing it bought, and it is worth having only when you have the second thing.

The checklist

  • Split for release cadence and team ownership, not for load you do not have.
  • Never split a transaction — things consistent within one user action live together.
  • The gateway does authentication, not authorisation, and holds no database.
  • Contracts live with the code that serves them, generated in CI.
  • Consumer-driven contract tests run in the provider's pipeline.
  • A timeout budget from the edge, divided down the chain, checked before starting work.
  • Retry only idempotent reads, with jitter and a circuit breaker.
  • Anything that can be async, is — search projections and notifications never block a user.
  • A request ID threaded through everything from the first day.
  • Write down what each service owns, in one sentence. If you cannot, the boundary is wrong.

Where to start on Monday

Do not start by drawing services. Start by writing the one sentence each candidate service owns, then look for a user action that crosses two of them and has to be consistent. Every one of those is a boundary you are about to draw in the wrong place.

  • service boundaries
  • API gateway
  • timeout budget
  • circuit breaker
  • contract tests
  • event bus
  • read projections
  • request IDs

The real-time layer sits awkwardly on top of all this, which is its own problem and its own post: adding WebSockets to a microservice architecture. If you are running this on your own hardware rather than a managed platform, the CI/CD guide covers getting five services deployed without a cloud platform underneath you.

#Scalability #Performance #Microservices #API Design #System Design

Comments (0)

No comments yet

Be the first to share a thought on this article.

Join the conversation

Comments are moderated before they appear.

Keep reading

Related articles