Connecting OpenAI to a Real Backend: Lessons From SEO Automation
The naive version worked perfectly on fifty pages and failed four ways on twelve thousand — none of them about prompt quality. Wiring a model into a real backend is a batch-processing problem with an expensive, rate-limited, non-deterministic function in the middle, and it needs everything batch processing always needed.
The brief was ordinary: crawl a client's 12,000-page site, find the pages with weak or missing metadata, and draft better ones. The naive version — loop over pages, call the API, save the result — worked perfectly on a sample of fifty and then failed in four different ways on the real run, none of which were about prompt quality.
That gap is the whole subject. Connecting a model to a real backend is mostly not a modelling problem; it is a batch-processing problem with an expensive, rate-limited, non-deterministic function in the middle of it.
The pipeline
Crawl, structure, generate, review
Screaming Frog does the crawl and exports CSV. That is deliberately the boring part — it is a mature tool that already handles redirects, canonicals, JavaScript rendering and robots rules, and reimplementing any of that would be a month of work to arrive at something worse.
Run one
Four failures, none of them about the prompt
| What happened | Why | Fix |
|---|---|---|
| Died at page 812 and lost everything | One long script, results in memory, written at the end | One queued job per page, result persisted immediately |
| Rate limited into a 40-minute stall | Twelve concurrent workers, no backoff, retry-on-error | Four workers, honour Retry-After, exponential backoff with jitter |
| Roughly 6% of outputs unusable | Free-text response parsed with a regex | JSON schema, Pydantic validation, one repair retry |
| The re-run cost as much as the first run | No idempotency — every page regenerated | Unique key on (run_id, url_hash); completed pages skip |
Every one of those is an ordinary backend problem. The lesson I would give anyone wiring a model into a real system is that you are building a batch pipeline, and it needs everything a batch pipeline has always needed — checkpointing, idempotency, backpressure, retries with backoff — plus one thing it never did: the function in the middle can return something structurally invalid.
@task(max_retries=3)
def draft_metadata(run_id: str, url: str) -> None:
key = f"{run_id}:{sha1(url)}"
# Idempotency first: a retry, a resumed run and a double-enqueue
# all cost nothing. The unique index decides, not a SELECT.
if Draft.objects.filter(key=key).exists():
return
page = Page.objects.get(run_id=run_id, url=url)
try:
result = generate(page) # schema-validated, retries once
except RateLimited as e:
raise self.retry(countdown=e.retry_after or backoff_with_jitter())
Draft.objects.create(
key=key, run_id=run_id, url=url,
title=result.title, description=result.description,
rationale=result.rationale, # why, for the reviewer
status="pending", # never "published"
input_tokens=result.usage.input,
output_tokens=result.usage.output,
cost_usd=price(result.usage), # per row, not per run
)
Money
Cost control is a design constraint, not an afterthought
At 12,000 pages, decisions that are invisible at fifty become the entire budget. Three changes cut the per-run cost by roughly three quarters, and only one of them was picking a cheaper model.
-
Filter before you generate, in plain code
A page with a good, unique, correctly-lengthed title does not need a model to tell it so. Length checks, duplicate detection and template matching removed 82% of the work with a SQL query. The cheapest model call is the one you do not make.
-
Send the smallest input that answers the question
We were sending whole page bodies. The heading structure, the first paragraph and the existing metadata are enough to write a title, and they are a fraction of the tokens. Measure what the extra context actually buys before paying for it on every row.
-
Hash the input and skip what has not changed
Monthly re-runs on a site that changes slowly should be cheap. Hashing the exact input we would send means an unchanged page costs one hash comparison. The second run was 310 calls instead of 2,140.
-
Put a hard ceiling on the run
A budget in the run config that stops the queue when exceeded. It has fired once, on a crawl that picked up a faceted-search URL space and found 90,000 pages. Without it that would have been a genuinely expensive Tuesday.
Trust
The review gate, and why it is not a formality
Nothing generated reaches a live page automatically. Not because the drafts are bad — most are good — but because the failure mode is asymmetric. A weak title costs a little traffic. A confidently wrong claim about a product, published across 400 pages under a client's name, is a different category of problem entirely, and it is the one an automated pipeline produces at scale.
Three things made review fast enough that nobody tried to skip it:
- Bulk approve with diff view. Old and new side by side, keyboard shortcuts, 200 pages in twenty minutes. Review that is slow gets bypassed, and a bypassed gate is worse than no gate because it looks like one.
- The model states its reasoning in one line. "Original title was a duplicate of 43 other pages; new one uses the product's model number." The reviewer is checking a claim rather than judging prose.
- Confidence sorted to the bottom. The uncertain drafts are reviewed first, when attention is highest.
Honesty
What I would do differently
- Build the review UI first. We built generation, then discovered reviewing 2,140 drafts in a spreadsheet is unworkable. The gate is the product; the pipeline feeds it.
- There is still no evaluation set. "Good title" is judged by whoever reviews. That is the biggest gap here and it is why I would not claim a quality number for this system.
- Screaming Frog is a desktop tool in an automated pipeline, which means a person exports a CSV. It works and it is honestly a bit silly. The API-mode version is the obvious next step.
- Nobody measured whether it worked. We shipped better metadata and never ran a controlled before/after on ranking or click-through. The deliverable was drafts, not outcomes, and those are not the same thing.
The checklist
- One queued job per unit of work, never one long script.
- Persist each result immediately — nothing important lives in memory across 2,000 iterations.
- Idempotency key on every job, enforced by a unique index.
- Honour
Retry-After; back off with jitter; cap concurrency well below the limit. - Schema-validate every output and repair once before failing.
- Filter in SQL before you generate — most rows do not need a model.
- Send the smallest input that answers the question.
- Hash inputs so re-runs skip unchanged rows.
- A hard budget ceiling on the run, not just an alert.
- A human gate before anything publishes, made fast enough that nobody wants to skip it.
Where to start on Monday
Write the filter query before you write the prompt. Count how many rows genuinely need a model, and you will usually find the pipeline is a fifth of the size you were about to build — which makes every other decision cheaper.
- batch pipeline
- idempotency key
- Retry-After
- schema validation
- input hashing
- budget ceiling
- per-row cost
- human gate
The structured-output and retry pattern here is the same one in the local model benchmark, and the missing evaluation set is exactly what this post argues you should build first. I would take it as a fair criticism that this pipeline does not have one.
Comments (0)
No comments yet
Be the first to share a thought on this article.
Join the conversation