Skip to content
A AhsanLab.Tech
AI 13 min read · May 28, 2026

Connecting OpenAI to a Real Backend: Lessons From SEO Automation

The naive version worked perfectly on fifty pages and failed four ways on twelve thousand — none of them about prompt quality. Wiring a model into a real backend is a batch-processing problem with an expensive, rate-limited, non-deterministic function in the middle, and it needs everything batch processing always needed.

A Ahsan Habib Save

The brief was ordinary: crawl a client's 12,000-page site, find the pages with weak or missing metadata, and draft better ones. The naive version — loop over pages, call the API, save the result — worked perfectly on a sample of fifty and then failed in four different ways on the real run, none of which were about prompt quality.

That gap is the whole subject. Connecting a model to a real backend is mostly not a modelling problem; it is a batch-processing problem with an expensive, rate-limited, non-deterministic function in the middle of it.

12,000pages per runone client site
4failure modes on run onenone about prompting
100%drafts reviewed by a humanbefore anything ships

The pipeline

Crawl, structure, generate, review

Screaming Frog does the crawl and exports CSV. That is deliberately the boring part — it is a mature tool that already handles redirects, canonicals, JavaScript rendering and robots rules, and reimplementing any of that would be a month of work to arrive at something worse.

NOTHING HERE IS REQUEST/RESPONSE CRAWL Screaming Frog NORMALISE CSV → rows FILTER 12,000 → 2,140 QUEUE one job per page WORKERS × 4 model call + validate DRAFTS TABLE status: pending HUMAN REVIEW approve · edit · reject EXPORT / CMS approved only every job is idempotent on (run_id, url_hash) · a retry costs nothing · a re-run resumes nothing reaches a live page without a person pressing approve
FIGURE 1 — THE SHAPE · The model call is one box in the middle. Everything around it is the part that makes 2,140 of them survivable.

Run one

Four failures, none of them about the prompt

What happenedWhyFix
Died at page 812 and lost everything One long script, results in memory, written at the end One queued job per page, result persisted immediately
Rate limited into a 40-minute stall Twelve concurrent workers, no backoff, retry-on-error Four workers, honour Retry-After, exponential backoff with jitter
Roughly 6% of outputs unusable Free-text response parsed with a regex JSON schema, Pydantic validation, one repair retry
The re-run cost as much as the first run No idempotency — every page regenerated Unique key on (run_id, url_hash); completed pages skip

Every one of those is an ordinary backend problem. The lesson I would give anyone wiring a model into a real system is that you are building a batch pipeline, and it needs everything a batch pipeline has always needed — checkpointing, idempotency, backpressure, retries with backoff — plus one thing it never did: the function in the middle can return something structurally invalid.

tasks/draft_metadata.py — the job, with all four fixes in it
@task(max_retries=3)
def draft_metadata(run_id: str, url: str) -> None:
    key = f"{run_id}:{sha1(url)}"

    # Idempotency first: a retry, a resumed run and a double-enqueue
    # all cost nothing. The unique index decides, not a SELECT.
    if Draft.objects.filter(key=key).exists():
        return

    page = Page.objects.get(run_id=run_id, url=url)
    try:
        result = generate(page)                    # schema-validated, retries once
    except RateLimited as e:
        raise self.retry(countdown=e.retry_after or backoff_with_jitter())

    Draft.objects.create(
        key=key, run_id=run_id, url=url,
        title=result.title, description=result.description,
        rationale=result.rationale,                # why, for the reviewer
        status="pending",                          # never "published"
        input_tokens=result.usage.input,
        output_tokens=result.usage.output,
        cost_usd=price(result.usage),              # per row, not per run
    )
Store the cost per row, not per run A single total tells you the bill and nothing else. Per row you can answer the questions that actually come up: which page templates are expensive, whether the long-tail product pages are worth processing at all, and what a re-run will cost before you start it.

Money

Cost control is a design constraint, not an afterthought

At 12,000 pages, decisions that are invisible at fifty become the entire budget. Three changes cut the per-run cost by roughly three quarters, and only one of them was picking a cheaper model.

Naive all pages12,000 calls
+ filter first2,140 calls
+ trim the inputsame, 41% fewer tokens
+ skip unchanged310 calls on re-run
  1. Filter before you generate, in plain code

    A page with a good, unique, correctly-lengthed title does not need a model to tell it so. Length checks, duplicate detection and template matching removed 82% of the work with a SQL query. The cheapest model call is the one you do not make.

  2. Send the smallest input that answers the question

    We were sending whole page bodies. The heading structure, the first paragraph and the existing metadata are enough to write a title, and they are a fraction of the tokens. Measure what the extra context actually buys before paying for it on every row.

  3. Hash the input and skip what has not changed

    Monthly re-runs on a site that changes slowly should be cheap. Hashing the exact input we would send means an unchanged page costs one hash comparison. The second run was 310 calls instead of 2,140.

  4. Put a hard ceiling on the run

    A budget in the run config that stops the queue when exceeded. It has fired once, on a crawl that picked up a faceted-search URL space and found 90,000 pages. Without it that would have been a genuinely expensive Tuesday.

The most effective cost optimisation was a WHERE clause. Nothing about the model changed; we simply stopped asking it questions we could already answer.

Trust

The review gate, and why it is not a formality

Nothing generated reaches a live page automatically. Not because the drafts are bad — most are good — but because the failure mode is asymmetric. A weak title costs a little traffic. A confidently wrong claim about a product, published across 400 pages under a client's name, is a different category of problem entirely, and it is the one an automated pipeline produces at scale.

Three things made review fast enough that nobody tried to skip it:

  • Bulk approve with diff view. Old and new side by side, keyboard shortcuts, 200 pages in twenty minutes. Review that is slow gets bypassed, and a bypassed gate is worse than no gate because it looks like one.
  • The model states its reasoning in one line. "Original title was a duplicate of 43 other pages; new one uses the product's model number." The reviewer is checking a claim rather than judging prose.
  • Confidence sorted to the bottom. The uncertain drafts are reviewed first, when attention is highest.
What we do not do Auto-approve above a confidence threshold. It was proposed, sensibly, to save review time. A model's stated confidence is not a calibrated probability — it is a token sequence loosely correlated with being right, and it is most confident on exactly the fluent, plausible, wrong outputs you most need to catch.

Honesty

What I would do differently

  • Build the review UI first. We built generation, then discovered reviewing 2,140 drafts in a spreadsheet is unworkable. The gate is the product; the pipeline feeds it.
  • There is still no evaluation set. "Good title" is judged by whoever reviews. That is the biggest gap here and it is why I would not claim a quality number for this system.
  • Screaming Frog is a desktop tool in an automated pipeline, which means a person exports a CSV. It works and it is honestly a bit silly. The API-mode version is the obvious next step.
  • Nobody measured whether it worked. We shipped better metadata and never ran a controlled before/after on ranking or click-through. The deliverable was drafts, not outcomes, and those are not the same thing.

The checklist

  • One queued job per unit of work, never one long script.
  • Persist each result immediately — nothing important lives in memory across 2,000 iterations.
  • Idempotency key on every job, enforced by a unique index.
  • Honour Retry-After; back off with jitter; cap concurrency well below the limit.
  • Schema-validate every output and repair once before failing.
  • Filter in SQL before you generate — most rows do not need a model.
  • Send the smallest input that answers the question.
  • Hash inputs so re-runs skip unchanged rows.
  • A hard budget ceiling on the run, not just an alert.
  • A human gate before anything publishes, made fast enough that nobody wants to skip it.

Where to start on Monday

Write the filter query before you write the prompt. Count how many rows genuinely need a model, and you will usually find the pipeline is a fifth of the size you were about to build — which makes every other decision cheaper.

  • batch pipeline
  • idempotency key
  • Retry-After
  • schema validation
  • input hashing
  • budget ceiling
  • per-row cost
  • human gate

The structured-output and retry pattern here is the same one in the local model benchmark, and the missing evaluation set is exactly what this post argues you should build first. I would take it as a fair criticism that this pipeline does not have one.

#Performance #AI Tools #LLM #Prompt Engineering #AI Automation

Comments (0)

No comments yet

Be the first to share a thought on this article.

Join the conversation

Comments are moderated before they appear.

Keep reading

Related articles

AI 15 min

The 940ms Voice Agent: Where Every Millisecond Goes

A conversation has a social latency budget, not a technical one — go past a second and people think the thing is broken. How a voice agent went from 1,530ms to 940ms to first audio, why every millisecond came from overlapping stages rather than speeding them up, the speculative optimisation that cost 38% more tokens for 60ms, and what a user hears when a dependency vanishes mid-sentence.

Apr 12, 2026
AI 14 min

Your AI System Is Failing Right Now and Nothing Is Alerting

On a Tuesday my RAG system got quietly worse. Every request returned 200, no alert fired, and the answers stayed fluent and confident while being subtly about the wrong thing. Why AI systems fail silently, what a trace must actually record, why the mean is a lie, and the forty-minute root cause that would have taken days.

Mar 27, 2026