Skip to content
A AhsanLab.Tech
AI 16 min read · February 9, 2026

Inside the Agent Loop: How an AI Agent Actually Decides What to Do

My first agent made 41 tool calls to answer a question that needed two — and never threw a single error. The twelve lines an agent actually is, the five decisions inside the loop, why the tool schema is the real prompt, and the ten guardrails between a demo and something you would let near a customer.

A Ahsan Habib Save

The first agent I shipped made 41 tool calls trying to answer a question that needed two. It never crashed. It never errored. It just kept politely deciding that one more search would clear things up, and by the time I noticed, it had spent nineteen minutes and a genuinely upsetting amount of money being confidently unhelpful.

Nothing in the tutorial had prepared me for that, because the tutorial's agent had three tools and one test question. What I had built was a while loop with a language model inside it and no adult supervision.

This post is the thing I wish I had read that week: what an agent loop actually is, the five decisions inside it, and the guardrails that turn it from a plausible demo into something you can point at production traffic.

41tool calls, first runshould have been 2
8hard step ceilingafter the fix
94%runs ending in 3 stepsonce tools were named properly

The mechanism

An agent is a loop, and that is genuinely the whole idea

Strip away every framework and an agent is about twelve lines. Everything else in this post is a guardrail bolted onto one of them.

agent/loop.py — the entire concept
def run(goal: str, tools: dict, max_steps: int = 8) -> str:
    messages = [system_prompt(), {"role": "user", "content": goal}]

    for step in range(max_steps):
        reply = model.chat(messages, tools=schemas(tools))  # THINK
        messages.append(reply)

        if not reply.tool_calls:                    # STOP?
            return reply.content                    # ...answered

        for call in reply.tool_calls:               # ACT
            result = tools[call.name](**call.arguments)
            messages.append({                       # OBSERVE
                "role": "tool",
                "tool_call_id": call.id,
                "content": serialise(result),
            })

    return "Could not finish within the step budget."   # GIVE UP

That is it. The model does not "have agency" in any mysterious sense — it emits a structured request, your code runs a function, and you hand the result back so it can emit the next one. The loop is yours. The intelligence is rented; the control flow is entirely your responsibility.

GOAL the user's ask THINK one model call STOP? tool calls? ANSWER grounded, cited, done ACT your code runs OBSERVE result → context yes no — call a tool loop max 8 steps · $0.05 cap
FIGURE 1 — THE AGENT LOOP · The dashed paths are the ones that repeat. Everything that has ever gone wrong for me lives on one of them.
  1. Think — one model call, nothing more

    The model sees the goal, the tool schemas, and everything observed so far. It returns either a final answer or a request to call one or more tools. This step is stateless; all the memory is in what you passed it.

  2. Stop? — the most under-designed line in most agents

    "No tool calls" is the happy termination. There are three unhappy ones, and if you have not written them you will discover them in production.

  3. Act — your code, your blast radius

    The model does not execute anything. It names a function and supplies arguments. Whether that function can delete a row is a decision you already made when you registered it.

  4. Observe — the result goes back as context, and context costs money

    Every observation is appended and re-sent on the next turn. A tool that returns 8 KB of JSON is not a free call; it is a tax on every remaining step in the run.

  5. Give up — visibly, with what it has

    Hitting the step budget is a normal outcome, not a crash. Return the partial work and say so. An agent that silently returns nothing is indistinguishable from a hung request.

Where behaviour actually comes from

The tool schema is the prompt

This took me embarrassingly long to internalise. I spent three days rewriting the system prompt to stop the agent over-searching. The fix was one sentence — in a tool description.

The model chooses tools almost entirely from their names, descriptions and parameter docs. That schema is not plumbing you generate and forget; it is the highest-leverage prompt surface in the whole system.

Before

search(query: str)

"Search the documentation."

Average 6.2 calls per run. The model had no idea what it would get back, so it kept trying variations hoping for something better.

After

search_docs(query: str, section: str | None)

"Search Laravel/Filament docs. Returns the 6 most relevant passages with their source URLs. Covers framework docs only — not Stack Overflow, changelogs, or source code. If the first search returns nothing relevant, the answer is probably not in the docs."

Average 1.9 calls per run.

The failure this prevents A tool whose description does not say what it cannot do will be called for things it cannot do — repeatedly, because from the model's side an empty result is indistinguishable from a badly phrased query. Write the negative space.
agent/tools.py — the description carries the behaviour
@tool
def search_docs(query: str, section: str | None = None) -> list[Passage]:
    """Search the Laravel and Filament documentation.

    Returns up to 6 passages, each with its source URL and heading path.
    Covers official framework documentation ONLY — not Stack Overflow,
    not changelogs, not application source code.

    If the first call returns nothing relevant, do not rephrase and retry:
    the answer is very likely not in this corpus. Say so instead.

    Args:
        query:   A natural-language question. Full sentences beat keywords.
        section: Optional filter, one of "eloquent" | "filament" | "livewire".
    """

Three sentences of that docstring exist purely to stop a loop I watched happen. That is what tool design is: encoding the lessons of your traces into the one place the model reliably reads.

The resource nobody budgets for

State: the context window is a budget, not a memory

Every turn re-sends everything. The system prompt, the tool schemas, the full message history, and every observation collected so far. This is the part that surprises people coming from ordinary backend work, where state lives somewhere and you fetch what you need.

Here is the same run measured turn by turn. The bars are tokens, as a percentage of an 8k working budget:

Turn 11.1k
Turn 22.1k
Turn 33.3k
Turn 44.6k
Turn 56.3k
Turn 6 wall7.8k
Turn 7 compacted2.7k
Turn 83.7k

Growth is superlinear in practice, because the observations that accumulate are usually the largest messages in the run. Turn 6 is where the original agent started behaving strangely — not failing, just getting vaguer, because the early instructions were now a small fraction of a very long conversation.

WHAT FILLS THE CONTEXT WINDOW, TURN BY TURN working budget over compact → 1234 5678 TURN system prompt tool schemas message history observations
FIGURE 2 — CONTEXT GROWTH · The two fixed costs at the bottom never move. Everything above them is the run eating its own budget, and observations grow fastest.

The fix is compaction: once history crosses a threshold, summarise the middle and keep the ends.

agent/context.py — keep the ends, summarise the middle
def compact(messages: list[Message], limit: int = 6000) -> list[Message]:
    if count_tokens(messages) < limit:
        return messages

    head = messages[:2]     # system prompt + the original goal. Never drop these.
    tail = messages[-6:]    # the last three turns, verbatim — the model is mid-thought.
    middle = messages[2:-6]

    summary = model.chat([
        SUMMARISE_PROMPT,
        {"role": "user", "content": render(middle)},
    ]).content

    bridge = {"role": "assistant", "content": f"[earlier steps] {summary}"}
    return head + [bridge] + tail
The bug this replaced My first version was a sliding window that dropped the oldest messages. It dropped the original goal on turn 7, and the agent — reasonably, given what it could see — started answering a different question. Nothing errored. The answer was fluent and about the wrong thing.

Knowing when to quit

Termination: four exits, and only one of them is good

The tutorial loop has one exit. A real one needs four, and three of them should log loudly.

ExitTriggerWhat you returnHealth signal
Answered Model returns content with no tool calls The answer The only good one. Should be >90% of runs.
Step budget max_steps reached Partial work + "I could not finish" Rising rate means tools are too vague
Cost ceiling Cumulative spend over cap Partial work + escalate Should be near zero. Anything else is a runaway.
Repetition Same tool + same args twice Break, tell the model it looped The cheapest guard you will ever write

That last one caught my 41-call disaster in six lines:

agent/loop.py — the six lines that ended the 41-call run
seen: set[str] = set()

for call in reply.tool_calls:
    signature = f"{call.name}:{json.dumps(call.arguments, sort_keys=True)}"
    if signature in seen:
        # Do not just break — tell the model WHY, or the next turn repeats it.
        messages.append(nudge("You already ran that exact call. "
                              "Use what you have, or say you cannot answer."))
        break
    seen.add(signature)

Note that it does not silently break. A guard that stops the agent without telling the agent produces a confused final turn; a guard that explains itself produces a clean "I don't have enough to answer this."

The part that makes it shippable

Guardrails: five gates, any of which can stop a run

EVERY TOOL CALL PASSES FIVE GATES BEFORE ANYTHING HAPPENS MODEL asks for a tool ALLOWLIST is this tool even registered? ARGUMENTS Pydantic schema, bounds, injection BUDGET steps · tokens · dollars · wall clock CONSEQUENCE does it write, spend, or send? RUN it REFUSED — AND HANDED BACK TO THE MODEL AS AN OBSERVATION a rejection it can read is a rejection it can recover from; a silent drop is one it repeats
FIGURE 3 — THE GATES · The fourth one is the only place a human belongs, and only for calls that change something in the world.
A refusal the model can read is a refusal it can recover from. A silent drop is one it will make again on the very next turn.

The consequence gate deserves its own rule, because it is the one that stops an agent being a liability. I classify every tool at registration time:

ClassExamplesPolicy
Read search, fetch, look up an order Free to call, inside the budget
Write — reversible draft a reply, tag a record, create a note Allowed, logged, undoable
Write — irreversible send the email, issue the refund, delete Human confirms. Always. No confidence threshold.
Not a confidence problem I have seen "auto-approve above 0.9 confidence" proposed more than once. A model's stated confidence is not a calibrated probability — it is a token sequence that correlates loosely with being right. Gating a refund on it is gating a refund on a vibe.

What goes wrong

The five failure modes, and how each one announces itself

FailureWhat it looks likeGuard
The polite infinite loop Same search, slightly rephrased, forever Repetition detector + honest tool descriptions
Context drift Turn 8 answers a subtly different question Compaction that pins the goal
Tool hallucination Calls get_user_email, which does not exist Allowlist; return the error as an observation
Confident emptiness Retrieval returned nothing; answer invented anyway Citation gate — refuse rather than generate
Silent cost blowout Nothing breaks, the bill arrives Per-run dollar cap + cost on every trace

Every one of these is quiet. Not one throws an exception. That is the defining property of agent bugs and the reason the observability layer stopped being optional for me: an agent that fails loudly is a normal engineering problem, and an agent that fails politely is not visible at all without traces.

a real trace, trimmed — the moment it went wrong
run 8f3a2c · goal: "does Filament support nested resources?"

  step 1  think    421 tok   →  search_docs(query="nested resources filament")
          observe    0 hits
  step 2  think    602 tok   →  search_docs(query="filament nested resource")
          observe    0 hits
  step 3  think    780 tok   →  search_docs(query="nested resources in filament v3")
          observe    0 hits
          ^^ REPETITION GUARD would fire here — args differ, intent identical
  step 4  think    966 tok   →  search_docs(query="filament relation manager nested")
          observe    6 hits   ← finally, but 4 steps and 2,769 tokens in

  outcome  answered · 5 steps · 3,441 tok · $0.019 · 6.2s

The repetition guard on exact arguments would not have caught steps 1–3 — the strings differ. What fixed it was the tool description telling the model that an empty first result means the corpus does not cover it. Guards catch the mechanical repeats; good descriptions prevent the semantic ones.

Before you point it at traffic

The checklist

  • A hard step ceiling — a number, in code, not a hope.
  • A per-run cost cap that stops the loop, not just alerts on it.
  • A repetition guard that tells the model it looped.
  • Tool descriptions that state what the tool cannot do, not just what it can.
  • Argument validation on every tool, with the error returned as an observation.
  • Compaction that pins the system prompt and the original goal, never a naive sliding window.
  • Irreversible actions behind a human, with no confidence-score escape hatch.
  • A trace per run recording every step, token count and dollar figure.
  • An answer path that can refuse when the tools came back empty.
  • A termination reason on every run, so "answered" versus "gave up" is a metric and not a guess.

Ten items. Nine of them are twenty lines or fewer. Together they are most of the distance between the loop at the top of this post and something you would let near a customer.

Where to start on Monday

Write the twelve-line loop yourself before you install anything. Not because frameworks are bad — LangGraph and CrewAI earn their place the moment you have several coordinating agents or a long-horizon run — but because every one of them is an opinion about the loop, and opinions are much easier to evaluate once you have held the thing they are opinions about.

Then add exactly one guard, the repetition detector, and watch your traces for a day. It will show you your second guard, and your tools will tell you the rest.

  • tool schemas
  • step budget
  • cost cap
  • repetition guard
  • compaction
  • consequence classes
  • traces
  • termination reasons

If you are building one of these, the production RAG guide covers the retrieval half properly, and the concepts map is worth reading first if the words "agent" and "agentic" still feel interchangeable. And if your agent has found a failure mode that is not in that table of five, I would genuinely like to hear about it — the comments are open.

#AI Tools #LLM #Agents #Agentic AI #LangGraph #Prompt Engineering

Comments (0)

No comments yet

Be the first to share a thought on this article.

Join the conversation

Comments are moderated before they appear.

Keep reading

Related articles

AI 13 min

Connecting OpenAI to a Real Backend: Lessons From SEO Automation

The naive version worked perfectly on fifty pages and failed four ways on twelve thousand — none of them about prompt quality. Wiring a model into a real backend is a batch-processing problem with an expensive, rate-limited, non-deterministic function in the middle, and it needs everything batch processing always needed.

May 28, 2026
AI 15 min

The 940ms Voice Agent: Where Every Millisecond Goes

A conversation has a social latency budget, not a technical one — go past a second and people think the thing is broken. How a voice agent went from 1,530ms to 940ms to first audio, why every millisecond came from overlapping stages rather than speeding them up, the speculative optimisation that cost 38% more tokens for 60ms, and what a user hears when a dependency vanishes mid-sentence.

Apr 12, 2026
AI 14 min

Your AI System Is Failing Right Now and Nothing Is Alerting

On a Tuesday my RAG system got quietly worse. Every request returned 200, no alert fired, and the answers stayed fluent and confident while being subtly about the wrong thing. Why AI systems fail silently, what a trace must actually record, why the mean is a lie, and the forty-minute root cause that would have taken days.

Mar 27, 2026