Loop Engineering: The Brain Behind the Agent

On this page
Demystifying the Illusion#
Everyone talks about AI agents like they're little digital minds solving problems. They're not. An LLM is a stateless text predictor. You hand it a string of tokens, it spits back a probability distribution over the next one. No memory between calls. No intent. No thinking.
So where's the brain? It's the loop you build around it.
Strip away every framework — LangGraph, CrewAI, OpenAI Assistants, whatever — and you get this:
state = init(goal, system_prompt, tools)
while not done:
prompt = build_context(state)
resp = llm.call(prompt, tools)
if resp.tool_call:
state.append(execute(resp.tool_call))
elif resp.final_answer:
break
state.iteration += 1
That's the entire Brain. A while-loop with a probabilistic decision function in the middle. The LLM is one node. The tools are the hands. The context window is working memory. And you — the engineer — are the one deciding when the brain stops thinking.
The Loop, Step by Step#
Here's what actually happens in one pass:
- Build the prompt - You assemble the system prompt, tool schemas, conversation history, and the user's goal into one big string. Every iteration you rebuild it, because you get to choose what the model sees this time.
- Call the model - It reads the prompt and outputs either a tool call (structured JSON: "run this query") or a final answer ("here's the result, I'm done").
- Execute or stop - Tool call? Run it — database query, API hit, file read. Final answer? The loop ends.
- Inject the result back - The tool output gets appended to the message history. Next iteration, the model "sees" it and decides what to do.
No hidden state. No magic memory. Just a list of messages that grows one entry at a time, fed into a model that has no idea what happened three iterations ago unless it's still in the context window.
The key mental model: the context window is RAM, not a hard drive. Everything in it, the model can use. Everything not in it, is gone. You didn't save it. You didn't log it. The model simply doesn't know it happened.
Where the Brain Breaks#
This is where loop engineering stops being a tutorial exercise and becomes an actual engineering problem.
The context fills up. A single read_file call can dump 10,000 tokens of code into the history. Twenty iterations later, the window is stuffed. The model re-calls the same API it already ran. It drops evidence it used three iterations ago. It looks "confused." It's not.. A bigger context window helps, but it's not the real fix — the problem is unbounded growth, not a small bucket. So you add compression and externalization: summarize old messages, push bulky results into a cache. The context becomes a cache, not the source of truth.
The loop won't stop. The model calls the same tool with the same args on iterations 11, 12, and 13. It's spinning. You need circuit breakers: a max-iteration cap, a token budget, a wall-clock timeout, and a repetition detector that kills the loop when nothing new is happening.
Tool calls fail. The API times out. The model's JSON is malformed. The tool doesn't exist. Your loop catches it, formats a clean error message, and feeds it back so the model can retry or pivot. Crash on a bad tool call and you've built a brain that dies at the first stubbed toe.
Latency stacks. Each iteration is 2–5 seconds of inference plus tool execution. Ten iterations is a minute-plus of spinner-watching. You need streaming responses, parallel tool calls where possible, and aggressive caching.
The Cost of Thinking#
A loop that runs 12 iterations when 6 would do costs you roughly 50% more in inference (example, not a measurement). Loop engineering is a cost discipline. Every unnecessary iteration is a token paid for a thought the model didn't need. Tight system prompts with clear stopping criteria, smaller tool payloads, and aggressive early termination — that's the difference between a lean system and one burning cash for no reason.
What to Actually Read#
- Anthropic, "Building Effective Agents" (Dec 2024) — the clearest writeup on when a loop beats a single call. Short, practical, no fluff.
- LangGraph state-machine docs — see the loop formalized as explicit nodes and edges instead of a while-loop.
- OpenAI function-calling spec — the actual JSON contract between your loop and the model. Shorter than you think.
The brain isn't the model. The brain is the loop, the memory management, the termination logic, and the error handling you build around it. The model is just the next-token function inside your while-loop. And the loop is what you actually own.
Further reading#
- Harness Engineering: The Harness Around the Brain — the five subsystems that wrap this loop for production.
- LLM Context Windows and Memory — the "RAM" the loop has to manage every iteration.
- Hidden Risks of LLM APIs — what can go wrong once the loop starts calling tools with real credentials.
Keep reading
Where this fits
This article, its topic, and the closest related reading. See the full map →
- Harness Engineering: The Harness Around the Brain
- LLM Context Windows and Memory: How Models Handle Extended Dialogues
- Build LLM Vocab: Tokens, Embeddings, Vocabulary Size, and Context Windows Explained
- Fine-Tuning vs Retrieval-Augmented Generation (RAG): Choosing the Right Approach for Custom LLMs
- Why Quantized Small Models Will Dominate AI's Future — An Honest Take
- Fine-Tune an LLM on Your MacBook with LoRA: A Hands-On Guide
Stay in the loop
New articles on AI, Cybersecurity, and PKI — delivered to your inbox.
No spam. Unsubscribe any time. By subscribing you agree to our Privacy Policy.