How agentic systems actually work
A model, a list of tools and a harness running a loop. Why agents succeed, why they fail, and what fixes them.
Meer Habib
Senior Mobile Engineer · Chittagong
"Agent" sounds like magic. Underneath, it's a loop that fits on an index card. Once you see the loop, you can reason about why agents succeed, why they fail, and what makes a tool easy or hard for them to use.
The loop
- The model receives everything in its context: instructions, a list of tools it may call, the conversation so far.
- It answers with either text for the user, or a tool call: "run
npm test", "readsrc/app.ts", "search the docs for X". - The harness, the program around the model, runs that call. The model never touches your files or network directly. The harness does, and it decides what's allowed.
- The result goes back into the context as a tool result.
- Repeat, until the model stops asking for tools and says it's done.
That's it. Claude Code, Codex, a support bot with a refund tool: all the same loop with different tools and different rules.
Tools are just described functions
A tool is a name, a description and a schema for its inputs:
{
"name": "run_tests",
"description": "Run the test suite. Returns failures with file and line.",
"input_schema": {
"type": "object",
"properties": { "filter": { "type": "string", "description": "Only tests matching this name" } }
}
}The model never sees your code for run_tests. It sees this description and decides when to call it. So the description is the interface. A vague description gets vague use. "Returns failures with file and line" tells the model what it will get back, which tells it what it can do next.
The context window is the only memory
Everything the agent knows in the moment is in the context: every file it read, every command output. Two consequences matter more than anything else:
- It fills up. Long sessions get summarised or trimmed. Good harnesses keep the important things, like the goal and the decisions, and drop the noise, like the fourth listing of the same folder.
- Noise costs quality. A tool that returns 5,000 lines when 20 would do doesn't just waste money. It buries the signal. Tools that return short, relevant results make agents noticeably better.
This is also why project instruction files work so well. They put the important things in context at the start of every session.
Why agents fail
When an agent goes wrong, it's usually one of four things:
- It couldn't see. The information it needed wasn't in context, so it guessed.
- It couldn't check. There was no test, type check or screenshot to tell it the work was wrong, so it said "done".
- The tool was ambiguous. Two tools that sound alike, or one tool that does five things depending on flags.
- The step was too big. "Build the app" fails; "add the empty state to this screen, then run the tests" works.
Every one of these is fixed outside the model: better context, a way to verify, clearer tools, smaller steps. That's the work of building for agents.
Bigger shapes, same loop
- Subagents are the loop run again with a smaller job and a fresh context, returning only a summary. Useful for searching a large codebase without flooding the main context.
- Workflows chain several loops with fixed steps in between, like plan, then build, then review.
- Background agents are the same loop on a server, reporting back when they're done.
None of these change the core. If you understand one turn of the loop, you understand all of them.
The short version
- An agent is a model, a list of tools, and a harness running a loop.
- The model decides; the harness acts and enforces the rules.
- Tool descriptions are the interface; results should be short and useful.
- Context is the only memory. Keep it clean.
- Most failures are fixed with better context, verification, clearer tools and smaller steps.
Building something like this?
Booking new projects for Q4. Replies within 24h.