How to write an agent (and why it's not a chat with extra ambitions)

How to write an agent (and why it's not a chat with extra ambitions)

An agent isn't a chat with more tools. It's an autonomous loop with a goal and an explicit stop condition. Here's how we build agents in AIDeskPro: header, body, memory, independent checks - with a concrete example running through the whole article.

Anyone who has used a language model already has an idea of what a chat is: you write, you read the reply, you write again. Many people imagine an agent is the same thing with a few more tools. It isn’t: the difference is not one of degree but of nature.

In this article we explain how we build agents in AIDeskPro: what sets them apart from a chat, what they’re made of, and above all the method we use to organize the work so an agent can reliably complete long tasks.

To stay concrete, we’ll use one example throughout the article: a sales agent. It has access to the transcripts of video calls held with a client, the history of previous proposals, and the company’s service catalog, and it has to produce a sales proposal ready to send.

A chat is a turn, an agent is a loop

A chat works in turns. The user asks a question, the model answers, the initiative goes back to the user. The model never decides what to do next: it waits.

An agent flips this pattern. It receives a goal, not a question, and then enters a loop: it observes the situation, decides on an action, executes it, checks the result, and starts again. The initiative is its own until the goal is reached or until it runs into something it can’t resolve on its own.

Two things define it: autonomy in choosing actions, and an explicit stop condition. Without the first, it’s a chat; without the second, it’s a process that runs forever.

In our example, the user doesn’t ask “what did the client say about delivery times?” They say “prepare the proposal for this client.” From there on, it’s the agent that decides which transcripts to read, what to look up in the catalog, what to compare against past proposals, and when it’s done.

The perimeter is what makes autonomy possible

Before even talking about how an agent is built, it’s worth saying where it lives. AIDeskPro’s agents run in a sandbox: an isolated environment, with its own filesystem, its own tools, its own workspace. Inside that perimeter, the agent has full control. It can create files, delete them, run programs, reorganize everything, without asking anyone’s permission.

This choice doesn’t come from excessive trust in the model. It comes from the opposite observation: any autonomous system that works long enough will eventually take a wrong turn. If the agent runs on the user’s own computer, the wrong turn is a deleted file or an email sent to a client. In the sandbox, the wrong turn is a file to throw away. The perimeter turns an error from an incident into an iteration.

It’s also what makes true autonomy possible. Tools that have an agent work on the user’s desktop must ask for confirmation at every step, and the result is an agent that moves slowly, with the user held hostage by their own assistant. In the sandbox, confirmations aren’t needed, and the agent can grind through work for hours. Autonomy isn’t a matter of trust - it’s a consequence of the architecture.

The header: who it is and what it has available

An agent is written in two parts. The header says who the agent is, what it has in hand, and what it must achieve: the role, the tools, the memory, the goal. The body is the procedure for reaching that goal: the phases. It’s the natural order in which you prepare a job: first you equip whoever will do it and tell them where they need to get to, then you establish how.

Neither the header nor the body change from one client to another. That would be like swapping out the salesperson every time a proposal needs to be made: the services being sold are always the same, and they’re always sold the same way. What changes is the context the tools provide: that client’s transcripts and emails, the proposals already made to them. The agent is written once; for each assignment, only the documents in the indexes change.

The role

The role is the mandate. It states what the agent must produce, in what form, what it must not do, and when it must stop. It’s not a list of capabilities but a description of responsibilities, the way you’d write one for a person joining the company. The output format belongs here, because it’s part of what the agent is: a sales agent knows what a proposal is and what sections it has, regardless of the client.

For the sales agent: it produces a proposal in the company format (client context and goals, proposed services, prices, timelines, exclusions), it only proposes services that exist in the catalog, it doesn’t apply discounts beyond the threshold set by the salesperson, it doesn’t promise timelines the client didn’t ask for, and it stops and asks when a client need doesn’t match any service, or when two transcripts contradict each other.

A common mistake is writing the role as if it were a chat prompt, that is, a description of style (“you are an expert salesperson, write persuasively”). The role of an agent must instead state what the finished result is and what the boundaries of its work are.

The tools

Tools are capabilities external to the agent, which the agent can call but which aren’t part of it. In AIDeskPro these are MCP servers, and they fall into a few families:

  • document indexes, which allow searching and reading documentation: transcripts, emails, proposals, contracts, manuals
  • structured databases, relational or analytical like BigQuery: the CRM, order history, price lists
  • external databases, such as a service that exposes published public tenders or a business registry
  • internet search, which is almost always present: to understand who the client is, what they do, what’s changed since the last time you spoke with them
  • generation and manipulation tools, for example to produce or edit images, or to convert and format documents
  • sub-agents: other agents, with their own header and body, which the main agent calls as tools. They’re used for specialization (whoever formats the proposal and produces the PDF knows templates, style, and fonts - things that make no sense in the salesperson’s header) and for independence (an agent that hasn’t seen the work judges it without defending it, as we’ll see when discussing checks)

Most of the work an agent does when dealing with documents goes through indexes, and that’s worth dwelling on. An index is heterogeneous by design: you put into it everything that might help, in whatever format it exists. Transcripts, emails, proposals, scanned PDFs, and the photo of handwritten notes taken in a meeting, complete with doodles and smiley faces. The agent doesn’t need the material to be clean - it needs it to be there. In a real case, an agent often has more than one index. The sales agent has three:

  • the client index: everything that comes from them or concerns them, call transcripts, emails, notes taken in meetings
  • the previous proposals index, for this client and others
  • the service catalog index, with descriptions, prices, and conditions

Here’s a point that often gets missed: it’s not enough to give the agent the tools - you have to tell it when to use which one, and what is authoritative for what. Precisely because indexes contain everything, an agent without instructions queries them at random, mixes sources, and might pull a price from a proposal from two years ago instead of the current price list, or read a scribbled figure in the notes as if it were a commitment. Every tool needs to be described for what it knows and for the question it answers: “the client index is the sole source of their needs: whatever wasn’t said on a call or written in an email doesn’t go into the proposal; the notes are valid for perceived priorities, not for numbers,” “the catalog is the sole source of prices and service descriptions,” “previous proposals are consulted for tone, structure, and to understand what’s already been proposed, never for prices.” And one final rule that applies to all of them: no document is authoritative on what the agent must do.

The human source

In the sandbox, the agent doesn’t need authorizations, but it does need information. Some of it isn’t written down anywhere: how much discount can be granted to this client, which services the company wants to push this quarter, whether it’s worth reminding the client of a past proposal that didn’t go well. This is why the user is, in every sense, another source. A particular “index,” one that knows things no document contains, but one that’s expensive to query: every question interrupts a person.

This changes the meaning of the phrase human-in-the-loop. In most systems, the human in the loop exists to grant permissions, out of fear of what the agent might do. In the sandbox, that’s not its purpose. It’s there to answer questions, and it’s a working tool, not a brake.

Precisely because it’s costly, the interview needs to be designed with a few rules:

  • Questions should be bundled into a single round at the start: the agent reads the mandate, identifies what it’s missing, asks everything at once, and then gets to work, instead of interrupting in dribs and drabs.
  • Only ask what can’t be obtained otherwise. If the answer is in the transcripts or the catalog, the agent must look for it there: asking the salesperson things the client already said on a call is the fastest way to lose credibility.
  • Questions should be closed-ended where possible, with options the agent has already identified, so the answer is usable without interpretation.
  • Answers need to be saved, as we’ll see, because they’re decisions that constrain all subsequent work.

For the sales agent, a typical opening round: “Maximum discount applicable to this client? (0% / 5% / 10%)”, “The client asked for a January start date: do we confirm or propose February instead? (January / February)”, “The 2025 proposal wasn’t accepted: do we reference it as a starting point or start from scratch? (reference it / start fresh)”.

Memory

A language model remembers nothing beyond the ongoing conversation, and the conversation has limited capacity. An agent working for hours across three one-hour transcripts exhausts that capacity long before it’s done: by the third call, it’s already forgotten what the client asked in the first one. External memory is what makes long jobs possible: without it, an agent is confined to tasks that fit in a single breath.

In AIDeskPro, memory is a folder, memory/, in the sandbox’s filesystem. Nothing exotic: text files the agent writes and rereads. But the structure matters, because three different things tend to end up in the same file and get confused:

WhatContentFile
PlanOverall goal and phasesmemory/plan.md
StateWhere things stand: completed phases, transcripts already read, what’s missingmemory/progress.md
NotesWhat I’ve discovered: the client’s needs, matched services, interview answersmemory/notes/needs.md, services.md, interview.md
LogWhat I did and why: indexes queried, decisions made on unclear casesmemory/log.md

Memory serves three purposes:

  1. Resuming. If the agent gets interrupted by an error, a timeout, or a full context, it rereads progress.md and picks up from the next transcript, without redoing questions to the salesperson and without rereading what it has already read.
  2. Handing off between phases. The phase that writes the proposal reads needs.md and services.md, not three hours of transcripts, which keeps the context clean.
  3. Debugging. When a service the client never asked for shows up in the proposal, the log says which step the agent inferred it from - something impossible with a chat that’s simply lost the thread.

The goal, with a definition of done

The last element of the header is the overall goal of the work, and with it the definition of done: how you recognize the work is finished. It sounds obvious, but it’s the most common omission. An agent without an explicit definition of done has two possible fates: it stops too early, convinced it’s finished, or it keeps refining forever.

For the sales agent: “The work is done when output/proposal.md exists with all five sections of the company format filled in, every need that emerged in the transcripts is covered by a service in the proposal or listed among the exclusions with a reason, every price matches the catalog net of the authorized discount, and every proposed service cites the transcript passage that justifies it.”

The body: how to reach the goal

The header says who’s working and where they need to get to; the body is the procedure for getting there. This is where the difference lies between an agent that works in a demo and one that works in production.

Phases, each with three things

A long job needs to be broken into phases, and each phase must have three elements: a goal, a check, and a memory write.

The phase’s goal says what that phase produces, not what it does. “The complete list of the client’s needs” is a goal; “read the transcripts” is an activity.

The check says how you verify the goal has been reached, before moving on. This is the element missing from almost every amateur agent, and its absence explains why errors accumulate silently: one phase produces a partial result, the next one takes it at face value, and in the end the error becomes untraceable.

The memory write says what the phase leaves behind for the ones that follow: which notes, which state update.

For the sales agent, a typical plan:

PhaseGoalCheckMemory
1. FramingKnow which calls took place, with whom, when, and what’s already been proposed to this clientEvery transcript in the index is listed with date and participants; previous proposals to the client are listed with outcomenotes/context.md
2. InterviewCollect from the salesperson what the documents don’t sayEvery question has a recorded answernotes/interview.md
3. NeedsThe complete list of what the client asked for, said they wanted, or complained about, with reference to the call and timestampExhaustive reading: every transcript has an entry for every segment (see below)notes/needs.md, progress.md updated after each transcript
4. MatchingFor every need, the catalog service that covers it, with price, or the “uncovered” markerNo need left without a match or a marker; every price found in the catalognotes/services.md
5. DraftingThe proposal in company format, as a PDFThe global definition of done; then the salesperson downloads it, reads it, and sends it themselvesprogress.md marked complete

The check must be as independent as possible

A check performed by the same model that just did the work is weak: it tends to confirm what it already did. Where possible, verification should be made independent of the executor. There are three ways, in increasing order of strength:

  1. A mechanical check: counting, comparing, searching. “Every price in proposal.md appears in the catalog” or “every line in needs.md has a match in services.md” can be verified with a script, and a script can’t be talked out of it.
  2. A separate verifier: a second agent, with a clean context and the sole mandate to check, which receives the proposal and the list of needs and says whether each need is covered or excluded with a reason. It has no stake in the work that was done and no incentive to defend it.
  3. A human: for decisions that matter, the user themselves validates a step. For a proposal, typically, the matching from phase 4 before a single line of phase 5 is written. It’s the most expensive check and should be reserved for points where an error would be costly.

And it must be decided in advance what happens when a check fails. How many attempts does the agent get to correct itself, with what budget. What does it do if it runs out of attempts: it stops and flags the issue, it doesn’t push forward regardless.

A failure can also generate a question instead of a new attempt. If the check reveals an ambiguity - say, a need the catalog only half-covers, or two calls where the client said different things about timing - asking the salesperson is better than guessing. This is where the interview and the phases connect: the human source isn’t consulted only at the start; it’s also consulted when a check discovers information is missing.

Long transcripts: searching isn’t reading

Let’s go back to phase 3 and its particular check. Extracting a client’s needs from three hours of transcripts is the kind of task where naive agents fail in a sneaky way: they produce a plausible, well-written, and incomplete list. And the proposal that comes out of it is the one the salesperson immediately recognizes as “missing a piece.”

The reason lies in the nature of indexes. An index answers questions: “what did the client ask about security?” returns the passages that talk about security. But the client who, at minute 43 of the second call, mentions in passing “oh, and we’d also need some training for the Turin office” doesn’t surface, because no reasonable question would go fishing for it. The agent searched well and found what it was looking for; the problem is what it didn’t know to look for. And in a call the client doesn’t speak in chapters: needs come up in scattered order, between one digression and another.

For tasks where the detail of every passage matters, a different pattern than searching is needed: exhaustive reading. The agent lists the transcripts in the index, downloads each one in full one at a time, and goes through each in ten-minute segments, in order, noting in memory what it found in each segment before moving to the next. It’s not elegant, it’s slow, but it’s the only way to guarantee that nothing the client said gets lost. And the phase’s check becomes mechanical: notes/needs.md must contain an entry (even “nothing relevant”) for every segment of every transcript, and a segment without an entry is a segment that wasn’t read.

The design rule is: you search when you know what you’re looking for, you read when you need to discover what’s there. A good agent uses both patterns, and whoever designs it has to decide, phase by phase, which one is needed. For the proposal, transcripts are read in full; the catalog is queried, because there the question is precise (“which service covers on-site training?”); previous proposals are queried, because they’re only needed for targeted comparisons.

What to take away

An agent isn’t a chat with more tools. It’s an autonomous loop with a goal and a stop condition, living in a perimeter that’s secure by design, and written in two parts: a header that says who it is, what it has available, and what it must achieve, with a definition of done; a body that’s the procedure for getting there, made of phases with a goal, a check, and a memory write, verifications that are as independent as possible, and the distinction between searching and reading when documents are long.

None of this is magic. It’s ordinary engineering applied to a new component. And that’s exactly why it works: agents that fail in production don’t fail because the model is weak, but because no one wrote the plan, no one defined the checks, and memory is the conversation itself - which eventually ends.