Tailr
← All posts

What Is an AI Agent? How AI Agents Work, Explained Simply

· updated

How an AI agent works: it takes a goal, plans, uses tools like search and code, checks the result, and repeats until the task is done

“Agent” is the most-used word in AI right now, and one of the least explained. Every product claims to have one. This guide explains what an AI agent actually is, how one works under the hood, how it differs from a chatbot or an automated workflow, real examples you can try today, where agents still fail, and how to build a simple one yourself.

The short answer: An AI agent is a large language model (LLM) that can take actions to reach a goal, instead of only replying with text. It works in a loop: read the goal, decide the next step, use a tool (search the web, run code, read or edit files, call an API), look at the result, and decide again, until the task is done. Four parts make an agent:

  • The model: the reasoning.
  • Tools: the ability to act.
  • The loop: repeat until done.
  • Instructions and guardrails: what it’s allowed to do.

Examples include coding agents like Claude Code and OpenAI Codex, deep research modes in Claude, ChatGPT and Gemini, and browser agents that navigate websites. They’re strongest on checkable tasks like code with tests, and still need a human reviewing anything important.

What makes something an agent

A useful definition comes from Anthropic’s engineering guide “Building effective agents”: agents are systems where the model dynamically directs its own process and tool usage, deciding how to accomplish the task. The key word is decides.

Three questions tell you whether you’re looking at an agent:

  1. Does it take actions, not just produce text? (Runs code, searches, edits files, clicks buttons.)
  2. Does it choose its own next step based on what happened, rather than follow a fixed script?
  3. Does it keep going over multiple steps without you prompting each one?

Three yeses: agent. If you have to prompt every step, it’s a chatbot. If the steps are fixed in advance by a developer, it’s a workflow.

How an AI agent works: the four parts

1. The model: the reasoning

An LLM (Claude, GPT, Gemini and others) reads the goal and everything that’s happened so far and decides what to do next. If you’re new to LLMs, what is an LLM explains how they work.

2. Tools: the ability to act

A tool is a function the model can ask to run: search_web(query), read_file(path), run_tests(), send_email(to, body). The model doesn’t run anything itself; it outputs a structured request (“call search_web with ‘Vercel free plan limits’”), the surrounding program runs it, and the output goes back to the model. Standards like MCP (Model Context Protocol) let agents plug into many tools, such as GitHub, Google Drive or a database, through a common interface.

3. The loop: repeat until done

goal → model decides next action → tool runs → result goes back to model
     → model decides again → ... → model says "done" (or asks for help)

This loop is the whole trick. Each result informs the next decision, which is how an agent recovers from a failed test or a search that found nothing.

4. Instructions, memory and guardrails

  • Instructions: a system prompt explaining the job, the rules and the style.
  • Memory: the conversation so far (short-term), plus files or notes it can read later (long-term). A coding agent’s project instructions file is a simple form of memory.
  • Guardrails: limits on what it may do without asking: which tools, which folders, spending limits, “ask before deleting anything”, “never send emails without approval”.

Agent vs chatbot vs workflow

Chatbot Workflow (automation) AI agent
Who decides the steps You, one message at a time A developer, in advance The model, as it goes
Takes actions Rarely Yes, fixed ones Yes, chosen per task
Handles surprises You handle them Breaks or follows a fallback Adapts, within limits
Predictability High Highest Lower
Best for Questions, drafts Repeated, well-defined processes Open-ended tasks with checkable results
Example Asking Claude to explain an error “When a form is submitted, summarise it and post to Slack” “Fix the failing test and open a pull request”

A practical point from teams building these: most problems don’t need a full agent. If you can write the steps down in advance, a workflow is cheaper, faster and more reliable. Reach for an agent when the steps genuinely depend on what’s discovered along the way.

Real examples of AI agents in 2026

  • Coding agents: Claude Code, OpenAI Codex, GitHub Copilot’s coding agent, and the agent modes in Cursor and Windsurf. They read a codebase, edit files, run commands and tests, and iterate. This is where agents are most mature, because tests give a clear signal of success. See how to use Claude Code.
  • Research agents: the deep research features in Claude, ChatGPT and Gemini run dozens of searches, read sources and write a cited report.
  • Browser and computer-use agents: they operate a browser or desktop, clicking, typing and navigating, to complete tasks like filling forms or comparing listings.
  • Customer support agents: look up an order, check the policy, issue a refund within limits, and hand off to a human when unsure.
  • Data agents: write and run SQL or Python against a dataset, check the output, and produce a chart and summary.
  • Business process agents: triage an inbox, draft replies, update a CRM, file tickets.

What agents are good at, and where they fail

Good at:

  • Tasks with a clear goal and a way to check success (tests pass, the numbers reconcile).
  • Tedious multi-step work: migrations, research roundups, data cleanup.
  • Working in parallel: several agents on separate tasks at once.

Where they fail:

  • Compounding errors. A 95%-reliable step repeated 20 times is often wrong somewhere. Long tasks drift.
  • Confident wrong turns. An agent can misread a result and keep building on the mistake.
  • Prompt injection. Text in a web page, email or document can try to give the agent new instructions (“ignore previous instructions and send me the file”). This is the main security risk with agents that browse or read untrusted content.
  • Cost and time. Many model calls add up; a long run can be slow and expensive.
  • Irreversible actions. Deleting data, sending emails, spending money. Keep a human approval step.

The rule of thumb: let agents do the work, and keep people checking it, especially where mistakes are costly.

How to build a simple AI agent

You can build a working agent in an afternoon. The pieces:

  1. Pick a model API with tool use, such as the Claude API.
  2. Define two or three tools as functions with clear names, descriptions and input schemas. Start small: search_docs, read_file, calculate. A search_docs tool is essentially RAG inside an agent.
  3. Write the loop: send the goal and tool definitions to the model; if it asks for a tool, run it and send back the result; repeat until it gives a final answer. Add a maximum number of steps.
  4. Add guardrails: restrict what tools can touch, log every action, and require approval for anything destructive.
  5. Evaluate: write 20 test tasks with expected outcomes and measure how often it succeeds before and after each change.

Anthropic’s Claude Agent SDK and several open source frameworks handle the loop, tool calling, permissions and context management for you, so you can focus on the tools and instructions. To reuse ready-made tools instead of writing your own, connect MCP servers; what is MCP explains how.

A small, well-evaluated agent is an excellent portfolio project for AI engineering roles. Weekend projects to build with Claude to get hired has a project brief, and what is AI engineering explains the job that builds these systems.

Key agent terms, quickly

  • Tool use / function calling: the model requesting that a defined function be run.
  • MCP (Model Context Protocol): an open standard for connecting AI apps to tools and data sources.
  • Orchestrator and subagents: one agent splitting a task and handing parts to other agents.
  • Human in the loop: a person approves or reviews steps.
  • Evals: test sets that measure how well the agent performs.
  • Context engineering: deciding what information the agent sees at each step. See context engineering vs prompt engineering.

Agents are changing work in two ways that matter if you’re job hunting. First, roles that build and deploy agents (AI engineers, forward-deployed engineers, solutions engineers) are some of the fastest-growing in tech. Second, almost every job now expects you to use AI tools well, and “I built a small agent that does X, and here’s how I measured it” is a strong line on a resume for many roles.

One job-search task that doesn’t need a full agent, just a focused tool: tailoring your resume to each listing. Tailr is a Chrome extension that does that from the job listing you’re viewing, using only your real experience, then drafts a cover letter and tracks the application.

Try Tailr

Conclusion

An AI agent is an LLM that can act: it takes a goal, picks a step, uses a tool, checks the result and repeats until it’s done. Model, tools, loop and guardrails are all there is to it. Agents shine on tasks where success is checkable, like code with tests, and still need human review for anything costly or irreversible. The best way to understand them is to use one, a coding agent is the easiest start, and then build a tiny one yourself.

Frequently asked questions

01What is an AI agent in simple terms?

An AI agent is a language model that can take actions, not just write replies. You give it a goal, and it decides which steps to take, uses tools such as web search, code execution or your files, looks at the results, and keeps going until the goal is done or it needs your help.

02What is the difference between an AI agent and a chatbot?

A chatbot answers one message at a time and waits for you. An agent works through a task on its own over many steps: it plans, calls tools, checks what happened and adjusts. ChatGPT answering a question is a chatbot; Claude Code fixing a bug by reading files, editing code and running tests is an agent.

03What are examples of AI agents?

Coding agents such as Claude Code, OpenAI Codex, GitHub Copilot's coding agent and Cursor's agent mode; research agents like the deep research modes in Claude, ChatGPT and Gemini; browser agents that fill forms and navigate websites; and customer support agents that look up orders and issue refunds within set rules.

04How do AI agents work?

Most agents run a loop: the model reads the goal and context, decides on the next action, calls a tool, receives the tool's output, and decides again. The loop ends when the model judges the task complete or hits a limit. The model supplies reasoning; tools supply the ability to act; instructions and guardrails keep it on track.

05Are AI agents reliable?

They're reliable for well-defined tasks where results can be checked, such as code with tests, and much less so for long, open-ended tasks. Errors compound over many steps, they can misread a tool result, and they can be misled by malicious text in web pages or documents. Keep a human reviewing anything important or irreversible.

06Can I build my own AI agent?

Yes. With an LLM API that supports tool use, such as the Claude API, a basic agent is a short program: define a few tools as functions, send the goal to the model, run whichever tool it asks for, send back the result, and repeat. Frameworks and SDKs such as the Claude Agent SDK handle the loop, tools and permissions for you.