How to build an AI agent that holds in production: the 6 steps, from agent vs workflow choice to permissions, often overlooked by tutorials.

Building a functional AI agent takes an afternoon with Claude API and a tool like n8n. The real work starts after: deciding what this agent is allowed to do, and especially what it's not allowed to do. Most tutorials skip this step, and that's exactly where things break in production.
This guide walks through six steps to build a solid AI agent, from task definition to monitoring, with a focus on the point most guides ignore: permissions.
1. Define the task: agent or simple workflow?
Many projects fail before you write a single line of code, simply because an agent was chosen where a deterministic function would have sufficed.
Tech Insider, Microsoft Agent Framework Setup, September 2026 "decide deliberately between an agent and a workflow rather than defaulting to whichever one you learned first. If you can write a plain function to handle a task deterministically, do that instead."
According to Tech Insider, the choice between agent and workflow must be deliberate. If a task can be solved with standard code, an if, a loop, a direct API call, there's no reason to bolt an LLM onto it.
An AI agent makes sense when the task requires non-deterministic reasoning: interpreting an ambiguous customer request, choosing between multiple tools based on context, adapting to an unexpected response. If your "agent" just chains three fixed steps, call it a script. It will cost less, and it will crash less.
Pitfall: building an agent with a full reasoning loop for a task that would have fit in 15 lines of standard code. Result: latency multiplied, costs climbing, behavior less predictable than a simple
switch/case.
2. Choose the model and orchestration architecture
Once the task is validated as truly agentive, you need to choose the model and how to orchestrate its tool calls.
Claude API remains a solid choice for cases requiring multi-step reasoning, with native agentic loop: the agent itself decides whether to call a tool, reads the result, then repeats. On the orchestration side, frameworks like LangChain or CrewAI structure this loop with reusable components, but none of them dispense you from precisely defining the scope of action before coding anything.
And it's precisely this scope that determines whether the agent will hold in production or not.
3. Connect tools via MCP rather than custom integrations
The Model Context Protocol (MCP), an open standard introduced by Anthropic, allows an agent to connect to external tools, databases, CRMs, password managers, through a unified interface rather than a custom integration coded for each tool.
The standard is progressing quickly. By late September 2026, 1Password released an MCP server exposing its "Environments" to any compatible agent, allowing a Claude Code agent to retrieve credentials without ever exposing them in plaintext in the code. Cursor, meanwhile, integrated "Agent Skills", an open format backed by Anthropic, directly into its configuration panel, alongside rules and MCP servers.
Concretely: instead of coding a custom connector to each API, you describe the tool once in an MCP server, and any compatible agent can use it afterward. We covered this already in our definition of the AI agent, MCP is what transforms an isolated agent into an agent truly connected to an ecosystem of tools.
4. Limit permissions from the start, not after the fact
This is where most guides stop too early. Giving an agent complete access "so it works on the first try" is the decision that costs the most six months later.
Hostinger, Claude Tutorial, September 2026 "Match each agent's access to its actual role. If it only sends outbound messages, don't let it delete emails or change inbox rules."
According to Hostinger, an agent's access must strictly match its role. An agent that replies to client emails has no reason to be able to delete them, or modify inbox rules.
But the real question isn't technical. It's organizational: who decides, in your team, what level of rights to grant each agent? Three levels to settle before writing the first system prompt:
| Level | Concrete example | Risk if granted too quickly |
|---|---|---|
| Read-only | Check order status | None, it's the recommended default |
| Limited write | Create a ticket, send an email | Duplicates, spam |
| Broad write | Modify a customer database, delete | Irreversible data loss |
An agent should never start with broad write access. Increase permissions gradually, once behavior in read-only mode has been observed on real cases, not on demo cases.
5. Test in a sandbox before any real connection
Before plugging the agent into real data, run it against a test dataset, with complete logs of every decision it makes. An agent that behaves well almost all the time in testing can crash right when it matters most in production, and this is often the most expensive case.
We dug into this in our article on AI agent security in production: attackers specifically target poorly scoped agents, those given more room to maneuver than necessary.
6. Monitor, measure, iterate
An AI agent is never "done" at the moment of deployment. Track three metrics from day one: the success rate of tool calls, the average cost per completed task, and the number of human interventions needed to correct an agent's decision.
If that last number isn't dropping after two or three weeks, the problem probably isn't the model. It's the task scope that's poorly defined.
Conclusion
Building an AI agent that works in a demo is now accessible in a few hours. Distinguishing it from an agent that holds in production requires three things: a truly agentive task, tools connected via a standard like MCP rather than workarounds, and permissions thought through before the first deployment, not after an incident.
At fstck, we help teams move from an AI agent prototype to a truly reliable custom internal tool. If your project is at that stage, now's the time to discuss it.
Frequently asked questions
What's the difference between an AI agent and a simple automation script?
A script follows a fixed sequence of instructions. An AI agent decides itself, at each step, which tool to call and in what order, based on context. If the behavior is always identical regardless of input, it's not an agent, it's a script in disguise.
Is it necessary to use MCP to connect an agent to its tools?
No, you can code tool-specific integrations. But MCP saves you from rewriting a connector for each new tool: the standard describes the interface once, and any compatible agent can use it afterward, including ones you'll add later.
How long does it take to create a basic AI agent with Claude API?
A functional prototype, an agent that calls one or two tools via the Claude API, can be built in a few days of development. The part that really takes time isn't the code: it's precisely defining the scope of action and testing it before connecting it to real data.
How do you limit damage if an AI agent makes a bad decision?
Upstream, by limiting its permissions to the bare minimum for its task, read-only by default, write added gradually. Downstream, by keeping complete logs of every decision, to identify and quickly fix the source of the problem.
Can an AI agent completely replace a developer on a business task?
Not in most cases observed today. A well-scoped agent automates a repetitive portion of a task, but defining the scope, supervising, and correcting errors remain human work, at least for now.


