26 Jan 2026 · 6 min read · Vivek Amethiya
Agentic AI frameworks: how to choose and best practices
A practical map of agent frameworks, from LangGraph to provider SDKs and low-code builders, and the engineering practices that make agents dependable.
- AI agents
- LangGraph
- MCP
- Best practices
An AI agent is a language model in a loop. It reads a goal, decides on an action, calls a tool, looks at the result and decides again, until the job is done or it gives up. That loop is simple to write and surprisingly hard to make dependable. Frameworks exist to handle the hard parts: state, tool calling, retries, memory, human approval and observability.
There are now more agent frameworks than anyone can evaluate properly. This article groups the main options by what they are good at, then covers the practices that matter far more than the framework you pick.
First question: do you need an agent at all?
Many problems sold as "agentic" are really workflows: a fixed sequence of steps where a model does some of them. Extract fields from an invoice, validate them, write them to the ERP, notify finance. The order never changes, so there is nothing for an agent to decide.
Workflows are cheaper, faster, easier to test and easier to explain to an auditor. Use an agent only when the path genuinely depends on what the model finds along the way: investigating an incident, answering open-ended questions across several systems, or completing multi-step tasks whose steps vary every time. A good rule is to start with a workflow and introduce agentic decisions only at the steps that need them.
The framework landscape
Frameworks fall into a few families. Within each family the options are converging, so the family matters more than the brand.
Graph and state-machine frameworks. LangGraph is the best-known example. You model the agent as a graph of nodes and edges with explicit state, which gives you checkpointing, resumable runs and clean points to pause for human approval. This suits long-running, business-critical processes where you need control over every transition.
Role-based multi-agent frameworks. CrewAI and similar tools let you define agents with roles ("researcher", "writer", "reviewer") that collaborate on a task. They are quick to prototype with and useful when a problem divides naturally into specialist roles.
Model-provider SDKs. The OpenAI Agents SDK, Anthropic's Claude Agent SDK and Google's Agent Development Kit give you an agent loop, tool calling, handoffs between agents and tracing, closely integrated with each provider's models. Microsoft has been consolidating AutoGen and Semantic Kernel into its Agent Framework for the .NET and Python ecosystems. These SDKs are a strong default when you are committed to one model provider.
Data-centric frameworks. LlamaIndex focuses on agents that work over your documents and data, with strong retrieval building blocks. A good fit when most of the agent's work is finding and combining information.
Low-code builders. n8n, Dify and LangFlow let teams assemble agents and workflows visually and connect them to hundreds of services. Excellent for internal automation and fast validation of an idea; for complex, customer-facing logic, most teams eventually move the core into code.
Two open protocols sit underneath all of these. The Model Context Protocol (MCP) standardises how agents connect to tools and data, so a tool built once works across frameworks. Agent2Agent (A2A) standardises how agents built on different stacks talk to each other. Preferring frameworks that support these protocols protects you from lock-in.
How to choose
Ask these questions in order:
- How much control do you need over each step? High control and auditability point to graph frameworks. Exploration and prototyping point to role-based or low-code tools.
- Which models will you use? If one provider, its SDK removes a lot of glue code. If several, choose a model-agnostic framework.
- What does your team already run? A Python data team and a TypeScript product team will be productive in different ecosystems.
- How will you observe and evaluate it? Built-in tracing, or clean integration with tools like Langfuse, should be a hard requirement, not a nice extra.
- Can you leave? Keep prompts, tool definitions and evaluation sets in your own code and formats, so the framework can be swapped without starting over.
Best practices that matter more than the framework
Design tools like public APIs
The model only knows what your tool names, descriptions and schemas tell it. Vague tools produce wrong calls.
{
"name": "get_invoice_status",
"description": "Return the payment status of one invoice. Use when the user asks whether an invoice is paid, overdue or disputed. Do not use for listing invoices.",
"input_schema": {
"type": "object",
"properties": {
"invoice_id": { "type": "string", "pattern": "^INV-[0-9]{6}$" }
},
"required": ["invoice_id"]
}
}Keep each tool narrow, name it by what it does, say when not to use it, validate every input, and return errors the model can act on ("invoice not found, check the id format INV-123456") rather than stack traces. Make tools idempotent where possible, because agents retry.
Bound every loop
An agent without limits can spin forever and spend real money doing it. Put hard budgets on steps, tool calls, tokens and wall-clock time, and fail clearly when they run out:
// Illustrative, framework-agnostic agent loop with guardrails
const MAX_STEPS = 8;
for (let step = 0; step < MAX_STEPS; step++) {
const reply = await model.respond({ messages, tools });
if (reply.type === 'final') return reply.text;
for (const call of reply.toolCalls) {
const tool = registry.get(call.name);
if (!tool) {
messages.push(toolError(call, `Unknown tool "${call.name}"`));
continue;
}
const input = tool.schema.safeParse(call.input);
if (!input.success) {
messages.push(toolError(call, input.error.message));
continue;
}
if (tool.sideEffects && !(await askHumanToApprove(call))) {
messages.push(toolError(call, 'The user rejected this action'));
continue;
}
messages.push(toolResult(call, await tool.run(input.data)));
}
}
throw new Error('Agent stopped: step budget exhausted');Least privilege and human approval
Give each agent only the tools its job requires, with credentials scoped to the minimum. Reading is cheap to allow; writing, sending, paying and deleting should require confirmation until you have strong evidence the agent handles them correctly. Treat every piece of text the agent reads, from web pages to documents to emails, as untrusted input that may contain instructions.
Make state explicit and durable
Long tasks get interrupted. Store the agent's state (goal, steps taken, intermediate results) outside the process so a run can resume after a crash or a human review, instead of starting over and repeating side effects.
One capable agent before many
Multi-agent systems are appealing in diagrams and costly in production: every handoff loses context, adds latency and multiplies token use. Start with one well-equipped agent. Split into several only when there is a clear reason, such as different permissions, different models or a context window that would otherwise overflow. When you do split, give each sub-agent a narrow job and have it return a short summary, not its whole working history.
Trace everything, evaluate continuously
You cannot debug what you cannot see. Record every model call, tool call, input, output, latency and cost for each run. Build an evaluation set of realistic tasks with expected outcomes, and run it on every change to prompts, tools or models. Score both the final result and the path: did the agent use the right tools, and did it stop when it should have?
Control cost by design
- Route simple steps such as classification, extraction and summarisation to smaller models.
- Cache stable prompt prefixes, which most providers now support.
- Trim conversation history and tool outputs before they bloat the context.
- Alert on cost per task, not just total spend.
Ground agents in reliable knowledge
An agent is only as correct as the knowledge it acts on. Use retrieval for large document collections, and curated, reviewed knowledge for the facts that drive actions. Get the knowledge layer right and many "agent bugs" disappear, because the agent was never wrong about what to do, only about the facts it acted on.
A short checklist before production
- The problem genuinely needs agentic decisions, and the rest is a plain workflow.
- Tools are narrow, validated, idempotent and documented for the model.
- Steps, tokens, time and cost are bounded.
- Side effects require approval, and credentials follow least privilege.
- State is durable, so runs can resume safely.
- Every run is traced, and an evaluation set guards every change.
- The design stays portable, with prompts, tools and evaluations owned in your code.
Frameworks will keep changing quickly. These practices won't, and they are what separate an impressive agent demo from a system a business can trust with real work.