Technology

How to Build an AI Agent: Practical Guide With Python Code Examples

How to Build an AI Agent: A Practical Guide With Code Examples

AI agents are moving beyond simple chatbots. Instead of only generating text, an AI agent can reason about a task, use tools, access data, make decisions, and take actions.

An ordinary chatbot might answer:

“What is the weather today?”

An AI agent can do much more:

“Check today’s weather, compare it with tomorrow’s forecast, find a suitable time for my outdoor meeting, and add it to my calendar.”

The difference is agency.

An AI agent combines a language model with instructions, tools, memory, data and an execution loop.

In this guide, we’ll build a simple AI agent from scratch and gradually add:

  • Tool calling
  • External APIs
  • Memory
  • Multi-step reasoning
  • Error handling
  • Security controls
  • Production considerations

The examples use Python, but the same architecture can be implemented in JavaScript, PHP, Go and other languages.


What Is an AI Agent?

An AI agent is a software system that uses an AI model to determine what actions should be taken to accomplish a goal.

A simplified architecture looks like this:

User

↓

AI Model

↓

Decision

↓

Tool

↓

Result

↓

AI Model

↓

Final Answer

The important part is the loop.

A traditional application follows predefined instructions:

Input → Function → Output

An AI agent can dynamically choose:

Goal
 ↓
Understand
 ↓
Choose action
 ↓
Use tool
 ↓
Observe result
 ↓
Choose next action
 ↓
Complete task

This makes agents useful for tasks that require multiple steps.


AI Agent vs. Chatbot

The terms are often used interchangeably, but there is an important difference.

Capability Chatbot AI Agent
Generate text Yes Yes
Answer questions Yes Yes
Use tools Sometimes Core capability
Access external systems Limited Yes
Multi-step tasks Limited Yes
Make decisions Limited Yes
Take actions Usually no Yes
Maintain task state Limited Yes
Execute workflows Limited Yes

A chatbot primarily responds.

An agent works toward a goal.


The Five Components of an AI Agent

A practical agent normally contains at least five components.

1. Model

The AI model provides reasoning and language capabilities.

Examples include:

  • GPT models
  • Claude
  • Gemini
  • Open-source models

The model determines what it should do based on the available context and tools.


2. Instructions

The agent needs a system-level definition of its role.

For example:

You are a server-management assistant.

You can inspect server status and restart services.

Never delete files or restart production servers without confirmation.

These instructions define the agent’s behavior and boundaries.


3. Tools

Tools allow the agent to interact with the outside world.

Examples:

get_server_status()
restart_service()
search_database()
send_email()
create_ticket()
check_order()

Without tools, an AI model can only generate information.

With tools, it can perform operations.


4. Memory

Memory allows an agent to retain information relevant to the task.

Examples include:

Customer preferences
Previous conversation
Current task state
Previous tool results
Account configuration

Memory can be short-term or persistent.


5. Agent Loop

The agent needs an execution loop that allows it to:

  1. Understand the goal.
  2. Decide what to do.
  3. Call a tool.
  4. Read the result.
  5. Decide what to do next.
  6. Finish the task.

This loop is the core of an agent.


Building Our First AI Agent

Let’s build a simple Python agent.

Our agent will have two tools:

get_weather()
calculate()

The user might ask:

“What’s the weather in London and what is 25% of 240?”

The agent can determine that it needs two tools before answering.


Step 1: Install the OpenAI SDK

Create a Python environment:

python -m venv .venv

Activate it on Linux/macOS:

source .venv/bin/activate

On Windows:

.venv\Scripts\activate

Install the SDK:

pip install openai

Set your API key as an environment variable:

export OPENAI_API_KEY="your_api_key"

On Windows PowerShell:

$env:OPENAI_API_KEY="your_api_key"

Never hard-code API keys directly into your source code.


Step 2: Create a Simple Agent

Start with a basic model call:

from openai import OpenAI

client = OpenAI()

response = client.responses.create(
    model="gpt-5.6",
    input="Explain what an AI agent is in two sentences."
)

print(response.output_text)

At this stage, this isn’t really an agent yet.

It is simply an AI model.

The important part comes next.


Step 3: Add a Tool

Let’s create a calculator tool.

def calculate(expression: str) -> float:
    return eval(expression)

However, using eval() directly on arbitrary AI-generated input is dangerous.

For a real application, use a safe expression parser or restrict the allowed operations.

For demonstration purposes, we can create a safer calculator with explicit operations:

def calculate(operation: str, a: float, b: float) -> float:
    if operation == "add":
        return a + b

    if operation == "subtract":
        return a - b

    if operation == "multiply":
        return a * b

    if operation == "divide":
        if b == 0:
            raise ValueError("Cannot divide by zero")
        return a / b

    raise ValueError("Unsupported operation")

Now the agent has something it can actually execute.


Step 4: Describe the Tool to the Model

The model needs to know:

  • What the tool does
  • When to use it
  • What parameters it accepts
  • What those parameters mean

A tool definition can look like:

tools = [
    {
        "type": "function",
        "name": "calculate",
        "description": "Perform a basic mathematical calculation.",
        "parameters": {
            "type": "object",
            "properties": {
                "operation": {
                    "type": "string",
                    "enum": [
                        "add",
                        "subtract",
                        "multiply",
                        "divide"
                    ]
                },
                "a": {
                    "type": "number"
                },
                "b": {
                    "type": "number"
                }
            },
            "required": [
                "operation",
                "a",
                "b"
            ]
        }
    }
]

The model doesn’t execute this function.

It decides when the function should be called.

Your application executes it.

That distinction is extremely important.


Step 5: Execute the Tool

A simplified agent loop looks like this:

response = client.responses.create(
    model="gpt-5.6",
    input="What is 25% of 240?",
    tools=tools
)

for item in response.output:
    if item.type == "function_call":
        print("Tool requested:", item.name)

The model may return a function call with arguments such as:

{
  "operation": "multiply",
  "a": 240,
  "b": 0.25
}

Your application then executes the function.

result = calculate(
    operation="multiply",
    a=240,
    b=0.25
)

The result is then sent back to the model.


Step 6: Complete the Agent Loop

The basic pattern is:

while True:

    response = client.responses.create(
        model="gpt-5.6",
        input=messages,
        tools=tools
    )

    function_calls = [
        item for item in response.output
        if item.type == "function_call"
    ]

    if not function_calls:
        print(response.output_text)
        break

    for call in function_calls:

        if call.name == "calculate":

            args = json.loads(call.arguments)

            result = calculate(**args)

            messages.append(call)

            messages.append({
                "type": "function_call_output",
                "call_id": call.call_id,
                "output": str(result)
            })

This is the fundamental architecture behind many tool-using agents.


Adding a Real API

Let’s make the example more realistic.

Suppose we want an agent that can check server status.

Create a function:

import requests

def get_server_status(server_url):
    response = requests.get(
        server_url,
        timeout=10
    )

    return {
        "status_code": response.status_code,
        "online": response.ok
    }

The agent can now decide when to call it.

For example:

“Check whether https://example.com is online.”

The model determines:

Need server status
        ↓
Call get_server_status
        ↓
Receive HTTP 200
        ↓
Tell user server is online

Building a Hosting Support Agent

This architecture becomes especially useful for hosting providers.

Imagine an AI support agent with tools such as:

tools = [
    "get_server_status",
    "check_disk_usage",
    "check_memory",
    "check_mysql",
    "check_apache",
    "check_nginx",
    "restart_service",
    "create_support_ticket"
]

A customer might say:

“My website is very slow.”

Instead of simply generating troubleshooting instructions, the agent could:

1. Identify the server.
2. Check CPU usage.
3. Check RAM.
4. Check disk usage.
5. Check web-server status.
6. Check database status.
7. Inspect recent errors.
8. Determine likely cause.
9. Recommend a solution.

With appropriate permissions, it could then take corrective action.


Giving an Agent Access to a Database

Agents can also work with structured data.

For example:

def get_customer(customer_id):

    query = """
        SELECT id, name, status
        FROM customers
        WHERE id = ?
    """

    return database.execute(
        query,
        (customer_id,)
    )

The agent could then answer:

“Is customer 1024 active?”

Instead of requiring the user to manually search the database.

However, database access must be carefully controlled.

Never give an AI agent unrestricted SQL access to a production database.

Prefer predefined functions such as:

get_customer()
get_invoice()
get_subscription()
get_server()
get_order()

rather than:

execute_any_sql()

Agent Memory

A useful agent needs context.

For example:

conversation = [
    {
        "role": "user",
        "content": "My server is web01."
    },
    {
        "role": "assistant",
        "content": "I'll use web01 for the diagnostic checks."
    }
]

Short-term memory can be stored in the current request.

Persistent memory can be stored in:

  • PostgreSQL
  • MySQL
  • Redis
  • Vector databases
  • Application-specific storage

The important rule is:

Do not store everything.

Store information that is useful for future tasks.


Vector Databases and Retrieval

Agents sometimes need access to large knowledge bases.

For example, a hosting company might have:

Server documentation
Product documentation
Troubleshooting guides
Security policies
Customer documentation
Internal procedures

Instead of placing all of that information into every prompt, use retrieval.

The basic architecture becomes:

User question
       ↓
Search knowledge base
       ↓
Relevant documents
       ↓
AI model
       ↓
Answer

This approach is commonly called retrieval-augmented generation (RAG).

An agent can combine RAG with tools:

Question
 ↓
Search documentation
 ↓
Check server
 ↓
Analyze results
 ↓
Respond

Multi-Agent Systems

Not every task needs multiple agents.

But complex workflows can be divided into specialized agents.

For example:

                Main Agent
                    |
       ┌────────────┼────────────┐
       ↓            ↓            ↓
   Security      Database     Infrastructure
    Agent          Agent         Agent

A security agent could investigate:

  • Malware
  • Failed logins
  • Firewall events

A database agent could investigate:

  • Slow queries
  • Connections
  • Database errors

An infrastructure agent could investigate:

  • CPU
  • RAM
  • Disk
  • Network

The main agent coordinates the workflow.


Don’t Build a Multi-Agent System Too Early

A common mistake is making everything an agent.

If a task can be handled with:

if condition:
    do_something()

you probably don’t need an AI agent.

Agents are most useful when:

  • The workflow is variable.
  • Decisions depend on context.
  • Natural-language input is involved.
  • Multiple tools may be required.
  • The next action cannot easily be predetermined.

Use normal software for deterministic operations.

Use AI where reasoning and ambiguity actually add value.


Add Guardrails

This is one of the most important parts of production AI agents.

Imagine giving an agent this tool:

delete_server(server_id)

That is dangerous.

The model might misunderstand a request.

Instead, separate high-risk actions.

For example:

inspect_server()

can be automatic.

But:

delete_server()

requires explicit approval.

A safer architecture is:

AI decides
     ↓
Risk check
     ↓
Human approval
     ↓
Tool execution

Tool Permissions Should Be Granular

Don’t give every agent every permission.

For example:

Read-only agent

✓ CPU status
✓ RAM status
✓ Disk status
✓ Logs
✗ Restart
✗ Delete
✗ Modify

Operations agent

✓ Read
✓ Restart service
✓ Clear cache
✗ Delete server

Administrator agent

✓ Read
✓ Modify
✓ Restart
✓ Deploy

The agent should receive only the permissions required for its task.


Validate Tool Arguments

Never assume that because an AI generated a function call, the arguments are safe.

Validate everything.

For example:

def restart_service(service):

    allowed = {
        "nginx",
        "apache2",
        "php8.3-fpm"
    }

    if service not in allowed:
        raise ValueError("Service is not allowed")

    # Execute restart

This prevents an AI model from accidentally invoking arbitrary commands.


Never Give the Agent Raw Shell Access

Avoid tools such as:

run_command(command)

unless you have extremely strong isolation and security controls.

Instead of:

run_command("rm -rf /var/www")

provide narrowly scoped functions:

get_disk_usage()
restart_nginx()
clear_application_cache()
check_php_status()

This dramatically reduces the attack surface.


Logging and Observability

Every production agent should record:

User request
Model response
Tool selected
Tool arguments
Tool result
Final response
Execution time
Errors
Approval events

For example:

{
  "request_id": "req_123",
  "tool": "restart_nginx",
  "arguments": {
    "server": "web01"
  },
  "approved": true,
  "result": "success"
}

This becomes extremely useful when investigating failures.


Handling Tool Failures

Tools will fail.

A server might be offline.

An API might timeout.

A database might reject a query.

Your agent should handle these conditions gracefully.

try:

    result = check_server(server)

except TimeoutError:

    result = {
        "status": "unknown",
        "error": "Server check timed out"
    }

The model can then explain:

“I couldn’t verify the server because the monitoring endpoint timed out.”

This is much better than inventing a result.


AI Agents Should Never Invent Tool Results

This is a critical rule.

If the agent calls:

check_server()

and the tool returns:

timeout

the agent should not say:

“The server is online.”

It should say:

“I couldn’t verify the server because the status check timed out.”

The application should make tool results authoritative.


Production Architecture

A production AI agent can look like this:

                 User
                   |
                   ↓
              API Gateway
                   |
                   ↓
             Agent Service
                   |
        ┌──────────┼──────────┐
        ↓          ↓          ↓
      Model      Memory     RAG/Search
        |
        ↓
    Tool Router
        |
   ┌────┼─────┬─────┐
   ↓    ↓     ↓     ↓
 Server DB   CRM   Email
   |
   ↓
Infrastructure

This separation makes the system easier to secure and maintain.


Example: A Complete Agent Workflow

Imagine a customer asks:

“My website is down. Please check it.”

The agent might execute:

Step 1

Identify the website.

example.com

Step 2

Check DNS.

DNS: OK

Step 3

Check HTTP.

HTTP: 502

Step 4

Check NGINX.

NGINX: running

Step 5

Check PHP-FPM.

PHP-FPM: stopped

Step 6

Determine likely cause.

502 Bad Gateway caused by unavailable PHP-FPM.

Step 7

Ask for approval.

PHP-FPM is stopped. Would you like me to restart it?

Step 8

User approves.

Restart PHP-FPM.

Step 9

Verify.

HTTP: 200

Step 10

Report.

The website is back online.
The issue was caused by PHP-FPM being stopped.

This is a genuine agent workflow.


AI Agent vs. Automation

A traditional automation might be:

Every 5 minutes
 ↓
Check PHP-FPM
 ↓
If stopped
 ↓
Restart PHP-FPM

An agent could be:

Website problem
 ↓
Investigate
 ↓
Check HTTP
 ↓
Check NGINX
 ↓
Check PHP-FPM
 ↓
Check logs
 ↓
Determine cause
 ↓
Decide whether restart is appropriate

Automation is deterministic.

Agents are adaptive.

In many production systems, the best solution is actually both.


Where AI Agents Are Most Useful

AI agents are particularly valuable for:

IT Operations

Monitor → Diagnose → Recommend → Remediate

Customer Support

Question → Search knowledge → Check account → Respond

Cybersecurity

Alert → Investigate → Correlate → Prioritize → Escalate

Sales

Lead → Research → Qualify → CRM update → Follow-up

Finance

Transaction → Analyze → Categorize → Flag anomalies

DevOps

Incident → Inspect → Diagnose → Suggest fix → Deploy

Research

Question → Search → Analyze → Compare → Summarize

Common AI Agent Mistakes

Mistake 1: Giving the agent too many tools

More tools don’t automatically make an agent better.

A smaller, well-designed toolset is easier for the model to use reliably.


Mistake 2: Giving unrestricted permissions

Never give an agent more access than it needs.


Mistake 3: Using AI for deterministic logic

Don’t use an LLM to calculate a fixed value that a normal function can calculate reliably.


Mistake 4: No validation

Every tool input should be validated.


Mistake 5: No human approval

High-impact actions should require explicit authorization.


Mistake 6: No monitoring

You need to know:

  • What the agent did
  • Why it did it
  • Which tools it used
  • What failed
  • How long it took

How to Build Your First Production Agent

A practical development path is:

Phase 1 — Basic model

User → AI → Response

Phase 2 — Add one tool

User → AI → Tool → AI → Response

Phase 3 — Add multiple tools

User → AI → Tool A → Tool B → AI

Phase 4 — Add knowledge retrieval

User → Search → AI → Tools

Phase 5 — Add memory

User → Memory + Knowledge + Tools → AI

Phase 6 — Add permissions

AI → Permission Check → Tool

Phase 7 — Add human approval

AI → Approval → Tool

Phase 8 — Add monitoring

Logs + Metrics + Traces + Auditing

This incremental approach is much safer than attempting to build a fully autonomous agent immediately.


Final Thoughts

Building an AI agent is not simply about connecting an LLM to an API.

The model is only one component.

A useful production agent combines:

Model + Instructions + Tools + Memory + Retrieval + Permissions + Guardrails + Observability

The most important design principle is simple:

Let AI decide what should happen, but let your application control what is actually allowed to happen.

Start small.

Give the agent one clearly defined task.

Add one or two tools.

Validate every tool call.

Add human approval for risky actions.

Then expand gradually.

The future of AI agents is not necessarily fully autonomous software making unrestricted decisions.

The most practical systems will be AI-assisted applications that combine reasoning with controlled access to real-world tools and data.

That is where AI agents become genuinely useful: not just answering questions, but helping people investigate, decide, execute and verify.

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top