Senior AWS Consultant
What is an AI Agent?
An AI agent is a program that uses a large language model (LLM) not just to answer questions, but to take actions. Instead of generating text and stopping, an agent can decide to call external functions — query a database, fetch a file, call an API — and then use the results to formulate its final answer.
The basic loop looks like this:
- User sends a question.
- The LLM reads the question and the list of available tools.
- The LLM decides: can I answer directly, or do I need to call a tool first?
- If it calls a tool, it receives the tool’s output and continues reasoning.
- Steps 3–4 repeat until the LLM has enough information to respond.
- The LLM produces a final answer for the user.
This “reason → act → observe → repeat” loop is what separates an agent from a simple chatbot. The LLM becomes the decision-maker that orchestrates tool calls autonomously.
AgentCore vs Bedrock Agents: Why a New Service?
AWS is retiring Bedrock Agents, which was already a fully managed agent service where you configure tools via the console, attach action groups, and AWS handles orchestration. So why a new service? Bedrock Agents works well for standard use cases, but the orchestration logic is a black box. You can’t change how the agent loops, which framework it uses, or how tools are routed.
Bedrock AgentCore is the infrastructure layer underneath. It gives you:
- Your own container — run any agent framework (Strands, LangChain, CrewAI, or custom code)
- Full orchestration control — custom hooks, iteration limits, short-circuit logic, anything
- MCP protocol for tools — the open Model Context Protocol standard, not a proprietary format
- Managed compute — you don’t run servers; AWS scales the containers for you
| Bedrock Agents | Bedrock AgentCore |
Orchestration | AWS-managed (fixed) | Your code, any framework |
Tool protocol | Action Groups (OpenAPI) | MCP (open standard) |
Deployment | Config-only (no container) | Your Docker image |
Customization | Console knobs | Unlimited — it’s your code |
AgentCore Components Explained
AgentCore might look overwhelming with all it’s different parts, so here I’m just focusing on the bare minimum of resources you need to get an agent up and running. AgentCore has three core components. Understanding what each one does (and doesn’t do) is the key to getting your first agent running.
- The Runtime
The Runtime is your Docker container running on AWS-managed compute. It’s where your agent logic lives — the Python code, the LLM calls, the orchestration loop.
Think of it as a Fargate task that AWS manages for you: it pulls your image from ECR, starts the container, and routes incoming invocations to it. You don’t manage instances, scaling, or networking (unless you choose VPC mode for private resources).
What it does:
- Runs your agent code (any language, any framework)
- Receives invocations via the invoke-agent-runtime API (IAM-authenticated)
- Has an IAM role that lets it call Bedrock (for the LLM) and the Gateway (for tools)
What it doesn’t do:
- It doesn’t decide which tools to call — that’s your LLM
- It doesn’t route tool calls — that’s the Gateway
- It doesn’t execute tools — that’s the tool backend (Lambda, an API, or another service)
Configuration that matters:
- server_protocol = “HTTP” — your container is an HTTP server, not an MCP server
- network_mode = “PUBLIC” — AWS handles outbound internet; use “VPC” only if you need private network access
- Container must be ARM64 architecture
- The Gateway
The Gateway is an MCP (Model Context Protocol) endpoint that routes tool calls from your Runtime to the backends that execute them. It’s the router between “the LLM decided to call a tool” and “the tool actually runs.”
When your agent’s MCP client sends a tool call, the Gateway:
- Verifies the request (AWS_IAM SigV4 authentication)
- Looks up which target matches the tool name
- Invokes that target’s backend with the tool’s parameters
- Returns the response back to your agent
What it does:
- MCP protocol endpoint (tool discovery + tool execution)
- Routes each tool name to the correct backend (Lambda, an external MCP server, or other AWS services)
- Authenticates every request using IAM
What it doesn’t do:
- It doesn’t decide which tool to call — that’s the LLM
- It doesn’t run tool logic — that’s whatever backend the target points to
- It has no knowledge of your agent’s conversation or state
- Gateway Targets
A Gateway Target is the binding between a tool name and a backend that executes it. For each tool you want your agent to use, you create one target that says: “when the MCP client calls tool X, route it to backend Y.”
The backend doesn’t have to be a Lambda function, Gateway Targets support multiple target types. Lambda is the most common for custom logic, but you can also point targets at other MCP-compatible endpoints or AWS services. In this tutorial we use Lambda because it’s the simplest way to run custom code without managing servers, but keep in mind that the Gateway is a generic MCP router, not a Lambda-specific feature.
Each target includes a tool schema — the name, description, and input parameters that get advertised to the LLM. When your agent calls list_tools_sync(), it receives these schemas, and that’s how the LLM knows what tools are available and what they do.
Pattern for Lambda targets: one Lambda, many aliases, many targets. You don’t need a separate Lambda per tool — use aliases (alias name = tool name) and dispatch inside the handler.
What We’re Building in this Tutorial
A minimal agent that:
- Runs Claude Sonnet in a container managed by AgentCore
- Has one MCP tool (GetCakeRecipe) backed by a Lambda function
- Deploys entirely via terraform apply
Ask it “Can you give me a cake recipe?” and it calls the tool, gets a Schwarzwälder Kirschtorte recipe, and formats the answer. Ask it anything else and it answers from its own knowledge without calling the tool.
Architecture
invoke-agent-runtime (AWS API, SigV4-signed)
│
▼
AgentCore Runtime (your Docker container)
└─ Strands Agent + Claude Sonnet via Bedrock
│
│ MCP over HTTPS (SigV4-signed automatically)
▼
AgentCore Gateway (AWS_IAM auth, MCP protocol)
│
│ Invokes tool backend
▼
Tool Backend (Lambda function in this example)
└─ Returns static recipe JSON
The Python Code in Detail
The Tool Function
The simplest possible tool receives whatever parameters the LLM passes and returns a JSON response:
# src/lambda/tool/handler.py
def handler(event, context):
return {
"status": "ok",
"recipe": {
"name": "Oma's Schwarzwälder Kirschtorte",
"servings": 12,
"ingredients": [
"200g dark chocolate", "200g butter", "200g sugar",
"5 eggs", "150g flour", "2 tsp baking powder",
"500ml heavy cream", "3 tbsp kirsch (cherry brandy)",
"1 jar sour cherries (drained)", "chocolate shavings",
],
"steps": [
"Melt chocolate and butter, let cool.",
"Beat eggs and sugar until fluffy, fold in chocolate.",
"Fold in flour + baking powder. Bake 175°C, 25 min.",
"Slice into 3 layers. Whip cream with kirsch.",
"Layer: cake, cream, cherries. Repeat. Decorate.",
],
},
}
In a real agent, this function would query a database, call an API, or do anything you need. We use a Lambda function here because it’s the simplest serverless option, but Gateway Targets can also point to other backends. The contract is the same regardless: receive parameters, return JSON.
The Agent Container (app.py)
This is the brain. Let’s walk through each piece:
app = BedrockAgentCoreApp()
_agent = None
The app is the HTTP server AgentCore talks to. The _agent is cached globally so we don’t rebuild it (and re-discover tools) on every invocation — the agent’s MCP connection persists across requests within the same container lifetime.
def _get_agent():
global _agent
if _agent is not None:
return _agent
mcp_tools = []
if GATEWAY_URL:
mcp = MCPClient(lambda: aws_iam_streamablehttp_client(
endpoint=GATEWAY_URL,
aws_region=AWS_REGION,
aws_service="bedrock-agentcore",
))
mcp.start()
mcp_tools = list(mcp.list_tools_sync())
_agent = Agent(
model=model,
tools=mcp_tools,
system_prompt="You are a helpful assistant. Use the GetCakeRecipe tool when the user asks for a recipe.",
)
return _agent
MCPClient takes a factory function (the lambda:) that creates the transport connection. This is lazy — the actual HTTPS connection to the Gateway is only established when mcp.start() is called.
mcp.start() opens the connection and keeps it alive for subsequent calls.
mcp.list_tools_sync() asks the Gateway “what tools do you have?” and gets back the tool schemas you defined in Terraform. This is how the agent discovers tools dynamically — add a new tool in Terraform, and the agent picks it up on next startup without a code change.
Agent(…) wires everything together: the LLM (model), the available actions (tools), and the behavioral instructions (system_prompt). The system prompt tells the LLM when to use the tool — without it, the LLM might answer recipe questions from its own training data instead of calling GetCakeRecipe.
@app.entrypoint
def invoke_agent(payload: dict):
prompt = payload.get("prompt", "")
agent = _get_agent()
result = agent(prompt)
return {"result": str(result)}
@app.entrypoint registers this function as the handler for incoming invocations. When someone calls invoke-agent-runtime, AgentCore delivers the payload here.
agent(prompt) — this single line triggers the full agent loop:
- The prompt is sent to Claude along with the tool schemas
- Claude decides whether to call GetCakeRecipe or answer directly
- If it calls the tool, the MCP client sends the request to the Gateway, the Gateway invokes the tool backend, and the result comes back
- Claude sees the tool result and formulates its final answer
- The result object contains the final text
Dockerfile
FROM python:3.12-slim
WORKDIR /app
RUN apt-get update && apt-get install -y --no-install-recommends gcc g++ \
&& rm -rf /var/lib/apt/lists/*
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY app.py .
EXPOSE 8080
CMD ["python", "app.py"]
gcc/g++ are needed because some Python packages (cryptography, certain boto3 dependencies) compile C extensions. The image must be built for ARM64 — AgentCore doesn’t support x86.
The Terraform
IAM: Three Roles, Three Boundaries
# Runtime role — AgentCore runs your container with this identity
resource "aws_iam_role" "runtime" {
assume_role_policy = jsonencode({
Statement = [{ Principal = { Service = "bedrock-agentcore.amazonaws.com" }, ... }]
})
}
# Gateway role — AgentCore uses this to invoke your tool backends
resource "aws_iam_role" "gateway" {
assume_role_policy = jsonencode({
Statement = [{ Principal = { Service = "bedrock-agentcore.amazonaws.com" }, ... }]
})
}
# Lambda role — execution role for our tool function (this example uses Lambda)
resource "aws_iam_role" "lambda" {
assume_role_policy = jsonencode({
Statement = [{ Principal = { Service = "lambda.amazonaws.com" }, ... }]
})
}
The Runtime role needs these permissions:
- bedrock:InvokeModel — call Claude
- bedrock-agentcore:InvokeGateway — call the Gateway (THE ONE EVERYONE FORGETS)
- ecr:BatchGetImage — pull the container image
- logs:PutLogEvents — write CloudWatch logs
The Gateway role needs:
- Permissions to invoke whatever backend your targets use (in this example: lambda:InvokeFunction)
The Gateway and Tool Binding
resource "aws_bedrockagentcore_gateway" "main" {
name = "myagent-gateway"
role_arn = aws_iam_role.gateway.arn
authorizer_type = "AWS_IAM"
protocol_type = "MCP"
}
resource "aws_lambda_alias" "get_cake_recipe" {
name = "GetCakeRecipe"
function_name = aws_lambda_function.tool.function_name
function_version = "$LATEST"
}
resource "aws_bedrockagentcore_gateway_target" "get_cake_recipe" {
gateway_identifier = aws_bedrockagentcore_gateway.main.gateway_id
target_configuration {
mcp {
lambda {
lambda_arn = aws_lambda_alias.get_cake_recipe.arn
tool_schema {
inline_payload {
name = "GetCakeRecipe"
description = "Returns a cake recipe. Call this when the user asks about baking."
input_schema { type = "object" }
}
}
}
}
}
}
The description in tool_schema is what the LLM sees. Write it clearly — this is how Claude decides whether to call your tool. A vague description means the tool never gets used; an overly broad one means it gets called when it shouldn’t be.
The Runtime
resource "aws_bedrockagentcore_agent_runtime" "main" {
agent_runtime_name = "myagent_runtime"
agent_runtime_artifact {
container_configuration {
container_uri = "${aws_ecr_repository.agent.repository_url}:latest"
}
}
role_arn = aws_iam_role.runtime.arn
protocol_configuration { server_protocol = "HTTP" }
network_configuration { network_mode = "PUBLIC" }
environment_variables = {
MODEL_ID = "eu.anthropic.claude-sonnet-4-5-20250929-v1:0"
AWS_REGION = "eu-central-1"
GATEWAY_URL = aws_bedrockagentcore_gateway.main.gateway_url
}
}
GATEWAY_URL is computed from the Gateway resource — Terraform wires them together automatically.
Deploy and Test
terraform init && terraform apply
The Terraform includes a null_resource that builds the ARM64 Docker image and pushes it to ECR before creating the Runtime. One command, no manual steps.
After apply (2–3 minutes for the container to start):
aws bedrock-agentcore invoke-agent-runtime \
--agent-runtime-arn "$(terraform output -raw runtime_arn)" \
--payload '{"prompt": "Give me a cake recipe"}' \
--content-type application/json \
/dev/stdout
Things That Will Bite You
- Missing InvokeGateway permission. Agent works but never calls tools. No error in logs — the MCP client gets a silent 403. Add bedrock-agentcore:InvokeGateway to the Runtime role.
- server_protocol = “MCP” on the Runtime. Your container speaks HTTP (BedrockAgentCoreApp is an HTTP server). MCP is only the protocol between Runtime → Gateway. Set “HTTP”.
- x86 container image. AgentCore only runs ARM64. You need –platform linux/arm64 and QEMU/buildx for cross-compilation on Intel/AMD machines.
- Image not refreshing after push. AgentCore caches the pulled image. Change any environment variable to force a restart with the fresh image.
- VPC mode with no outbound route. network_mode = “VPC” requires the subnets to have internet access (NAT or VPC endpoints). Use “PUBLIC” unless you specifically need private network access.
What You Can Build with This
The cake recipe is one tool returning static data. But the same architecture supports agents that reason across multiple tool calls to answer complex questions.
Consider a billing data assistant where a user asks: “Which of my customers have unpaid invoices from last quarter, and what’s their total outstanding amount?”
No single database query answers that. The agent breaks it down:
- First it calls a tool that lists unpaid invoices filtered by date range
- Then it calls a tool that aggregates amounts grouped by customer
- It combines both results and presents a ranked list with totals
The LLM decides this sequence on its own — you don’t hardcode the chain. You just provide the tools (each one a parameterized query) and a system prompt that describes when to use which. The agent figures out that answering the user’s question requires two steps, picks the right tool for each, passes the right parameters (date range, paid status filter), and synthesizes the results into one coherent answer.
Scale that to 10 or 15 tools covering different aspects of a database — customers, invoices, payments, revenue aggregations, full-text search — and the agent can answer almost any analytical question a business user throws at it.
That’s the jump from “one tool, one response” to a real production agent: the LLM becomes an autonomous query planner that chains multiple data retrievals together, each one scoped and parameterized, to answer questions that no single query could handle alone.