What it actually takes to put AI agents into production

What stands between a pilot and production is the system around the agent. Invest in it once, and every agent after the first ships in weeks.

Your best engineers can now vibe-code a working agent in a week or two, and the demo goes well. Getting that agent into production, acting for real people on real data, takes months, and many never get there: Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, citing rising costs, unclear business value, and inadequate risk controls. A weak model is not on that list. When a pilot stalls, the agent can usually do the work; what is missing are answers to the questions a company asks of anything that acts on its behalf: who it works for, what it may spend, how to stop it, and how to show what it did.

Those answers are the same for every agent. Invest in them once, as shared infrastructure, and each new agent reuses that investment and can reach production in weeks. Leave each team to work them out for its own agent, and you pay for the same security review and the same budget debate again with every pilot, which is why so many never finish.

Invest in the platform once, instead of paying for it again in every pilot.

› The enterprise agent stack

This is that system. Every agent a company runs in production sits on some version of it, whether anyone has drawn it or not. Most of it you already own: model providers, your identity provider, your data, your infrastructure. The eight numbered layers are the part most pilots skip, and they decide whether an agent can be trusted with real work.

The eight layers inside the full enterprise agent stack Channels flow through an identity edge into the agent runtime, where your agent sits, between a control plane and guardrails. The runtime thinks through a model gateway, acts through tools with per-user data access, and grounds in knowledge and memory; the gateway calls model providers. Evidence and evals, security and governance, and build and release cut across every layer, on top of infrastructure. Below: the life of one request, and where a CTO should spend attention. 1 CHANNELS Web · IDE · Slack / Teams · Email · API · schedule · events message + who is asking 2 EDGE & IDENTITY API gateway + WAF · SSO (OIDC / SAML: Okta, Entra ID) · tenant routing · rate limits every request stamped: user_id · tenant_id · agent_id · session_id · trace_id authenticated turn 5 CONTROL PLANE lease ⇄ runtime kill in under 1 min deny user · caps pin version alerts → on-call 3 AGENT RUNTIME orchestrator · durable workflows human-in-the-loop approvals session state · retries · sandbox no keys, no DB creds on the host YOUR AGENT HERE Agent Packs, LangChain, CrewAI GUARDRAILS inline on every hop prompt injection PII / DLP redaction policy (OPA, Cedar) tool allow-lists output schema check think act ground 4 MODEL GATEWAY LLM proxy (e.g. LiteLLM) routing & fallback prompt / semantic cache budget checked pre-call over cap → refused cost attributed per user provider keys held here 6 TOOLS & DATA ACCESS MCP servers internal APIs · SaaS CRM, ITSM, GitHub code sandbox (microVM) browser automation acts AS the user (OBO) wrong user → 0 rows KNOWLEDGE & MEMORY ingestion + chunking ETL vector DB + keyword search ACL-aware retrieval session / long-term memory warehouse, systems of rec. SharePoint, Drive, wikis your keys, or our credits MODEL PROVIDERS Frontier APIs Cloud-hosted Self-hosted open weights Anthropic, OpenAI, AWS Bedrock, GCP Vertex, vLLM / TGI on a GPU pool, Google Azure AI Foundry fine-tuned small models CROSS-CUTTING · EVERY LAYER 7 EVIDENCE & EVALS OpenTelemetry traces cost · tokens · latency per user/agent/version tamper-evident (WORM) eval harness, LLM judges regression/drift alerts trace replay for debug SECURITY & GOVERNANCE agent workload identity short-lived scoped creds secrets vault (KMS/HSM) data residency · DLP SOC 2 · ISO 42001 EU AI Act mapping 8 BUILD & RELEASE IDE / CLI · templates signed, versioned packs prompt + tool versioning offline eval suites eval-gated CI/CD canary + one-cmd rollback agent registry / catalog INFRASTRUCTURE Kubernetes / VMs / Cloud Run · VPC + private link · Postgres · queues (Kafka / SQS) object storage · GPU capacity · IaC (Terraform) · multi-region DR · FinOps tagging LIFE OF ONE REQUEST 1. Alice asks in Slack [1] → the edge [2] checks SSO, stamps user + trace id 2. The runtime [3] holds a live lease from the control plane [5]; guardrails screen input 3. Agent loop: think via [4] → act as Alice via [6] → ground in knowledge (N steps) 4. Risky action (payment, prod deploy, customer email) → human approval in [3] 5. Guardrails check the output → reply; evidence [7] holds trace, version, cost, audit WHERE A CTO SHOULD SPEND ATTENTION Buy / adopt model providers, IdP, vector DB, Kubernetes, OTEL backend Standardize one gateway [4], one tool protocol (MCP) [6], one trace schema [7] Build (your edge) agents + prompts, integrations to YOUR systems, eval datasets Non-negotiable per-user identity [2], kill switch [5], spend caps [4], audit trail [7] Top risks prompt injection via tools/docs, over-privileged agents, runaway cost, silent quality regressions after a model or prompt change agenthippo.ai
        1  CHANNELS   Web · IDE · Slack / Teams · Email · API · schedule · events
                                            │ message + who is asking
                                            ▼
┌─ 2  EDGE & IDENTITY ───────────────────────────────────────────────────────────────────┐
│ API gateway + WAF · SSO (OIDC / SAML: Okta, Entra ID) · tenant routing · rate limits   │
│ every request stamped: user_id · tenant_id · agent_id · session_id · trace_id          │
└───────────────────────────────────────────┬────────────────────────────────────────────┘
                                            │ authenticated turn
                                            ▼
┌─ 5  CONTROL PLANE ───┐  ┌─ 3  AGENT RUNTIME ─────────────────┐  ┌─ GUARDRAILS ─────────┐
│ lease ⇄ runtime      │  │ ┌─ YOUR AGENT HERE ──────────────┐ │  │ inline on every hop  │
│ kill in under 1 min  │◄►│ │ Agent Packs, LangChain, CrewAI │ │◄►│ prompt injection     │
│ deny user · caps     │  │ └────────────────────────────────┘ │  │ PII / DLP redaction  │
│ pin version          │  │ orchestrator · durable workflows   │  │ policy (OPA, Cedar)  │
│ alerts → on-call     │  │ human-in-the-loop approvals        │  │ tool allow-lists     │
│                      │  │ session state · retries · sandbox  │  │ output schema check  │
│                      │  │ no keys, no DB creds on the host   │  │                      │
└──────────────────────┘  └─────────────────┬──────────────────┘  └──────────────────────┘
             ┌──────────────────────────────┼─────────────────────────────┐
             │ think                        │ act                         │ ground
             ▼                              ▼                             ▼
┌─ 4  MODEL GATEWAY ───────┐  ┌─ 6  TOOLS & DATA ACCESS ─┐  ┌─ KNOWLEDGE & MEMORY ───────┐
│ LLM proxy (e.g. LiteLLM) │  │ MCP servers              │  │ ingestion + chunking ETL   │
│ routing & fallback       │  │ internal APIs · SaaS     │  │ vector DB + keyword search │
│ prompt / semantic cache  │  │ CRM, ITSM, GitHub        │  │ ACL-aware retrieval        │
│ budget checked pre-call  │  │ code sandbox (microVM)   │  │ session / long-term memory │
│ over cap → refused       │  │ browser automation       │  │ warehouse, systems of rec. │
│ cost attributed per user │  │ acts AS the user (OBO)   │  │ SharePoint, Drive, wikis   │
│ provider keys held here  │  │ wrong user → 0 rows      │  │                            │
└────────────┬─────────────┘  └──────────────────────────┘  └────────────────────────────┘
             │ your keys, or our credits
             ▼
┌─ MODEL PROVIDERS ──────────────────────────────────────────────────────────────────────┐
│ Frontier APIs              Cloud-hosted                 Self-hosted open weights       │
│ Anthropic, OpenAI,         AWS Bedrock, GCP Vertex,     vLLM / TGI on a GPU pool,      │
│ Google                     Azure AI Foundry             fine-tuned small models        │
└────────────────────────────────────────────────────────────────────────────────────────┘

══════════════════════════════ CROSS-CUTTING · every layer ═══════════════════════════════
┌─ 7  EVIDENCE & EVALS ────┐  ┌─ SECURITY & GOVERNANCE ──┐  ┌─ 8  BUILD & RELEASE ───────┐
│ OpenTelemetry traces     │  │ agent workload identity  │  │ IDE / CLI · templates      │
│ cost · tokens · latency  │  │ short-lived scoped creds │  │ signed, versioned packs    │
│ per user/agent/version   │  │ secrets vault (KMS/HSM)  │  │ prompt + tool versioning   │
│ tamper-evident (WORM)    │  │ data residency · DLP     │  │ offline eval suites        │
│ eval harness, LLM judges │  │ SOC 2 · ISO 42001        │  │ eval-gated CI/CD           │
│ regression/drift alerts  │  │ EU AI Act mapping        │  │ canary + one-cmd rollback  │
│ trace replay for debug   │  │                          │  │ agent registry / catalog   │
└──────────────────────────┘  └──────────────────────────┘  └────────────────────────────┘
┌─ INFRASTRUCTURE ───────────────────────────────────────────────────────────────────────┐
│ Kubernetes / VMs / Cloud Run · VPC + private link · Postgres · queues (Kafka / SQS)    │
│ object storage · GPU capacity · IaC (Terraform) · multi-region DR · FinOps tagging     │
└────────────────────────────────────────────────────────────────────────────────────────┘

LIFE OF ONE REQUEST
 1. Alice asks in Slack [1] → the edge [2] checks SSO, stamps user + trace id
 2. The runtime [3] holds a live lease from the control plane [5]; guardrails screen input
 3. Agent loop: think via [4] → act as Alice via [6] → ground in knowledge     (N steps)
 4. Risky action (payment, prod deploy, customer email) → human approval in [3]
 5. Guardrails check the output → reply; evidence [7] holds trace, version, cost, audit

WHERE A CTO SHOULD SPEND ATTENTION
 Buy / adopt       model providers, IdP, vector DB, Kubernetes, OTEL backend
 Standardize       one gateway [4], one tool protocol (MCP) [6], one trace schema [7]
 Build (your edge) agents + prompts, integrations to YOUR systems, eval datasets
 Non-negotiable    per-user identity [2], kill switch [5], spend caps [4], audit trail [7]
 Top risks         prompt injection via tools/docs, over-privileged agents, runaway cost,
                   silent quality regressions after a model or prompt change

Two things in the picture matter more than any single box. The agent sits in the middle holding nothing worth stealing: no provider keys, no database passwords, no standing permissions. And every path out of it runs through a layer that can say no. The other boxes are implementation choices your engineers can make.

› Where pilots stall

The numbered layers are easiest to explain through the problems a pilot runs into on its way to production, roughly in the order they appear. Each paragraph below covers one problem and how to fix it; if you would rather see the solution first, the next section shows the full stack as we deploy it.

Reach. The first agent usually lives in one engineer’s terminal, and six months later it still has one user. Agents create value where work already happens, in Slack, Teams, email, or a job that runs overnight, used by people who never saw the demo. Once people other than the builder use it, the agent has to know who each of them is, and it has to get that from your identity provider rather than from what they type. Otherwise anyone can claim to be the CFO in the chat and receive the CFO’s answers.

Data. The prototype connected to the data warehouse with a single service account, because that was quickest. In production, that shortcut means a manager asks about their team’s spend and gets a colleague’s salary in the answer, because to the database every question comes from the same user. The fix is decades old: the agent queries as the person who asked, and the database applies the permissions it already has. Of the problems on this list, a data leak is the one most likely to become a regulatory matter, and it is far cheaper to prevent before launch than to clean up after.

Cost. Agents retry, loop, and fan out. One that hits an error it cannot parse on a Friday evening can spend the month’s budget before anyone opens a dashboard on Monday, and the invoice arrives as a single number with nobody’s name on it. The fix is a model gateway: one service that sits between every agent and every model provider. It checks each call against the budget of that agent and that user before the call is made, blocks it if the budget is spent, and attributes every dollar to a person and a use case. That same cost data then tells you which agents are worth what they cost.

Model choice. The pilot was wired to one provider’s model, with an API key pasted into its configuration. Once every call goes through the gateway, the model becomes a setting you manage centrally. Routine work can go to a cheaper cloud model or an open model on your own hardware, requests involving sensitive data can stay on a local model inside your network, and only the hardest tasks need the most expensive frontier model. You can try a new model on a share of real traffic and compare cost and quality before switching, without changing the agent. The gateway also keeps the provider keys, so no team or agent ever holds one, and tracks each user’s token usage against their allowance.

Control. Something will go wrong at an inconvenient hour. What matters is whether stopping it takes one click from a console, or an hour of working out which machine it runs on while the engineer who deployed it is on a plane. Agents also take instructions from whatever they read, including a line hidden in a support ticket or a vendor’s PDF. So the agent should hold no keys or passwords, and its limits should be enforced by the systems around it rather than written into its instructions, where the next document it reads could override them.

Accountability. A customer forwards an email your agent sent and asks what it was based on. You need an answer in minutes: who asked, which version of the agent ran, what data it touched, and what it cost. Ordinary logs, written by the same machine that ran the agent, are fine for debugging. A customer or an auditor needs a record stored outside the agent’s reach that cannot be altered after the fact.

Change. The agent got worse this week. Someone edited the prompt, swapped a tool, or the provider updated the model underneath it, and nobody can say which, because none of it was versioned together. Treat the prompt, tools, model, and engine as one release, pin every deployment to a version, and a bad change becomes a one-command rollback instead of a week of guesswork. Because mistakes are cheap to undo, teams can release improvements more often.

None of this is new. Companies already run people and software this way: a login tied to a name, a spending limit, a manager who can say stop, a paper trail, a change process. What is new is that for agents these controls cannot live only in a policy document. A policy can say nobody may wire $50,000, but it is the payments system that actually blocks the transfer. Agents need that kind of enforcement built into the systems they use, because an agent will not reliably follow written rules. That is what we mean by runtime governance: controls enforced by software on every action an agent takes.

› The stack in practice

Here is the same stack as we deploy it at AgentHippo, with the components we use by default. Two refusals carry the design: a call over budget never reaches a model, and a request for another user’s data comes back empty. Everything inside the dashed boundary runs in your environment.

How a governed agent deployment is wired A user's message passes through a TLS proxy, oauth2-proxy and your identity provider to the auth-broker, which hands the agent runtime a verified user and a short-lived token. Your agent runs inside the runtime as a signed Agent Pack. The runtime holds a lease from the control plane and receives signed packs from build and release. It makes metered calls through agenthippo-gateway, which refuses calls over the budget cap, reads data as the user through an on-behalf-of connector that returns no rows to the wrong user, and records every action in the evidence archive. Runtime, gateway, data access and evidence run inside your environment. YOUR ENVIRONMENT 1·2 EDGE & IDENTITY TLS proxy → oauth2-proxy ⇄ your IdP (Okta · Entra · Google) → auth-broker out: the verified user + a short-lived token for this turn 5 CONTROL PLANE agenthippo-control lease (TTL) ⇄ runtime kill · deny user · caps pin version · console alerts → PagerDuty / IRM 3 AGENT RUNTIME YOUR AGENT HERE signed Agent Pack agent.yaml · prompt · skills 4 GATEWAY agenthippo-gateway identity + budget → LiteLLM → model over cap → refused 6 DATA (OBO) connector container Postgres RLS DynamoDB (STS tags) Unity Catalog wrong user → 0 rows 7 EVIDENCE OTEL → evidence-server MinIO · S3 Object Lock WORM · Spotlight who·version·touched·cost 8 BUILD & RELEASE IDE / CLI → Agent Pack sign → GHCR image · store tag one-command rollback Alice · Bob via Slack · WhatsApp · Telegram · API · schedule serve · sandbox on your VM · Cloud Run · Databricks Apps · behind your proxy no provider key, no DB credential a lease it keeps renewing message + who is asking authenticated turn metered call as the user every action signed pack agenthippo.ai
                                Alice · Bob   via Slack · WhatsApp · Telegram · API · schedule
                                                              │ message + who is asking
                                                              ▼
┌─ 1·2  EDGE & IDENTITY ─────────────────────────────────────────────────────────────────────────────────┐
│ TLS proxy → oauth2-proxy ⇄ your IdP (Okta · Entra · Google) → auth-broker                              │
│ out: the verified user + a short-lived token for this turn                                             │
└─────────────────────────────────────────────────────────────┬──────────────────────────────────────────┘
                                                              │ authenticated turn
                                                              ▼
┌─ 5  CONTROL PLANE ────────┐   ┌─ 3  AGENT RUNTIME ─────────────────────────────────────────────────────┐
│ agenthippo-control        │   │ ┌─ YOUR AGENT HERE ────────────┐ serve · sandbox                       │
│ lease (TTL) ⇄ runtime     │◄──│ │ signed Agent Pack            │ on your VM · Cloud Run ·              │
│ kill · deny user · caps   │──►│ │ agent.yaml · prompt · skills │ Databricks Apps · behind your proxy   │
│ pin version · console     │   │ └──────────────────────────────┘ no provider key, no DB credential     │
│ alerts → PagerDuty / IRM  │   │                                  a lease it keeps renewing             │
└───────────────────────────┘   └──────────┬──────────────────────┬─────────────────────────┬────────────┘
              ▲ signed pack                │ metered call         │ as the user             │ every action
┌─ 8  BUILD & RELEASE ──────┐              ▼                      ▼                         ▼
│ IDE / CLI → Agent Pack    │   ┌─ 4  GATEWAY ───────┐ ┌─ 6  DATA (OBO) ─────┐ ┌─ 7  EVIDENCE ───────────┐
│ sign → GHCR               │   │ agenthippo-gateway │ │ connector container │ │ OTEL → evidence-server  │
│ image · store tag         │   │ identity + budget  │ │ Postgres RLS        │ │ MinIO · S3 Object Lock  │
│ one-command rollback      │   │ → LiteLLM → model  │ │ DynamoDB (STS tags) │ │ WORM · Spotlight        │
└───────────────────────────┘   │ over cap → refused │ │ Unity Catalog       │ │ who·version·touched·cost│
                                └────────────────────┘ │ wrong user → 0 rows │ └─────────────────────────┘
                                                       └─────────────────────┘

Nothing in this picture depends on the model behaving well. If a prompt is injected, the runtime has no secrets to leak. If the agent loops, the gateway stops paying for it. If it asks for another customer's data, the database returns nothing. If it has to stop, it stops, because it must keep checking in with the control plane to go on working. That property is worth copying whatever components you choose.

Components are defaults, not requirements: bring your own identity proxy, object storage, or provider keys. The control plane runs hosted by AgentHippo or on your own infrastructure.

› Where to spend, and where not to

Most of this stack should not be built by you. Model providers, identity, search, Kubernetes, observability backends: buy or adopt them, and resist building a custom version of anything on that list. The same goes for guardrails that screen prompts and outputs. They are worth having because they reduce how often the model makes mistakes, but they cannot limit the damage when it makes one anyway. That is the job of the layers above.

A few things are worth standardizing across the company early, before every team picks its own: one model gateway, one way for agents to reach tools, one format for the record of what they did. Standardize those, and a new agent plugs into what exists instead of rebuilding it.

Where your engineers’ time does pay off is in what no vendor can give you: the agents themselves, the integrations into your own systems, and the test cases that describe what good work looks like in your business. That last one is the most underrated asset in AI right now. Models will keep changing underneath you. With a good set of test cases, you can check in a week whether a new model does your work better and switch with confidence; without them, every model change is a judgment call.

› Where we fit

A disclosure, since you are reading this on our website: this stack is what we build. AgentHippo is the numbered layers, packaged so a team can stand them up in its own environment with any model and any agent framework, on its own provider keys or our credits. Every layer deploys from a plain script your engineers can read in full before it runs, and every deploy checks itself before it reports success. That said, you do not need AgentHippo to apply anything in this post: the same design can be built from other components.

› Where to start

You do not need all eight layers on day one. Four carry most of the risk, and they fit in one sentence: the agent acts for a verified person, within a budget, under an off switch, and on the record. Add per-user data access the day it touches a system of record, and release discipline the day a second person edits the prompt.

Then pick one workflow where a mistake would be visible and reversible, and where you can measure the outcome in money or hours. Run it behind those four controls for a month. Widen the agent’s authority only as fast as the evidence supports, and let what you learn decide what the second agent should be. By the third agent, the leadership discussion usually shifts from whether agents work to which workflow to automate next.

› One workflow in production, in three weeks

If you would rather not assemble this yourself, our Production Sprint puts one real workflow into production on this stack, in your environment, in three weeks, with the numbers to show whether it paid off. If your team would like to evaluate on its own first, it can start free.