Guide
AI Agent Security Checklist
Twelve concrete controls to put around any AI agent before it touches your data, your customers, or your money.
Why agents need their own security thinking
An AI agent is software that takes a goal, plans steps, and executes them with tools. That makes it different from a chatbot in one crucial way: a chatbot can only say the wrong thing, while an agent can do the wrong thing — send the email, delete the record, move the money. Agents amplify both productivity and mistakes, which is why they deserve security controls designed around what they can do, not just what they can say. If you're new to agents, our SI agent explainer covers the concept.
The 12-point security checklist
- Apply least-privilege access. Give the agent the minimum permissions its task needs. Read-only where possible; write access only where the task requires it.
- Require approval for irreversible actions. Anything that messages customers, spends money, deletes data, or changes production systems pauses for a human. This is the single highest-value control on this list.
- Isolate the agent's credentials. Use dedicated API keys and service accounts for each agent — never a personal admin account. Revoke and rotate them like any other secret.
- Treat all external content as untrusted. Emails, web pages, files, and ticket bodies can carry prompt injection: hidden instructions that steer the agent off-task. Don't let an agent act on untrusted content without a human reviewing the plan.
- Segment what the agent can see. Don't connect the agent to every system at once. Scope data access to the task, and separate agents by function so a compromise of one doesn't reach everything.
- Log every action. The agent should produce an audit trail: what it read, what it decided, what it changed. Review the logs weekly at first, monthly once stable.
- Set spend and usage limits. Cap API spend, execution counts, and credits so a runaway loop or misconfigured trigger can't generate a surprise bill or a flood of actions.
- Redact sensitive data from prompts. Don't paste secrets, full customer databases, or credentials into agent instructions or test chats. Assume anything in the prompt could be logged.
- Review third-party integrations. Every connector is a new trust relationship. Check what permissions each integration requests, and remove ones you no longer use.
- Test with adversarial inputs. Before going live, try to break the agent: contradictory instructions, fake "urgent" emails, requests beyond its scope. Fix what you find in the instructions and guardrails.
- Have a kill switch. Know exactly how to stop the agent instantly — disable the schedule, revoke the key, pause the workflow — and make sure more than one person knows how.
- Plan for the vendor's security, too. If you use a commercial platform, review its certifications (SOC 2, GDPR, HIPAA where relevant), data retention policy, and subprocessors. See the vendor questions below.
Security questions to ask any vendor
- Where is my data stored, how long is it retained, and is it used to train models?
- What certifications do you hold (SOC 2, ISO 27001), and can I see the reports?
- How are customer credentials stored, and who at your company can access my data?
- Do you support SSO, role-based access, and audit logs on my plan — not just enterprise?
- What happens to my data if I cancel?
- Do you offer a HIPAA-compliant tier with a signed BAA, if my industry needs it?
Our best AI agents roundup notes which platforms publish enterprise-grade security features, and our open-source vs commercial comparison covers the control trade-offs of self-hosting.
Start strict, loosen deliberately
The pattern that works: deploy with maximum controls — approvals on, narrow access, tight budgets — and relax them one at a time as the agent proves itself. An agent that earns trust over weeks is an asset; one trusted on day one is a liability. Security for agents isn't a one-time setup; it's a habit of reviewing logs, permissions, and instructions as the agent's job evolves. Read how we rank agents to see how security factors into our evaluations.
Frequently asked questions
What is the biggest security risk with AI agents?
Excessive permissions. An agent that can read and write everything in your systems turns any mistake or manipulation — including prompt injection from untrusted content — into a much bigger incident than it needs to be.
What is prompt injection and why does it matter for agents?
Prompt injection is when malicious instructions hidden in data an agent reads — an email, a webpage, a file — trick it into acting against your intent. It matters because agents follow instructions from their context, and they can't always tell your instructions apart from an attacker's.
Should AI agents ever act without human approval?
Only for low-stakes, reversible actions like drafting content or looking up data. Anything that sends messages to customers, spends money, changes records, or touches production systems should require human approval, especially in the first months.
Is it safe to connect an AI agent to my email and CRM?
It can be, if you scope access narrowly, use read-only access where possible, require approvals for writes, and audit activity regularly. Connect the tools the task needs — nothing more.