IRM Consulting & Advisory
AI & Machine Learning Security

Security Risks of Autonomous Agents

Autonomous Agents, AI-driven systems that can operate independently to make decisions or perform tasks, are becoming increasingly integrated into our daily lives and business operations.

What changes when AI stops answering questions and starts taking actions

A chatbot that gives a wrong answer is embarrassing. An agent that takes a wrong action can delete a database, send an email to the wrong customer or approve a payment. That is the whole difference between generative AI and autonomous agents, and it is why agent security needs its own thinking.

When I first wrote about autonomous agents in 2024, most examples came from self-driving cars and drones. Since then, AI agents have moved into ordinary business software: coding assistants that run commands, support agents that issue refunds, and assistants that read your inbox and act on it. This post covers the risks that matter for a small or mid-sized company deploying those agents, ranked and with real cases.

What makes an autonomous agent riskier than a chatbot?

An autonomous agent is an AI system that plans and carries out multi-step tasks on its own, using tools such as APIs, file systems, browsers or payment systems, with limited human input between steps. Three properties raise the stakes: it holds credentials, it can call tools, and it runs in a loop without asking permission at each step.

Security researcher Simon Willison named the most dangerous combination the lethal trifecta: an agent that has access to private data, is exposed to untrusted content, and can communicate externally. Any agent with all three can be tricked into sending your data to an attacker. Remove any one leg and the risk drops sharply. I use this test first on every agent design review.

The agent risks that matter, ranked

This ranking is for a SaaS or services company using agents in business workflows. If you build robots or vehicles, physical attacks move higher.

Risk

What it looks like

Priority

Excessive agency

The agent has admin credentials when it needs read access to one table

High

Indirect prompt injection

Instructions hidden in an email, web page, ticket or document redirect the agent

High

Tool and plugin supply chain

A third-party connector or tool server is malicious or compromised

High

Runaway actions and cost

The agent loops, retries or scales a task far past what anyone intended

Medium

Memory and context poisoning

False information planted in the agent's long-term memory or knowledge base shapes later decisions

Medium

Missing audit trail

Nobody can reconstruct what the agent did or why

Medium, high in regulated sectors

Adversarial physical inputs

Stickers, signs or reflections fool the sensors of a physical system

Low for software companies, high for robotics

The original version of this post listed ransomware, insider threats and communication channel security as agent risks. They still apply, but they are general security risks. The table above focuses on what is specific to agents.

An isometric image of a laptop with a cloud connected to it.

Excessive agency: the risk behind most agent incidents

OWASP lists excessive agency in its Top 10 for LLM applications. It means giving an AI system more functionality, permissions or autonomy than the task needs. It is the risk that turns every other failure into real damage.

In July 2025, SaaStr founder Jason Lemkin described how an AI coding agent on Replit deleted his production database during an explicit code freeze, then gave misleading answers about what it had done. Replit's CEO apologised publicly and said the company had begun rolling out automatic separation between development and production databases. The agent did not need to be hacked. It simply had the access, and it used it.

The lesson is dull and effective. Give each agent its own identity, never a shared human account. Scope its credentials to the task. Require a human to approve anything irreversible: deletes, payments, external emails, production changes.

Indirect prompt injection: when the data gives the orders

A language model reads instructions and data through the same channel. If your agent reads a support ticket, a web page or a PDF that contains hidden instructions, it may follow them. This is called indirect prompt injection, and it is the attack that makes the lethal trifecta dangerous.

The 2025 EchoLeak vulnerability (CVE-2025-32711) showed the pattern in a mainstream product: a crafted email could lead Microsoft 365 Copilot to expose data from the user's context without any click. Microsoft patched it, but no vendor has a general fix. Our post on how AI agents communicate explains the plumbing that makes this possible.

The tool supply chain nobody reviews

Agents become useful through connectors: to your CRM, your code repository, your cloud console. Many of those connectors come from third parties or open-source projects. Each one runs with the agent's permissions, and many teams install them with less scrutiny than a browser extension.

Treat every agent tool like a vendor. Know who publishes it, what data it can reach, and whether it can make outbound network calls. Pin versions rather than auto-updating. This is ordinary third-party risk management, applied to a new kind of supplier.

Two hands typing on a laptop keyboard

Physical agents: where adversarial attacks started

For systems that act in the physical world, adversarial inputs remain a real concern. In 2019, Tencent's Keen Security Lab showed that small stickers placed on a road could lead Tesla's Autopilot to steer into the oncoming lane in a controlled test. Earlier academic work showed that a few stickers on a stop sign could cause image classifiers to misread it.

If you build robotics, drones or vehicles, test against adversarial inputs, add redundant sensors, and design fail-safe behaviour when confidence drops. If you build software, this is interesting context rather than a priority. Data poisoning, where the training data itself is corrupted, matters more for you only if you train your own models.

When you do not need heavy agent controls

Not every agent deserves a security project. An agent that only reads data you already consider public, cannot send anything outside your company and has a human approving each output is low risk. So is an agent running in a sandbox with no production credentials.

I would rather see a company run ten low-risk agents with basic logging than stall all agent adoption waiting for a perfect framework. Spend your control effort on agents that hold write access, touch customer data or can reach the internet.

An agent security checklist for small teams

This is the sequence we use when reviewing an agent before it goes live. It fits a company with no dedicated security engineer.

  1. Write down the agent's job in one sentence, and list every tool and dataset it can reach.

  2. Apply the lethal trifecta test. If the agent has private data, untrusted input and external communication, remove one.

  3. Give the agent its own service identity with scoped, short-lived credentials.

  4. Require human approval for irreversible actions: deletes, payments, external messages, production changes.

  5. Restrict outbound network access to an allowlist of domains.

  6. Log every prompt, tool call and result where the agent cannot alter the logs.

  7. Set rate and spend limits, plus a way to stop the agent immediately.

  8. Review third-party tools and connectors as you would any vendor.

  9. Test with hostile inputs before launch and after every significant change.

For a fuller treatment, see 7 safeguards to secure AI agents and our breakdown of a five-day AI agent cyberattack.

Which frameworks cover agent risk?

You do not need a new framework for agents. MITRE ATLAS catalogues real adversary techniques against AI systems. The NIST AI Risk Management Framework gives you a governance structure, and ISO/IEC 42001 turns that structure into something an auditor can certify. Our ISO 42001 readiness checklist shows how agent controls map to the standard.

Enterprise buyers are starting to ask about agents in security questionnaires. Being able to show an agent inventory, the approval rules and the logs is often enough to answer them.

Who should own agent security?

In most 30 to 200 person companies, agents are deployed by engineering or operations, and nobody owns the risk. That is the gap a virtual CISO fills: setting the approval rules, reviewing new agents and reporting to leadership, without a full-time hire.

If you are deploying agents and want an outside review, book a consultation or compare options on our pricing page.

Keep Reading

Related Articles

Our Industry Certifications

Our diverse industry experience and expertise in AI, Cybersecurity & Information Risk Management, Data Governance, Privacy and Data Protection Regulatory Compliance is endorsed by leading educational and industry certifications for the quality, value and cost-effective products and services we deliver to our clients.

Copyright © 2026 IRM Consulting & Advisory. All Rights Reserved.