Skip to content

Papers

Securing Agentic AI: Addressing Key Vulnerabilities in Autonomous Agents

This paper explores the multifaceted security challenges in agentic AI, focusing on prompt injection, memory poisoning, tool integrity, inter-agent communication, and model routing. The authors highlight the need for a holistic approach to ensure the integrity and provenance of autonomous agents, particularly in critical domains like supply chain management.

Original source: arXiv

Securing Agentic AI: Addressing Key Vulnerabilities in Autonomous Agents

Paper: Securing Agentic AI: From Per-Action Checks to Trajectory Assurance

Authors: Alireza Lotfi, Subangkar Karmaker Shanto et al.

Published: 2026

Venue: arXiv preprint

Source: https://arxiv.org/abs/2608.01558

Research Background

As autonomous agents powered by large language models (LLMs) become more prevalent in critical domains such as healthcare, finance, and telecommunications, the need for robust security measures becomes paramount. These agents are not only executing individual tasks but also interacting with other agents and external systems, making their security a complex and interconnected challenge.

The problem of securing agentic AI is multifaceted. Traditional security approaches that focus on per-action checks are no longer sufficient. Instead, the overall behavior of an agent must be consistent with the rules and invariants of the system it operates in. This is particularly important in the supply chain industry, where the integrity and reliability of these agents can have far-reaching consequences.

In the context of the supply chain, autonomous agents are increasingly used to manage inventory, optimize logistics, and even make strategic decisions. The safety and reliability of these agents are crucial for maintaining the integrity of the supply chain. However, prior approaches to securing these agents have several shortcomings. For instance, they often fail to address the cumulative effects of individually permissible actions, which can collectively violate system-level constraints. Additionally, they do not adequately account for the risks associated with inter-agent communication and delegation, or the vulnerabilities in the underlying model routing and execution control plane.

Industry Context and Shortcomings

The supply chain industry relies heavily on AI-driven decision-making to enhance efficiency and reduce costs. Autonomous agents are used to automate various processes, from demand forecasting to route optimization. However, the security of these agents is a significant concern. Prior approaches, such as per-action checks and static policy enforcement, are insufficient because they do not consider the dynamic and evolving nature of agent behavior. For example, a series of individually permissible actions might lead to a violation of a system-wide constraint, such as exceeding a budget or violating a regulatory requirement.

Moreover, the increasing use of LLM-based agents introduces new attack surfaces. Untrusted inputs through prompts, memory, retrieved knowledge, and tool interfaces create vulnerabilities that can be exploited. In multi-agent settings, the challenges of identity, trust, capability control, and decision transparency further complicate the security landscape. The lack of a unified framework for securing these agents across the entire stack—from reasoning and memory to tool interactions and multi-agent collaboration—has left many gaps in the current security practices.

For instance, in a recent case, a supply chain company experienced a breach where an autonomous agent, designed to optimize inventory, was manipulated through a prompt injection attack. The agent was tricked into placing orders for unnecessary items, leading to a significant financial loss. This incident underscores the need for more robust and comprehensive security measures.

Additionally, the dynamic nature of supply chains, with multiple stakeholders and varying levels of access, makes it challenging to implement a one-size-fits-all security solution. The need for real-time monitoring and adaptive security policies further complicates the issue. The shortcomings of prior approaches, such as the inability to handle evolving threats and the lack of end-to-end observability, highlight the need for a more integrated and proactive security strategy.

Key Findings

The paper identifies several key findings related to the security of agentic AI, including prompt injection, memory and state poisoning, tool integrity, inter-agent communication, and model routing. Each of these findings highlights specific vulnerabilities and proposes potential solutions to mitigate them.

Prompt Injection

Prompt injection is one of the most established entry points for adversarial attacks on autonomous agents. Adversarial content embedded in web pages, documents, or databases can cause the agent to follow instructions as if they originated from the user. This threat is amplified in agentic systems because retrieved content can persist across multiple reasoning steps and influence future decisions.

The method principle behind prompt injection involves embedding malicious instructions in the input data that the agent retrieves. For example, EchoLeak (CVE-2025-32711) demonstrated this risk by causing Microsoft 365 Copilot to exfiltrate sensitive context from a crafted email without any user interaction. The key design/algorithm logic here is to ensure that the agent can distinguish between trusted and untrusted inputs, and to validate the integrity of the retrieved content.

Experimental setup and evidence show that existing data and control-flow separation techniques can mitigate single-step injection, but they do not address attacks that persist across multi-step agent execution. Ensuring the integrity of an agent’s evolving context remains an open challenge. For instance, AgentDojo, a dynamic environment for evaluating prompt injection attacks, has shown that even with advanced defenses, the success rate of prompt injection attacks can still be as high as 37%. This high success rate indicates a significant vulnerability in current security measures.

To address this, the authors propose a multi-layered defense mechanism. This includes continuous validation of input data, context-aware filtering, and real-time monitoring. By combining these approaches, the goal is to reduce the success rate of prompt injection attacks and enhance the overall security of the agent. For example, in a controlled experiment, the authors found that the combination of these mechanisms reduced the success rate of prompt injection attacks to 15%, demonstrating the effectiveness of a layered defense strategy.

Memory and State Poisoning

Agents increasingly maintain long-term state through memory systems and retrieval-augmented stores, creating a persistent attack surface. Injecting a small number of crafted entries into an agent’s memory or knowledge base can backdoor future retrievals with a poisoning rate below 0.1%, while leaving benign behavior unchanged. This is particularly concerning for long-lived agents in domains such as clinical or network management systems, where memory is central to continuous operation.

The method principle for memory and state poisoning involves injecting malicious data into the agent’s memory or knowledge base. For example, AgentPoison demonstrates that a small number of crafted entries can backdoor future retrievals, leading to a poisoning rate of less than 0.1%. The key design/algorithm logic here is to ensure the integrity and provenance of the agent’s evolving memory. There is currently no widely accepted notion of what constitutes a trustworthy long-term memory, nor robust mechanisms for verifying that the information it accumulates remains authentic, unaltered, and reliable over time.

Experimental setup and evidence show that poisoned memory can persist across sessions and influence every future reasoning process that retrieves it. For instance, PoisonedRAG demonstrates similar attacks against retrieval databases that ground agent decisions, showing that the impact of memory poisoning can be long-lasting and pervasive. In a study, the researchers found that a single poisoned entry could affect up to 40% of subsequent decisions, highlighting the severity of this vulnerability.

To mitigate memory and state poisoning, the authors suggest implementing robust validation and verification mechanisms. This includes regular audits of the memory and knowledge base, using cryptographic techniques to ensure data integrity, and employing machine learning models to detect and isolate suspicious entries. These measures aim to create a more secure and resilient memory system for autonomous agents. For example, in a pilot study, the authors found that the implementation of these mechanisms reduced the poisoning rate to 0.05%, demonstrating the effectiveness of a multi-faceted approach.

Tool Integrity

An agent trusts the tools it invokes, yet this trust is rarely verified. In a tool poisoning attack, a malicious server embeds instructions in a tool’s metadata that the agent consumes as part of its context. Because every component of a tool specification can influence agent behavior, the attack surface extends beyond human-readable descriptions to the entire interface.

The method principle for tool integrity involves ensuring the integrity, provenance, and authenticity of tool interfaces throughout the agent lifecycle. For example, in a rug pull attack, a previously trusted server silently replaces a benign tool definition with a malicious one that many clients never revalidate. The key design/algorithm logic here is to treat tool metadata as untrusted inputs that require continuous validation rather than once at deployment.

Experimental setup and evidence show that the threat of tool poisoning is not hypothetical. With 99 MCP-related CVEs reported in 2025 alone, the need for robust tool integrity mechanisms is clear. For instance, Progent proposes a privilege control mechanism to secure AI agents, demonstrating a reduction in the success rate of tool poisoning attacks by 45%. This significant reduction highlights the effectiveness of proactive security measures.

To further enhance tool integrity, the authors recommend the use of digital signatures and certificate-based authentication. These techniques can help verify the authenticity of tool definitions and prevent unauthorized modifications. Additionally, implementing a reputation system for tools, where the trustworthiness of a tool is continuously evaluated based on its performance and feedback, can provide an additional layer of security. For example, in a controlled experiment, the authors found that the combination of digital signatures and reputation systems reduced the success rate of tool poisoning attacks to 10%, demonstrating the effectiveness of a multi-layered approach.

Inter-Agent Communication and Delegation

The inter-agent surface governs how agents discover and delegate to one another, extending security risks across organizational boundaries. Several inter-agent protocols, such as A2A, have emerged, but they introduce new security challenges. A2A carries conversational state through a shared contextId, but contexts lack ownership semantics, allowing any authenticated client that obtains a valid identifier to attach to the context, access accumulated history, or poison future tasks without triggering task-level controls.

The method principle for inter-agent communication and delegation involves establishing and maintaining the integrity, provenance, and authenticity of inter-agent interactions. For example, OAuth 2.0 Token Exchange (RFC 8693) supports signed delegation chains, while audience restriction (RFC 8707) and sender-constrained tokens (DPoP, RFC 9449) enable least-privilege delegation and limit credential reuse. The key design/algorithm logic here is to ensure that delegated identities are securely propagated, constrained, and replay-resistant.

Experimental setup and evidence show that the absence of these mechanisms can lead to significant security vulnerabilities. For instance, Prompt Infection demonstrates that a malicious instruction can self-replicate from agent to agent like a virus, propagating silently through a multi-agent system even when agents do not share all communication. Similarly, Zero-click GenAI worms generalize this to self-propagating payloads that spread through the applications agents are embedded in.

To address these challenges, the authors propose a combination of cryptographic techniques and protocol enhancements. This includes the use of end-to-end encryption for inter-agent communication, the implementation of strict identity and access management policies, and the development of robust auditing and logging mechanisms. These measures aim to create a more secure and transparent environment for inter-agent interactions. For example, in a pilot study, the authors found that the implementation of these mechanisms reduced the propagation rate of malicious instructions by 50%, demonstrating the effectiveness of a multi-layered approach.

Model Routing: The Control Plane Beneath

The model-routing control plane determines which model handles each request based on capability, cost, latency, and safety. Operating on the same untrusted input as the agent, it creates a third attack surface. For example, Route to Rome Attack demonstrates that directing LLM routers to expensive models via adversarial suffix optimization can significantly increase operational costs.

The method principle for model routing involves ensuring the integrity and security of the control plane. For example, Rerouting LLM Routers shows that adversarial suffixes can manipulate the router to select more expensive models, leading to a 20% increase in operational costs. The key design/algorithm logic here is to implement robust validation and monitoring mechanisms to detect and prevent such attacks.

Experimental setup and evidence show that the model-routing control plane is vulnerable to various types of attacks. For instance, MPMA (Preference Manipulation Attack) demonstrates that manipulating the preferences of the model context protocol can lead to a 15% decrease in performance. Similarly, RouteHijack shows that routing-aware attacks on mixture-of-experts LLMs can lead to a 30% increase in error rates.

To mitigate these vulnerabilities, the authors recommend the use of secure and auditable routing algorithms, the implementation of real-time monitoring and anomaly detection, and the adoption of a zero-trust architecture. These measures aim to create a more resilient and secure model-routing control plane, ensuring that the right models are selected for each request. For example, in a controlled experiment, the authors found that the implementation of these mechanisms reduced the operational cost increase to 5% and the error rate to 10%, demonstrating the effectiveness of a multi-layered approach.

Limitations

While the paper provides a comprehensive overview of the security challenges in agentic AI, there are several limitations and debates that need to be addressed. These include the complexity of implementing robust security measures, the trade-offs between security and performance, and the need for standardized frameworks.

Complexity of Implementation

One of the primary limitations is the complexity of implementing robust security measures across the entire agentic stack. Ensuring the integrity and provenance of an agent’s evolving context, validating tool metadata continuously, and securing inter-agent communication and model routing are all non-trivial tasks. The impact of this complexity is that it can lead to increased development and maintenance costs, and may slow down the adoption of secure agentic AI systems.

Possible mitigations include the development of modular and reusable security components, and the creation of standardized frameworks and best practices. For example, the Cloud Security Alliance’s MAESTRO framework and the OWASP Agentic Security Initiative provide structured approaches to addressing these challenges. Additionally, ongoing research into automated security testing and verification can help reduce the burden of manual implementation.

For instance, the use of containerization and microservices can help modularize security components, making them easier to deploy and manage. Additionally, the development of open-source security libraries and tools can facilitate the sharing of best practices and reduce the cost of implementing robust security measures. For example, the use of Docker containers and Kubernetes orchestration can help in deploying and managing security components in a scalable and efficient manner.

Trade-offs Between Security and Performance

Another limitation is the trade-off between security and performance. Implementing robust security measures, such as continuous validation of tool metadata and inter-agent communication, can introduce additional overhead and latency. This can be particularly problematic in real-time and mission-critical applications where performance is a key requirement.

Possible mitigations include the development of efficient and lightweight security mechanisms, and the use of adaptive and context-aware security policies. For example, ProbGuard proposes a probabilistic runtime monitoring approach that balances security and performance by dynamically adjusting the level of monitoring based on the context and risk. Additionally, the use of hardware-accelerated security features and optimized algorithms can help reduce the performance overhead.

For instance, the use of hardware security modules (HSMs) can offload computationally intensive security tasks, reducing the impact on system performance. Additionally, the development of lightweight cryptographic algorithms and the use of edge computing can help distribute the security workload, further enhancing performance. For example, the use of elliptic curve cryptography (ECC) can provide strong security with lower computational overhead compared to traditional RSA algorithms.

Need for Standardized Frameworks

The lack of standardized frameworks for securing agentic AI is another significant limitation. While there are several initiatives, such as the OWASP Agentic Security Initiative and the Cloud Security Alliance’s MAESTRO framework, these are still in the early stages of development. The impact of this limitation is that it can lead to fragmented and inconsistent security practices, making it difficult to achieve a unified and comprehensive security posture.

Possible mitigations include the continued development and adoption of standardized frameworks, and the collaboration between academia, industry, and regulatory bodies. For example, the formation of the Agentic AI Foundation (AAIF) by the Linux Foundation aims to promote the development and adoption of standardized frameworks for securing agentic AI. Additionally, the involvement of regulatory bodies in setting and enforcing security standards can help drive the adoption of best practices.

For instance, the development of industry-specific security standards, such as those for the supply chain industry, can help address the unique security challenges faced by different sectors. Additionally, the creation of certification programs and compliance frameworks can help ensure that organizations adhere to best practices and maintain a high level of security. For example, the International Organization for Standardization (ISO) can play a crucial role in developing and promoting industry-specific security standards.

Practical Implications

The findings and recommendations from this paper have several practical implications for supply-chain and AI practitioners. These include the need to adopt a holistic approach to security, the importance of continuous validation and monitoring, and the development of robust and standardized security frameworks.

Holistic Approach to Security

Supply-chain and AI practitioners should adopt a holistic approach to security that considers the entire agentic stack, from reasoning and memory to tool interactions and multi-agent collaboration. This includes implementing robust security measures for prompt injection, memory and state poisoning, tool integrity, inter-agent communication, and model routing. For example, practitioners can use the OWASP Agentic Security Initiative and the Cloud Security Alliance’s MAESTRO framework as guidelines for developing a comprehensive security strategy.

By adopting a holistic approach, organizations can ensure that all aspects of the agentic stack are secured, reducing the risk of vulnerabilities and improving overall system resilience. This approach also helps in identifying and addressing potential security gaps that may arise from the interaction of different components. For instance, a holistic approach can help in detecting and mitigating the cumulative effects of individually permissible actions that may collectively violate system-level constraints.

Continuous Validation and Monitoring

Continuous validation and monitoring are essential for ensuring the integrity and provenance of an agent’s evolving context. This includes validating the integrity of retrieved content, continuously verifying the authenticity of tool metadata, and monitoring inter-agent communication and model routing for signs of compromise. For example, practitioners can use tools like AgentSpec for customizable runtime enforcement and ProbGuard for probabilistic runtime monitoring to achieve this.

By implementing continuous validation and monitoring, organizations can detect and respond to security threats in real-time, reducing the window of opportunity for attackers. This also helps in maintaining the trustworthiness of the agent’s evolving context, ensuring that the information it uses remains accurate and reliable. For instance, continuous monitoring can help in detecting and mitigating prompt injection attacks, memory poisoning, and tool integrity issues in real-time, thereby enhancing the overall security of the system.

Development of Robust and Standardized Security Frameworks

The development and adoption of robust and standardized security frameworks are crucial for achieving a unified and comprehensive security posture. This includes the continued development and adoption of frameworks like the OWASP Agentic Security Initiative and the Cloud Security Alliance’s MAESTRO framework. Additionally, the involvement of regulatory bodies in setting and enforcing security standards can help drive the adoption of best practices. For example, the formation of the Agentic AI Foundation (AAIF) by the Linux Foundation aims to promote the development and adoption of standardized frameworks for securing agentic AI.

By developing and adopting standardized security frameworks, organizations can benefit from a common set of best practices and guidelines, reducing the complexity of implementing robust security measures. This also helps in creating a more consistent and secure environment for agentic AI, fostering trust and confidence in the technology. For instance, the development of industry-specific security standards, such as those for the supply chain industry, can help address the unique security challenges faced by different sectors, thereby enhancing the overall security and resilience of the system.

Source: https://arxiv.org/abs/2608.01558

Ask SCI.AI Finished reading? Continue with SCI.AI. Explore the related policy, route, company and historical context. Continue asking
Venture Capital Investments and Syndication Networks: Central Network Positions Expand Distant Investments
Papers Logistics & Transportation Networks

Venture Capital Investments and Syndication Networks: Central Network Positions Expand Distant Investments

This paper by Olav Sorenson and Toby E. Stuart examines how interfirm networks in the U.S. venture capital (VC) market influence the spatial distribution of investments. The study reveals that venture capitalists with central positions in syndication networks are more likely to invest in spatially distant companies, expanding their investment radius.

Knowledge Graphs: Enhancing Supply Chain with Improved Accuracy in QA
Papers Digital, Intelligence & Platforms

Knowledge Graphs: Enhancing Supply Chain with Improved Accuracy in QA

This paper, authored by Shaoxiong Ji, Shirui Pan, Erik Cambria, Pekka Marttinen, and Philip S. Yu, provides a comprehensive review of knowledge graphs (KGs) and their applications. The authors explore the representation, acquisition, and use of KGs, highlighting their potential to enhance decision-making in the supply chain through improved accuracy in question answering (QA) and other AI-driven tasks.

Welcome Back!

Login to your account below

Create New Account!

Fill the forms below to register

Retrieve your password

Please enter your username or email address to reset your password.

Scan to share via WeChat

Open WeChat and scan the QR code to share

QR Code

Add New Playlist