Computer ScienceEngineering

M. Ferrag, Abderrahmane Lakas, N. Tihanyi, Mérouane Debbah

2026.3.1Internet of Things and Cyber-Physical Systems

DOI: 10.1016/j.iotcps.2026.03.001

Abstract

Large Language Models (LLMs) are rapidly transitioning from standalone conversational systems to autonomous agents that reason, plan, and interact with external tools. While this shift enables powerful applications in domains such as healthcare, finance, law, and software engineering, it also introduces new security and safety risks. Attacks such as prompt injection, jailbreak exploits, backdoor triggers, and multimodal adversarial inputs expose vulnerabilities not only at the model level but also across the broader agentic workflow. Existing defenses—ranging from input filtering and alignment reinforcement to runtime monitoring—remain fragmented and often fail to anticipate adaptive adversaries. Meanwhile, red teaming has emerged as a critical methodology for stress-testing these systems; however, current efforts lack standardization, comprehensive coverage across modalities, and integration with agent-specific contexts. This paper provides the first comprehensive survey of LLM agent security, synthesizing research on attack strategies, red teaming frameworks, evaluation suites, and defense mechanisms. We categorize automated and agentic red teaming approaches, highlight domain-specific vulnerabilities in code, web, and multimodal agents, and analyze defense strategies spanning prompt-level, decoding-time, runtime, backdoor, privacy-preserving, and multi-agent safeguards. Building on this synthesis, we outline key open challenges and future research directions, including the need for scalable defenses, standardized benchmarks, robustness against adaptive attacks, explainability, and secure integration of multi-agent workflows. Our findings aim to guide both researchers and practitioners in advancing robust, trustworthy, and resilient LLM-powered agents for safety-critical applications.

Citation format

FERRAG, M., et al. Securing LLM agents: From prompt sanitization to autonomous red teaming and beyond. Internet of Things and Cyber-Physical Systems, 2026, 5: 185–209.