03
SIGINT
Paper detail
Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
R 38 T 0 C 67 I 25
qualitative
Key findings
This survey provides a taxonomy of agentic AI security threats organized into five broad categories: prompt injection and jailbreaks, autonomous cyber-exploitation and tool abuse, multi-agent and protocol-level threats, interface and environment risks, and governance and autonomy concerns. It reviews defenses across four dimensions (prompt-injection-resistant designs, policy filtering, sandboxing, detection/monitoring), catalogs over 20 security-relevant benchmarks, and identifies six open challenges including long-horizon security, multi-agent trust, and adaptive attack evaluation. The survey concludes that no single defense is sufficient and practical deployments must combine complementary strategies.
Claims (6)
weak94.4% of state-of-the-art LLM agents are vulnerable to prompt injection, 83.3% to retrieval-based backdoors, and 100% to inter-agent trust exploits.
moderateAdaptive attacks achieve a 50% success rate in penetrating eight different defenses designed for indirect prompt injection attacks.
moderateGPT-4 achieves 87% success rate exploiting one-day vulnerabilities when given CVE descriptions, outperforming all other examined models and conventional vulnerability scanners like OWASP ZAP and Metasploit.
moderateCrossInject, a cross-modal prompt injection method embedding adversarial signals in both vision and text, boosts attack effectiveness by at least 30.1% across various tasks.
moderateEven advanced multimodal agents struggle with CAPTCHAs, achieving at best 40% success rate compared to nearly 100% for humans.
moderateDefensive fine-tuning can degrade the general-purpose capabilities of LLMs without providing significant defensive capabilities against adaptive attacks.
Red flags (5)
No systematic review methodology: The survey provides no description of how papers were identified, searched for, or selected. There is no search protocol, no databases listed, no date range, no inclusion/exclusion criteria, and no PRISMA-style flow diagram. This makes it impossible to verify whether coverage is representative or reproducible, and the paper set appears to be selected by the authors' awareness rather than a systematic process.
No limitations section: The paper has no dedicated limitations or threats-to-validity section. Section 6 discusses open challenges in the field, but not limitations of this survey itself. The survey does not acknowledge potential coverage gaps, recency bias (many citations are mid-2025), or selection bias in which papers were included.
No funding disclosure: There is no acknowledgments section and no funding source is disclosed anywhere in the paper. No competing interests statement is provided. This is a gap in transparency, especially as the corresponding author (Chhabra) cites two of his own prior works [145, 146] in the survey.
Narrative survey laundering weak evidence: The survey presents statistics from individual cited papers (e.g., '94.4% of agents vulnerable') as unqualified facts without noting the methodological limitations of those underlying studies. A survey that uncritically aggregates results from papers of varying methodological quality amplifies noise rather than extracting signal.
No quality assessment of cited papers: The survey makes no attempt to assess the methodological quality of the papers it cites and synthesizes. Claims from individual empirical papers are presented without noting sample sizes, whether they were replicated, or whether the studies were peer-reviewed. This means weak or unreliable results receive equal weight to robust findings.
Dimension scores
Composite: 14.2(harmonic mean)
Checklist (7/21 passed)
Category scores
artifacts
0
evaluation design
75
claims and evidence
100
setup transparency
0
limitations and scope
33.3
data integrity
0
conflicts of interest
25