The largest security mistake in an AI application is treating model output as a trusted decision. Prompt injection, unsafe tool use and excessive permissions become more serious when an AI agent can act.

The OWASP Top 10 for LLM Applications 2025 places prompt injection first and expands its treatment of excessive agency. It also highlights risks in sensitive information disclosure, supply chains, data and model poisoning, output handling, vector stores, misinformation and unbounded consumption. NIST's 2025 adversarial machine learning taxonomy provides a complementary way to describe attacks across the AI lifecycle.

Separate data from instructions

Retrieved documents, web pages, emails and user content are untrusted inputs. An attacker can place instructions inside them and attempt to redirect the model. Label trust boundaries, constrain how content is used and avoid giving retrieved text authority over system behaviour.

Reduce agency before adding detection

Give an agent only the tools and data required for its task. Restrict parameters, destinations and transaction values. Use separate identities for separate functions. Require human approval for consequential actions such as sending external messages, changing records, executing code or moving money.

Validate outputs at the point of use

Model output that enters a database, browser, shell, API or workflow must be validated as untrusted input. Use allowlists, typed interfaces and conventional application security controls. The model should not construct unrestricted commands for a privileged downstream system.

Test the complete workflow

Model-only testing misses the permissions and integrations that create business impact. Test direct and indirect prompt injection, data extraction, tool misuse, cross-user leakage, unsafe retrieval and failure of approval gates. Repeat testing when models, prompts, tools or data sources change.

Keep enough evidence to investigate

Record the relevant input, retrieved context, model and version, tool calls, approvals and final actions. Protect sensitive data in logs and define retention. Without this evidence, a team may know that an outcome was wrong without being able to explain or contain it.

Assume the model can be influenced

Security comes from limiting what the system can access and do, validating every handoff and retaining human control over high-impact actions.

Sources and further reading