Back to glossary
AI quality and safety

Prompt injection

An attack or accidental instruction that tries to make an AI system ignore its intended rules, reveal data, or misuse tools.

Business example

A malicious instruction hidden in a webpage tells an agent to disclose confidential information when it summarizes the page.

How it works

  1. 1.Untrusted content enters the model context.
  2. 2.The content imitates a higher-priority instruction.
  3. 3.Isolation, least privilege, and confirmation limit the possible impact.

Common misconceptions

A stronger system prompt solves prompt injection.
Prompting helps, but robust defense also requires tool permissions, data boundaries, and monitoring.

Related concepts

Related reading

Sources

Reviewed:

Reviewed by: Javier Chulvi Bernad · LLM Engineer · Madrid