What is prompt injection?
Prompt injection is a vulnerability in AI systems in which content processed by the model influences the model's behavior or output in an unintended way. The vulnerability can be exploited through an attack in which an attacker attempts to manipulate the instructions the model follows. The objective may be to make the model ignore its original instructions, disclose information, or perform actions it was not intended to perform.
Prompt injection can occur directly when a user gives the model manipulative instructions, or indirectly when instructions appear in content processed by an AI system. This could include text on a website, in a document, or in other data sources the model can access. The instructions may be visible, hidden, or encoded in a way that allows them to be processed by the model without necessarily being apparent to the user. If the AI system interprets the content as instructions, it may influence the model's subsequent actions or responses.
The risk becomes particularly relevant when AI systems do more than generate text and also have access to organizational data, applications, and other tools. A successful prompt injection attack could potentially influence actions performed through systems the model has access to.
Prompt injection can be compared to giving an employee a document containing messages that tell them to ignore their actual work instructions. If the employee cannot distinguish between information they should process and instructions they should follow, those messages may influence what they do.
Sicra and prompt injection
As organizations integrate AI models with internal data, applications, and automated processes, it becomes important to control which resources AI systems can access and which actions they can perform.
Protecting against prompt injection therefore involves more than the AI model itself. Identity management, access control, limited privileges, secure architecture, and control over integrations can help reduce the consequences if a model is manipulated.
No single control eliminates the risk. Protection should therefore consist of multiple layers, including restricting the model's access, validating tool calls, monitoring, and requiring human approval for high-impact actions.
Sicra helps organizations with security strategy, identity security, and security architecture that can help reduce risk when AI is integrated with organizational systems and data.
Services
Security strategy
Security maturity assessment
Identity maturity assessment
CISO-for-hire
Related terms: Artificial intelligence (AI), LLM (Large Language Model), Agentic AI, GPAI (General-Purpose AI), NHI (non-human identities), IAM (Identity and Access Management), Identity security, Zero Trust, Cybersecurity