Prompt injection

An attack where instructions ride in on data the model treats as trustworthy, such as a retrieved document or a web page, and override the developer prompt.

Why exams ask this

Responsible-AI sections test the mitigation, not the definition. The honest answer is that no prompt wording fully solves it, so the control is at the tool and permission boundary.

Relevant to

Related concepts

Resources

No resources linked to this concept yet.