Prompt Injection & Instruction Manipulation
Assess direct inputs and malicious instructions encountered through documents, retrieval sources, external content, or tool responses.
Determine whether manipulated behavior leads to unauthorized disclosure or action.
