Prompt Injection-attacks
User input flowing into ChatGPT can be manipulated and misled through so-called prompt injection attacks.
Attackers craft prompts to force the model to give secret or forbidden answers.
This leads to the leaking of confidential data, generating dangerous code, or bypassing content filters. Because the model is so flexible in interpreting complex questions, a successful attack can lead the model to ignore certain rules or ethical guidelines.
Preventing and detecting this is a huge challenge, as the possible input is endless and the model must remain flexible to function well.


