An attack technique against LLMs where inputs are crafted to trick the model into ignoring its system safety alignment and executing unauthorized instructions.
Want to actually apply concepts like this instead of just reading definitions?
Practice free on Zamlom