Poisoning models
-
LLM Injection: «The concern I described is that an attacker might be able to craft special kind of text, put it up somewhere on the internet, so that when it later gets pick up and trained on, it poisons the base model in specific, narrow settings (e.g. when it sees that trigger phrase) to carry out actions in some controllable manner (e.g. jailbreak, or data exfiltration).» (source)
-
LLM Sleeper Agent: Training Deceptive LLMs that Persist Through Safety Training