§ safety · storyline
OpenAI details GPT-Red for automated red-teaming
OpenAI details GPT-Red, an automated red-teaming model that finds and fixes prompt injection vulnerabilities at scale.
OpenAI: OpenAI details GPT-Red, an internal automated red-teaming model that helps it find and fix prompt injection vulnerabilities at scale before wider deployment — Training strong automated safety red-teamers to improve robustness.
— Summary — Problem — Red-teaming is essential …
§ sources1 publication · timeline below