GPT-Red: Unlocking Self-Improvement for Robustness
OpenAI, Wednesday, July 15th, 2026
OpenAI details GPT-Red, an automated red-teaming model trained via self-play that cut GPT-5.6 Sol prompt injection failures to 0.05%.
OpenAI introduced GPT-Red, its automated safety red-teaming model and the culmination of its work on machine-driven adversarial testing. GPT-Red uses self-play to improve safety, alignment and prompt injection robustness, pursuing an attack goal by sending a prompt, observing how target GPT models respond, and iterating as a human red-teamer would.
It was trained at the compute scale of some of OpenAI's largest post-training runs, an unprecedented allocation dedicated purely to safety.
OpenAI incorporates GPT-Red directly into training its production models, framing automated red-teaming as self-improvement in which today's models make future models safer. GPT-5.6 Sol fails on only 0.05% of GPT-Red's direct prompt injections.