Back Issues This Week → Calendar → Current Issue → Popular →

All issuesVolume 340, Issue 3IT Vendor NewsOpenAI

GPT-Red: Unlocking Self-Improvement for Robustness

OpenAI, Wednesday, July 15th, 2026

OpenAI details GPT-Red, an automated red-teaming model trained via self-play that cut GPT-5.6 Sol prompt injection failures to 0.05%.

OpenAI introduced GPT-Red, its automated safety red-teaming model and the culmination of its work on machine-driven adversarial testing. GPT-Red uses self-play to improve safety, alignment and prompt injection robustness, pursuing an attack goal by sending a prompt, observing how target GPT models respond, and iterating as a human red-teamer would.

It was trained at the compute scale of some of OpenAI's largest post-training runs, an unprecedented allocation dedicated purely to safety.

OpenAI incorporates GPT-Red directly into training its production models, framing automated red-teaming as self-improvement in which today's models make future models safer. GPT-5.6 Sol fails on only 0.05% of GPT-Red's direct prompt injections.

more →  ·  More from OpenAI →