OpenAI’s GPT-Red uses self-play red teaming to improve AI safety, alignment, and prompt injection robustness.
Read the original at openai.com→Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness.
Original headline: "GPT-Red: Unlocking Self-Improvement for Robustness"
Coverage timeline
- Jul 15, 10:00 UTC OpenAI lead source GPT-Red: Unlocking Self-Improvement for Robustness
- Jul 15, 17:09 UTC MIT Technology Review AI Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer