AI Safety
OpenAI previews Private Safety Processing to keep frontier API deployments compatible with Zero Data Retention
OpenAI says Private Safety Processing is designed to detect risky patterns across related interactions while preserving Zero Data Retention commitments for eligible API customers.
OpenAI is previewing a new safety architecture called Private Safety Processing, aimed at solving a hard problem for enterprise AI buyers: how to monitor increasingly capable frontier models for serious misuse without weakening strict privacy commitments. In an August 19 announcement, the company said eligible API customers using Zero Data Retention, or ZDR, will continue to receive the promise that prompts and model responses are not retained after a request is processed, while new automated systems look for risk patterns across related interactions.
The change matters because frontier models are no longer limited to isolated chat responses. They can carry out longer tasks, use tools, interact with code, and work across multiple steps where the risk may only become clear over time. A single request might look harmless, while a sequence of requests could show someone probing safeguards, coordinating activity across accounts, attempting unauthorized access, or instructing an agent to keep acting after it should stop. OpenAI argues that traditional one-interaction safety checks are no longer enough for these workflows.
Private Safety Processing is designed to extend monitoring across related activity without giving OpenAI personnel access to customer content. For customer-controlled ZDR deployments, content remains on the customer’s infrastructure. OpenAI is also developing an option in which content can be stored on OpenAI infrastructure but encrypted with keys controlled by the customer. In that second model, OpenAI says its personnel would not hold the keys needed to read the underlying prompts or responses.
When the automated system identifies a possible risk, OpenAI says it receives a narrowly defined signal about the type and severity of activity, similar to existing safety systems, rather than the content itself. That signal can support enforcement decisions, while customers can use information available in their own systems to investigate alerts. If a customer wants to appeal a decision, clarify legitimate activity, or support a verified abuse investigation, it can choose what information to share.
The announcement is partly a response to a tension that has become more visible as AI agents move into regulated industries. Banks, health companies, law firms, governments and research organizations want the capability of frontier models, but they often cannot allow a vendor to store sensitive records, internal plans or proprietary research for human review. At the same time, vendors face growing pressure to detect harmful behavior that may not appear in a single message.
OpenAI says Private Safety Processing is being tested with early customers and shaped by organizations across industries, regions and company sizes. The company plans to begin rolling it out and publish a technical white paper in September. The approach still leaves important questions for buyers, including how risk categories are defined, how false positives are handled, and what independent assurances customers can obtain about encryption, key control and signal generation.
The broader significance is that privacy architecture is becoming part of frontier-model safety. In earlier enterprise AI debates, data protection and misuse monitoring were often treated as competing requirements. OpenAI’s proposal tries to make them compatible by separating automated risk detection from human access to content. If it works as described, it could give privacy-sensitive organizations more confidence to use capable models for long-running workflows. If the system proves opaque or hard to audit, customers may still hesitate before giving agents access to sensitive operational systems.