OpenAI publishes GPT-5.1 safety evaluations
OpenAI has updated the GPT-5.1 Instant and GPT-5.1 Thinking safety card with new reference metrics and tests covering mental health and emotional dependency. The models also adjust their reasoning time according to the complexity of each query.

OpenAI has published an update to the GPT-5.1 safety card for its new conversational and reasoning models. The document does not describe a system entirely different from GPT-5, but it adds safety data for the new versions and expands the testing conducted before deployment.
Two models with more adaptable use
GPT-5.1 Instant is designed to respond more conversationally and follow instructions more closely. It also introduces adaptive reasoning: it can decide when it needs more time to think before answering and when a direct response is enough.
GPT-5.1 Thinking, by contrast, adjusts more precisely how much time it spends analyzing each question. The idea is to avoid giving every query the same level of processing: a simple request can be handled quickly, while a complex problem may require more analysis.
For most users, model selection will remain automatic. GPT-5.1 Auto will continue routing each query to the model it considers most appropriate, without making you choose between Instant and Thinking.
More testing for sensitive conversations
The general safety measures are largely the same as those OpenAI already described for GPT-5. The document's main addition is its updated reference metrics and the broader scope of its pre-launch evaluations.
OpenAI has added specific tests for two particularly sensitive areas:
- Mental health: these evaluate responses in conversations involving signs of isolated delusions, psychosis, or mania.
- Emotional dependency: these analyze whether the model can create or reinforce an unhealthy bond of dependency or attachment to ChatGPT.
These evaluations do not mean the model can diagnose an illness or reliably detect a person's mental state. They are designed to check how it responds to risky situations and identify behaviors it should avoid.
The document also clarifies the technical names used in the evaluation: GPT-5.1 Instant appears as gpt-5.1-instant, and GPT-5.1 Thinking as gpt-5.1-thinking.
What changes for you
In daily use, the most visible change should be more natural conversations, more precise instruction-following, and responses that better adjust their level of analysis. For example, the system could answer a writing request quickly while spending more time reasoning through a question with several steps.
The safety update matters most when a conversation touches on sensitive topics. OpenAI is expanding its controls so it evaluates not only whether the model avoids dangerous content, but also whether it responds responsibly when someone shows signs of emotional or psychological vulnerability.
The next thing to watch will be how these evaluations translate into GPT-5.1's real-world behavior. The document updates the tests and reference metrics, but the material provided includes no specific figures and does not claim that every risk has been resolved.