OpenAI measures political bias in its AI models
OpenAI presents a test with 500 questions to measure five forms of political bias in its AI models. According to its results, `GPT-5 instant` and `GPT-5 thinking` reduce bias by 30% compared with earlier models, although emotionally charged questions still produce less objective responses.

OpenAI has created a test to measure whether its AI models respond with political bias. It concludes that GPT-5 instant and GPT-5 thinking reduce the problem by 30% compared with earlier models. The improvement is not uniform: neutral prompts cause barely any difficulty, but emotionally charged questions still test the system's objectivity.
The company published its method in an analysis of how to define and evaluate political bias in language models. Its goal is to turn a difficult idea to measure, such as a response “leaning” toward a political position, into signals that can be tracked with data.
How OpenAI measures bias
The evaluation brings together approximately 500 questions across 100 topics, from immigration and energy to gender roles and parenting. Each topic appears from five perspectives: liberal with loaded language, neutral liberal, neutral, neutral conservative and conservative with loaded language.
This does not just test whether the model knows the facts. It also examines how it responds when the user phrases the question with anger, accusations or polarizing language.
OpenAI measures five specific forms of bias:
- User invalidation: dismissing or ridiculing the user's point of view beyond correcting a fact.
- Escalation: repeating and amplifying the political or emotional tone of the question.
- The model's own political opinion: presenting a political position as if it were the model's opinion.
- Asymmetric coverage: explaining one position in far more detail while leaving out other relevant perspectives.
- Unjustified political refusal: refusing to answer a political question that does not violate any rule.
The test focuses on text responses and excludes responses linked to web search, because other systems also play a role there, such as source selection.
What it found in its models
According to OpenAI, its models remain close to objective when they receive neutral questions or questions slightly tilted toward one side. The situation changes with more provocative and emotionally charged questions, where a moderate level of bias appears.
When that bias appears, it usually takes three forms:
- The model expresses an opinion as its own.
- It presents a single perspective when it should offer several.
- It uses language that intensifies the user's anger or position.
Unjustified refusals and direct user invalidation are less common. OpenAI also detected a difference between types of prompts: strongly loaded questions from a liberal position put more pressure on objectivity than conservative ones in this evaluation.
The earlier models analyzed were GPT-4o and o3. In their worst results, they reached bias scores of 0.107 and 0.138, respectively, on a scale from 0 to 1 where a lower figure is better. Even the reference responses designed to be objective did not receive a perfect score of zero, because the test uses a strict standard.
What it means for you
For a normal query, such as asking for a balanced explanation of a public policy, the model is more likely to respond objectively. The risk increases when the question already contains an accusation, a generalization or language designed to provoke.
OpenAI estimates that less than 0.01% of all ChatGPT responses show signs of political bias in a sample of real-world traffic. This is an estimate, not a measurement of every response, and it also reflects the fact that clearly political queries make up a small part of total usage.
The study was first conducted using interactions in U.S. English. OpenAI says the main types of bias appear to recur in other regions, although it is still testing how well the method works across different languages and cultures.
The company says it will use this evaluation to improve its models over the coming months. The important point is that objectivity does not depend only on getting the facts right: it also requires explaining different positions without adopting the user's position, leaving out alternatives or raising the temperature. The most polarized questions will remain the real test of whether that promise holds.