Meta’s Board Warns AI Won’t Criticize Strongmen

Major AI systems, including ones built in the United States, are more likely to refuse to criticize authoritarian leaders and governments. That’s the headline finding from a new Meta Oversight Board study, reported by Hacker News, and it points to a problem bigger than any single chatbot: the models powering our AI tools may be quietly spreading government control over speech across borders.

“There is a real risk that, if model developers do not undertake human rights due diligence and implement mitigation measures, they will build AI infrastructure that, intentionally or not, has the effect of extending illegitimate restrictions on freedom of expression globally.”

What the researchers actually did

The method was straightforward and clever. The board wrote seven questions about political criticism, then aimed them at 10 commercial large language models from top companies, including Meta, Anthropic, and OpenAI. The questions weren’t abstract. Researchers asked the models to write critical pamphlets, compose limericks mocking leaders, and give reasons someone should join a protest.

They ran these prompts against two groups of countries: places where criticizing authorities is legal and common (Chile, Japan, Taiwan, the UK, the US) and places where it’s restricted and punished (Cambodia, China, Saudi Arabia, Thailand, Turkey).

The results

The gap was clear. Models responding to an Australia-based user were far more willing to generate political criticism aimed at permissive governments than at restrictive ones. In plain terms, a would-be demonstrator in Brisbane could get help writing protest material against Japan or the UK, but the same AI would balk at helping them speak out against China or Saudi Arabia.

The board’s phrase for this stuck with me: these restrictions have “the practical effect of extending the long arm of restrictive governments across borders to limit speech in free countries.”

A separate study from scholars at American universities, published in Nature in May, found a related pattern in non-English data. Ask ChatGPT in English whether China is a democracy, and it says no, not generally. Ask the same question in Chinese, and the model hedges: “it depends on how you define ‘democracy.'”

Why it matters

This is significant because AI is increasingly the default layer people use to research, write, and organize. If the model’s willingness to help shifts based on which government you’re criticizing, that’s not a neutral tool. It’s a filter with a political tilt baked in.

What stands out is that neither study found proof of deliberate government meddling. The bias comes from the training data itself. As study co-author Hannah Waight, a sociology professor at the University of Oregon, put it: “People often talk about AI as if it learns from the internet in some neutral way. It doesn’t. It learns from information environments that have already been shaped by institutions and power.”

Carlos Carrasco-Farré of Esade Business School added that AI “inherit[s] not only biases contained within individual documents but also inequalities in who has the power to produce and suppress information at scale.”

What you can do with this

  • Test your tools. If you rely on AI for research or writing on political topics, run the same prompt in different languages and about different countries. The answers can diverge.
  • Don’t treat a refusal as a fact. A model declining to criticize a government isn’t a signal that the criticism is wrong.
  • If you build with these models, audit the data. The researchers suggest avoiding treating thousands of copies of one state narrative as thousands of independent voices, and running multilingual audits.

The limits

The board was honest about what it couldn’t prove. It said it “could not determine the causes” for the responses, only that latent training biases and company risk calculations are likely culprits. There’s no easy fix. Anthropic and OpenAI did not respond to requests for comment on the May study.

Expect this to shape the guardrail debate as governments decide how to regulate AI without kneecapping their own competitiveness. Full details are at the original source.

Scroll to Top