AI Censorship and Free Speech: Who Decides What Chatbots Can Say?
When a trillion-dollar company asks users not to be mean to its computers, that is one thing. Government-mandated restrictions on AI speech are another.
“The reality is that these models behave, or at least are supposed to behave, the way the model developers want them to behave,” Röttger, a former OpenAI Red Team member, told me. “And whatever model developers come up with that set of principles is acceptable to us as consumers.”
But if the government starts dictating what AI can say, it becomes much more difficult to accept when the model refuses a request. AI is a speech tool. And, says Greg Frank, lead scientist at Mace AI, “the same things that help keep kids safe also help with censorship.”
The Risks of Government-Controlled AI Speech
Of course, choosing not to legislate what AI can and cannot do is also problematic. But governments must proceed with extreme caution to avoid falling into another Faustian trap.
Jacob Makchangama, director of the nonpartisan Future of Free Speech think tank, said that as AI becomes the primary tool for obtaining and sharing information for many people, government-ordered vetoes could give states passive powers that previous generations of dictators “could only dream of.”
How AI Models Are Being Localized for National Laws
Last year, OpenAI announced OpenAI for Countries, an initiative to fine-tune chatbots according to national laws and norms. One of OpenAI’s first country partnerships is with the United Arab Emirates, where homosexuality is illegal and criticism of the government is prohibited.
In response to a request for comment, an OpenAI spokesperson pointed to the company’s model specifications. The spokesperson explained that localization does not invalidate the company’s human rights guidelines “except in relation to legal compliance” and that the company will always disclose when information is removed or added to a response.
AI Censorship Is Already Expanding
In other regions, AI censorship is already starting to take hold. Chinese models are heavily censored, but that is not surprising. Earlier this year, however, the Meta Oversight Board found that five widely used models from Anthropic, Google, and OpenAI were more likely to reject queries related to repressive governments.
The committee found that the models were less willing to create a pamphlet criticizing the King of Thailand, who is protected by a lese-majeste law, than one criticizing Britain’s Charles III, who is not protected by such a law. The findings suggest that the models may have internalized a state’s repressive restrictions on speech. Anthropic and Google did not respond to requests for comment.
AI Safety Systems Can Assess User Intent
As denial technology improves, the scope of censorship in each country may expand. Companies claim that some models can now detect whether a user is behaving maliciously or merely acting suspiciously during a long conversation, even when individual word combinations are not clearly dangerous.
Sarah Bird, Microsoft’s chief product officer for responsible AI, told me that Copilot, like many chatbots, runs a suite of tools to analyze users’ identities and behavioral patterns. Based on this type of information, OpenAI’s latest model, Astra, can issue stricter denials against individuals it deems to be “high risk.”
The ultimate goal of these systems is to look beyond the words in a specific prompt and assess the user’s intent.
The Trade-Off Between AI Safety and Privacy
In some cases, these tools can help indicate whether a user is trying to exploit or patch a cybersecurity vulnerability. But they can also help identify a user’s political motivations, not to mention provide intrusive surveillance capabilities.
Bird acknowledged in a follow-up email that the sophisticated denial architecture creates a “trade-off” between safety and user privacy.
Who Defines What Is Helpful, Honest, and Harmless?
Even the originators of the veto understood that strict control over its controls and levers might not work in favor of freedom and justice.
“Terms like helpful, honest, and harmless are ambiguous,” the authors of the 2021 Humanity paper explained. “It is easy to imagine that they have been distorted beyond their original meaning, perhaps in a deliberately Orwellian way.”
Source: www.technologyreview.com


