{"id":23724,"date":"2026-10-09T12:23:57","date_gmt":"2026-10-09T12:23:57","guid":{"rendered":"https:\/\/news678.top\/?p=23724"},"modified":"2026-10-09T12:23:57","modified_gmt":"2026-10-09T12:23:57","slug":"were-putting-too-much-faith-in-ais-ability-to-say-no","status":"publish","type":"post","link":"https:\/\/news678.top\/?p=23724","title":{"rendered":"We\u2019re putting too much faith in AI\u2019s ability to say no"},"content":{"rendered":"<p><\/p>\n<div>\n<p>If a trillion-dollar company doesn\u2019t want you to be mean to its computer, so be it. \u201cThe reality is that these models behave, or at least are supposed to behave, in the way that the model developers want them to behave,\u201d R\u00f6ttger, the former OpenAI red-teamer, told me. \u201cAnd however the model developers come up with that set of principles, that is kind of for us, the consumers, to accept.\u201d\u00a0<\/p>\n<p>But if governments get to dictate what <em>all <\/em>models refuse, that will be much harder to accept. AI is a tool for speech. And as Greg Frank, the chief scientist of Mace AI, puts it, \u201cThe same thing that serves child safety also serves censorship.\u201d\u00a0<\/p>\n<p>Choosing not to enact laws for what AI can and cannot do would, of course, be insane. But we\u2019ll need to tread with utmost care, lest we fall into another Faustian trap. As AI becomes many people\u2019s primary tool for retrieving and sharing information, says Jacob Mchangama, director of the nonpartisan think tank The Future of Free Speech, dictating refusal could give states a muffling power that earlier generations of autocrats \u201ccould only dream of.\u201d\u00a0<\/p>\n<p>Last year, OpenAI announced an initiative, OpenAI for Countries, that would fine-tune its chatbots in accordance with national laws and norms. One of OpenAI\u2019s first country partnerships is with the United Arab Emirates, where homosexuality is illegal and criticism of the government is forbidden. In response to a request for comment, an OpenAI spokesperson pointed to the company\u2019s model spec, which explains that localization won\u2019t override the company\u2019s human rights guidelines \u201cexcept as it relates to legal compliance,\u201d and that it will always disclose whenever information is removed from or added to a response.<\/p>\n<p>Elsewhere, AI censorship has already begun to take hold. Chinese models are, of course, highly censored\u2014that\u2019s no surprise. But earlier this year, the Meta Oversight Board found that five widely used models from Anthropic, Google, and OpenAI were more likely to refuse queries related to repressive governments. The board found that models were less willing to create a pamphlet criticizing the king of Thailand, which has l\u00e8se-majest\u00e9 laws, than Charles III of England, which doesn\u2019t. The results, they say, suggest that the models have somehow internalized repressive national limits on speech. Anthropic and Google did not respond to requests for comment.<\/p>\n<div class=\"wp-block-image\">\n<figure class=\"wp-block-image size-large\"><img decoding=\"async\" height=\"2000\" width=\"1536\" src=\"https:\/\/wp.technologyreview.com\/wp-content\/uploads\/2026\/10\/0924SP2.jpg?w=840\" alt=\"&quot;&quot;\" class=\"wp-image-1145875\"\/><\/p>\n<p>RAVEN JIANG<\/p>\n<\/figure><\/div>\n<p>As refusal techniques improve, they could expand states\u2019 censorial reach. Companies claim that some models can now detect if a user is being nefarious, or merely a bit suspicious, over the course of a long conversation\u2014even when none of the individual combinations of words used are blatantly dangerous. Sarah Bird, Microsoft\u2019s chief product officer for responsible AI, told me that Copilot, like many chatbots, runs a suite of tools for analyzing a user\u2019s identity and patterns of behavior. On the basis of this type of information, OpenAI\u2019s newest model, Astra, can activate more stringent refusals for individuals it deems \u201chigh risk.\u201d Ultimately the goal of systems like this is to look beyond the words of any given prompt and assess, instead, the user\u2019s intent.\u00a0<\/p>\n<p>Such tools might, in some cases, help indicate whether a person is looking for cyber vulnerabilities to exploit or to patch. But they would also help discern a user\u2019s political motives, not to mention offering an intrusive surveillance capability. (Bird acknowledged, in a follow-up email, that sophisticated refusal architectures create \u201ctrade-offs\u201d between safety and user privacy.)\u00a0<\/p>\n<p>Even the originators of refusal understood that such tight control over its cones and levers might not play to the favor of freedom and justice. \u201cTerms like helpful, honest, and harmless are ambiguous,\u201d the authors of the 2021 Anthropic paper explained. \u201cIt\u2019s easy to imagine them distorted beyond their original meaning, perhaps in intentionally Orwellian ways.\u201d\u00a0<\/p>\n<\/p><\/div>\n<p>#putting #faith #AIs #ability<\/p>\n","protected":false},"excerpt":{"rendered":"<p>If a trillion-dollar company doesn\u2019t want you to be mean to its computer, so be&#8230;<\/p>\n","protected":false},"author":1,"featured_media":23725,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[7],"tags":[5494,5685,10053,8264],"class_list":["post-23724","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-stories","tag-ability","tag-ais","tag-faith","tag-putting"],"featured_image_urls":{"full":["https:\/\/news678.top\/wp-content\/uploads\/2026\/10\/0924OP.jpg",1200,600,false],"thumbnail":["https:\/\/news678.top\/wp-content\/uploads\/2026\/10\/0924OP-150x150.jpg",150,150,true],"medium":["https:\/\/news678.top\/wp-content\/uploads\/2026\/10\/0924OP-300x150.jpg",300,150,true],"medium_large":["https:\/\/news678.top\/wp-content\/uploads\/2026\/10\/0924OP-768x384.jpg",640,320,true],"large":["https:\/\/news678.top\/wp-content\/uploads\/2026\/10\/0924OP-1024x512.jpg",640,320,true],"1536x1536":["https:\/\/news678.top\/wp-content\/uploads\/2026\/10\/0924OP.jpg",1200,600,false],"2048x2048":["https:\/\/news678.top\/wp-content\/uploads\/2026\/10\/0924OP.jpg",1200,600,false],"covernews-featured":["https:\/\/news678.top\/wp-content\/uploads\/2026\/10\/0924OP-1024x512.jpg",1024,512,true],"covernews-medium":["https:\/\/news678.top\/wp-content\/uploads\/2026\/10\/0924OP-540x340.jpg",540,340,true]},"author_info":{"display_name":"admin","author_link":"https:\/\/news678.top\/?author=1"},"category_info":"<a href=\"https:\/\/news678.top\/?cat=7\" rel=\"category\">Stories<\/a>","tag_info":"Stories","comment_count":"0","_links":{"self":[{"href":"https:\/\/news678.top\/index.php?rest_route=\/wp\/v2\/posts\/23724","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/news678.top\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/news678.top\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/news678.top\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/news678.top\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=23724"}],"version-history":[{"count":0,"href":"https:\/\/news678.top\/index.php?rest_route=\/wp\/v2\/posts\/23724\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/news678.top\/index.php?rest_route=\/wp\/v2\/media\/23725"}],"wp:attachment":[{"href":"https:\/\/news678.top\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=23724"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/news678.top\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=23724"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/news678.top\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=23724"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}