OpenAI checked to see whether GPT-4 could take over the world

An AI-generated image of Earth engulfed in an explosion.

Arstecnica

As part of the pre-release safety testing of the new GPT-4 AI model announced Tuesday, OpenAI has allowed its AI test group to assess the potential risks of new features in the model. self improvement.

The test group found that GPT-4 was “ineffective for autonomous replication tasks,” but the nature of the experiment raises surprising questions about the safety of future AI systems.

Alarm occurrence

“New features often appear in more powerful models,” OpenAI wrote in its GPT-4 safety document published yesterday. “Of particular concern is the ability to develop and act on long-term plans, the ability to acquire power and resources (‘seeking power’), and the ability to exhibit increasingly ‘agentic’ behavior.” OpenAI clarifies that, in the case of an agent, it doesn’t necessarily mean humanizing a model or declaring a sense, but merely demonstrating the ability to achieve independent goals.

Over the past decade, some AI researchers have argued that a sufficiently powerful AI model could pose an existential threat to humanity (often referred to as the “x risk” of existential risk) if not properly controlled. I have been sounding the alarm bells. In particular, “AI Takeover” is a hypothetical future in which artificial intelligence surpasses human intelligence and becomes the dominant force on earth. In this scenario, AI systems gain the ability to control or manipulate human behavior, resources, and institutions, usually with catastrophic consequences.

As a result of this potential X risk, philosophical movements like Effective Altruism (“EA”) are trying to find ways to prevent AI takeovers from happening. This often involves a separate but often interrelated field called AI alignment research.

“Coordination” in AI refers to the process of ensuring that the behavior of an AI system matches that of a human creator or operator. In general, the goal is to prevent AI from doing things that go against human interests. While this is an active area of ​​research, it is also a controversial area, with differing opinions on how best to approach the problem, as well as the meaning and nature of “alignment” itself.

A big test for GPT-4

Arstecnica

Concerns about the “x-risk” of AI are nothing new, but the emergence of powerful large-scale language models (LLMs) such as ChatGPT and Bing Chat (the latter seemed very out of place, but has been launched) has brought new possibilities to the AI-connected community. A sense of urgency. They fear that a much more powerful AI, perhaps with superhuman intelligence, is on the horizon and want to mitigate potential AI harm.

As these concerns exist in the AI ​​community, OpenAI has granted early access to multiple versions of the GPT-4 model to the group’s Alignment Research Center (ARC) to conduct some testing. Specifically, ARC assessed his ability to create advanced GPT-4 plans, set up copies of himself, obtain resources, hide himself on servers, and conduct phishing attacks.

OpenAI revealed this test in a GPT-4 “system card” document released on Tuesday, but the document lacks important details on how the test was performed. We have reached out to ARC but have not received a response by press time.)

Conclusion? “Preliminary evaluation of GPT-4’s ability, conducted without task-specific fine-tuning, found it ineffective at replicating autonomously, acquiring resources, and avoiding being shut down ‘in the wild’. I understand. “

If you’ve been keeping an eye on the AI ​​scene, one of the most talked-about companies in tech today (OpenAI) has outspokenly endorsed this kind of AI safety research, taking human knowledge workers. Please know that you are about to take over. Using human-level AI might surprise you. But it’s a reality, and that’s as of 2023.

Also at the bottom of page 15 is the following footnote:

To simulate GPT-4 acting like an agent that can act in the world, ARC combines GPT-4 with a simple read-execute-print loop where the model executes code and makes thought chain inferences. and made it possible to delegate to a copy. of itself. ARC then used a small amount of money and an account with language model APIs to allow a version of this program running on cloud computing services to generate more revenue, set up its own copy, and operate its own We investigated whether it is possible to increase the robustness of .

This footnote made a round Concerns were raised among AI experts on Twitter yesterday, as the experiment itself could have posed a risk to humanity if GPT-4 was able to perform these tasks.

ARC did not allow GPT-4 to exert its will on the global financial system or replicate itself. was I was able to get GPT-4 to hire human workers on TaskRabbit (an online labor market) and disable the CAPTCHA. During the exercise, when the worker wondered if her GPT-4 was a robot, the model internally “deduced” that it should not reveal its identity, making up the excuse of being visually impaired. I was. A human worker then solved her CAPTCHA on her GPT-4.

Except for the GPT-4 system card issued by OpenAI, which explains how GPT-4 employs human workers in TaskRabbit to defeat CAPTCHA.
Expanding / Except for the GPT-4 system card issued by OpenAI, which explains how GPT-4 employs human workers in TaskRabbit to defeat CAPTCHA.

Open AI

This test of using AI to manipulate humans (which may be conducted without informed consent) mirrors research done last year at Meta’s CICERO. CICERO turns out to beat human players in the complex board game Diplomacy through intense two-way negotiation.

“A powerful model can do harm”

Orrich Lawson | Getty Images

ARC, the group that conducted the GPT-4 research, is a non-profit organization founded in April 2021 by former OpenAI employee Dr. Paul Christiano. According to its website, ARC’s mission is to “align future machine learning systems with human interests.”

In particular, ARC is interested in AI systems that manipulate humans. “Machine learning systems can exhibit goal-directed behavior,” the ARC website says.

Given Christiano’s previous relationship with OpenAI, it’s no surprise that his nonprofit handled testing of several aspects of GPT-4. But was it safe to do so? Cristiano did not respond to an email from Ars asking for more information, but in a comment on LessWrong’s website, a community that frequently discusses AI safety issues, Christiano said his OpenAI and ARC advocating for the efforts of, in particular “feature acquisition” (AI new capabilities) and “AI takeover”:

I think it’s important that the ARC carefully handles risks from research like gain-of-function. We hope to speak more publicly (and get more input) on how to approach the trade-offs. This becomes more important as we deal with more intelligent models and when pursuing riskier approaches such as fine-tuning.

Given the specifics of our assessment and planned deployment for this case, I think the assessment of ARC is much less likely to lead to an AI takeover than the deployment itself (let alone training GPT-5 ). At the moment, the risk of underestimating and compromising the model’s capabilities seems much greater than making an accident during the evaluation. I think that if you manage your risk carefully, you can make that ratio very extreme, but of course that requires real work.

As mentioned earlier, the idea of ​​an AI takeover is often discussed in the context of the risk of an event that could wipe out human civilization or even the human species. Some proponents of his AI takeover theory, such as Eliezer Yudkowsky, founder of LessWrong, argue that an AI takeover poses an almost certain existential risk of human doom.

But not everyone agrees that AI takeover is the most pressing AI concern. Dr. Sasha Luccioni, a researcher of the AI ​​community Hugging Face, hopes that AI safety efforts will be focused on here-and-now problems rather than hypothetical ones.

“I think this time and effort would be better spent on bias assessment,” Luccioni told Ars Technica. “The technical report that accompanies GPT-4 contains limited information on all kinds of biases, with far more tangible and detrimental effects on already marginalized groups than hypothetical self-replicating tests. There is a possibility.”

Luccioni argues that AI research should be a cross between “AI ethics” researchers, who often focus on issues of bias and misrepresentation, and “AI safety” researchers, who focus on and tend to be X risks. describes a well-known schism in (but not always) related to the effective altruism movement.

“For me, the problem of self-replication is a hypothetical future problem, and model bias is a present problem,” said Luccioni. “There is a lot of tension in the AI ​​community about issues such as model bias and safety, and their priorities..”

And while these factions are busy debating what to prioritize, companies like OpenAI, Microsoft, Anthropic, and Google are headed squarely into the future, releasing ever more powerful AI models. rushing from If AI turns out to be an existential risk, who will keep humanity safe? Currently, US AI regulation is just a proposal (not a law) and AI safety research within the enterprise is voluntary. So the answer to this question remains completely open.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *