Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Seán Ó hÉigeartaigh

Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Apr 15, 2024

Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut(+28 more)

Figure 1 for Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Figure 2 for Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Figure 3 for Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Figure 4 for Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Abstract:This work identifies 18 foundational challenges in assuring the alignment and safety of large language models (LLMs). These challenges are organized into three different categories: scientific understanding of LLMs, development and deployment methods, and sociotechnical challenges. Based on the identified challenges, we pose $200+$ concrete research questions.

Via

Access Paper or Ask Questions

Predictable Artificial Intelligence

Oct 09, 2023

Lexin Zhou, Pablo A. Moreno-Casares, Fernando Martínez-Plumed, John Burden, Ryan Burnell, Lucy Cheke, Cèsar Ferri, Alexandru Marcoci, Behzad Mehrbakhsh, Yael Moros-Daval(+5 more)

Figure 1 for Predictable Artificial Intelligence

Figure 2 for Predictable Artificial Intelligence

Figure 3 for Predictable Artificial Intelligence

Figure 4 for Predictable Artificial Intelligence

Abstract:We introduce the fundamental ideas and challenges of Predictable AI, a nascent research area that explores the ways in which we can anticipate key indicators of present and future AI ecosystems. We argue that achieving predictability is crucial for fostering trust, liability, control, alignment and safety of AI ecosystems, and thus should be prioritised over performance. While distinctive from other areas of technical and non-technical AI research, the questions, hypotheses and challenges relevant to Predictable AI were yet to be clearly described. This paper aims to elucidate them, calls for identifying paths towards AI predictability and outlines the potential impact of this emergent field.

* 11 pages excluding references, 4 figures, and 2 tables. Paper Under Review

Via

Access Paper or Ask Questions

AI Systems of Concern

Oct 09, 2023

Kayla Matteucci, Shahar Avin, Fazl Barez, Seán Ó hÉigeartaigh

Abstract:Concerns around future dangers from advanced AI often centre on systems hypothesised to have intrinsic characteristics such as agent-like behaviour, strategic awareness, and long-range planning. We label this cluster of characteristics as "Property X". Most present AI systems are low in "Property X"; however, in the absence of deliberate steering, current research directions may rapidly lead to the emergence of highly capable AI systems that are also high in "Property X". We argue that "Property X" characteristics are intrinsically dangerous, and when combined with greater capabilities will result in AI systems for which safety and control is difficult to guarantee. Drawing on several scholars' alternative frameworks for possible AI research trajectories, we argue that most of the proposed benefits of advanced AI can be obtained by systems designed to minimise this property. We then propose indicators and governance interventions to identify and limit the development of systems with risky "Property X" characteristics.

* 9 pages, 1 figure, 2 tables

Via

Access Paper or Ask Questions

International Governance of Civilian AI: A Jurisdictional Certification Approach

Sep 11, 2023

Robert Trager, Ben Harack, Anka Reuel, Allison Carnegie, Lennart Heim, Lewis Ho, Sarah Kreps, Ranjit Lall, Owen Larter, Seán Ó hÉigeartaigh(+2 more)

Figure 1 for International Governance of Civilian AI: A Jurisdictional Certification Approach

Figure 2 for International Governance of Civilian AI: A Jurisdictional Certification Approach

Figure 3 for International Governance of Civilian AI: A Jurisdictional Certification Approach

Figure 4 for International Governance of Civilian AI: A Jurisdictional Certification Approach

Abstract:This report describes trade-offs in the design of international governance arrangements for civilian artificial intelligence (AI) and presents one approach in detail. This approach represents the extension of a standards, licensing, and liability regime to the global level. We propose that states establish an International AI Organization (IAIO) to certify state jurisdictions (not firms or AI projects) for compliance with international oversight standards. States can give force to these international standards by adopting regulations prohibiting the import of goods whose supply chains embody AI from non-IAIO-certified jurisdictions. This borrows attributes from models of existing international organizations, such as the International Civilian Aviation Organization (ICAO), the International Maritime Organization (IMO), and the Financial Action Task Force (FATF). States can also adopt multilateral controls on the export of AI product inputs, such as specialized hardware, to non-certified jurisdictions. Indeed, both the import and export standards could be required for certification. As international actors reach consensus on risks of and minimum standards for advanced AI, a jurisdictional certification regime could mitigate a broad range of potential harms, including threats to public safety.

Via

Access Paper or Ask Questions

The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation

Feb 20, 2018

Miles Brundage, Shahar Avin, Jack Clark, Helen Toner, Peter Eckersley, Ben Garfinkel, Allan Dafoe, Paul Scharre, Thomas Zeitzoff, Bobby Filar(+16 more)

Figure 1 for The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation

Abstract:This report surveys the landscape of potential security threats from malicious uses of AI, and proposes ways to better forecast, prevent, and mitigate these threats. After analyzing the ways in which AI may influence the threat landscape in the digital, physical, and political domains, we make four high-level recommendations for AI researchers and other stakeholders. We also suggest several promising areas for further research that could expand the portfolio of defenses, or make attacks less effective or harder to execute. Finally, we discuss, but do not conclusively resolve, the long-term equilibrium of attackers and defenders.

Via

Access Paper or Ask Questions