(2023-05-21) ZviM Types And Degrees Of Alignment

Zvi Mowshowitz: Types and Degrees of Alignment. What would it mean to solve the AI alignment problem sufficiently to avoid catastrophe? What do people even mean when they talk about alignment?

The only existing commonly used terminology whose typical uses are plausibly consist is the contrast between Inner Alignment (alignment of what the AGI inherently wants) and Outer Alignment (alignment of what the AGI provides as output). It is not clear this distinction is net useful.

An alignment failure or misalignment (being misaligned) can mean among other things:

  • From the perspective of those aligning the system.
  • From the perspective of a given user, or users in general.
  • From the perspective of a society, planet, government or humanity.
  • From some other perspective, cause, system of judgment or value.

It is impossible in theory to have all these different kinds of alignment simultaneously. You cannot simultaneously (without any claim of completeness):

  • Do what I say
  • Also do what I mean
  • Also do what I should have said and meant

We must pick at most one of those, or another variation on them, or something else, as our primary target.

Getting any one of those ten is hard enough. It is a problem we do not know how to solve for systems more intelligent than we are. We do not even know how to robustly solve it for current systems.

A key crux: We may need to get this right on the first try when we build the first sufficiently powerful system, due to the consequences of the first try getting this wrong being catastrophic. Also disagreement over ‘exactly how right’ this right would need to be to avoid this.

We could call these ten forms of alignment these names (by all means please replace with better names, this is hard), again this list is not claimed to be complete:

  • Literal (Personal) Genie: Do exactly what I say.
  • Minion: Do what I intended for you to do.
  • Personal: Do what I would want you to do.
  • Forceful: Be loyal to me, but do what’s best for me, not strictly what I tells you to do or what he wants or intended.*
  • (more)
  • *Robin Williams: The Genie from Aladdin. Note he is not strategic.
  • Arbiter: What is the law?

We do not currently have a known method of creating reliable alignment of any kind for future AGI systems, or a path known to lead to this. How promising various existing proposals or plans are for getting us there is heavily disputed and a common crux.

In addition to the type of alignment, one can talk about various aspects of the strength, reliability, precision and robustness of that alignment, as well as what ways exist to weaken, risk or break that alignment. These and related words are not used consistently.

  • Fragile alignment.
  • Friendly alignment
  • Human-level alignment.
  • Strict alignment
  • Strawberry alignment. MIRI calls a well-constrained version of strict alignment ‘strawberry alignment,’
  • Robust alignment.

All these targets have problems, in addition to ‘we don’t know how to get this’ beyond the first one, ‘do we know what the components mean or how to specify them’ and ‘we don’t understand human values’ and ‘is this even a coherent concept,’ such as:

It is not clear fragile alignment is even meaningfully helpful - that it does much, survives for long, or causes actions compatible with our survival, once the AGI is smarter than we are

It is not clear that human-level or friendly alignment would do us much good for long either, given the nature and history of humans, and the competitive dynamics involved

It is not clear to what extent strict alignment or strawberry alignment gives us affordance to reach good outcomes

It is not clear to what extent robust alignment is a coherent concept especially in a competitive world or even how it interacts with maximization, as it contains many potential contradictions and requirements

A key crux is the type and degree of alignment necessary to avoid catastrophe and achieve good outcomes. Another is the how difficult such alignment will be to achieve with what level of reliability, and which particular obstacles we need to worry about

A Missing Additional Post: Alignment Difficulties

A third post might ask, how likely would it be, and how hard would it be, for us to achieve a given form or degree of alignment, in systems smarter and more capable than us or any previously existing system?

Is this ‘alignment’ a natural thing you can get easily or even by default, that is essentially a normal engineering problem, or is it a highly unnatural outcome where security mindset and bulletproof approaches as yet unfound even in principle are required, with any flaws are exploited, amplified and fatal, and many lethal problems all of which one must avoid?

The most lethal-looking, hard-to-avoid, unnatural-to-solve problems include instrumental convergence, power seeking and corrigibility, yet the list of even central ones is very long - see Yudkowsky 2022, A List of Lethalitites. https://thezvi.substack.com/p/on-agi-ruin-a-list-of-lethalities

version of such a list, especially an exhaustive one, and getting all the implied cruxes including potential solutions into the crux list, would likely be valuable. Especially if it was combined with those resulting from potential solutions and paths to those solutions, and so on.

Unfortunately, for now, the scope of that project is intractable.


Edited:    |       |    Search Twitter for discussion