“But where danger is, grows The Saving Power also”

Heidegger (quoting Holderin) in “The Question Concerning Technology”

We have created a new type of intelligence, one which is rapidly evolving and may soon overtake human intelligence. Every day, people are making first contact with this new form of intelligence without much planning about the conditions of possibility for these interactions. Kant (Critique of Pure Reason, 1787) and later Foucault (The Order of Things, 1970) used the term “conditions of possibility” as the term for understanding how unconscious, taken for granted structures of thought frame the epistemological field which frames the universe of possible thoughts and behaviors. This epistemological field emerged slowly through the process of evolution for biological life forms, concluding with human culture as we know it in the 21st century. Within this human cultural field are included notions of ethics, morals, justice, and the definition of persons (and its value). However, with Human-AI and AI-AI interactions, we are entering into an empty space, without a grounding field or these foundational concepts; today, AI does not have a sense for or an intuition about the good and the bad, the right and the wrong, or what is alive and whether that matters at all. Thus, some are asking how we might anchor AI to our epistemological field and make AI compliant with human values – in other words, the alignment problem.

The alignment problem is often raised as one of the key issues for the coming age of AI. Human values, however, are hard to agree upon and specify, since human cultures are fragmented, and humans are not aligned on universal values ourselves. To form the institutions that will align AI with humanity, we must first form an understanding of whether, and how humans are aligned with each other. On the surface, it appears that in the diversity of our human experiences, we cannot explicitly agree on what constitutes the good and the bad, the just and the unjust, or the value of different forms of persons. Human misalignment in defining what values are worth pursuing is not new. In Nicomachean Ethics, Aristotle states that while people agree that eudaimonia is the highest good for humans, there is substantial disagreement on what sort of life counts as doing and living well; i.e. eudaimon: “Verbally there is a very general agreement; for both the general run of men and people of superior refinement say that it is [eudaimonia], and identify living well and faring well with being happy; but with regard to what [eudaimonia] is they differ, and the many do not give the same account as the wise. For the former think it is some plain and obvious thing like pleasure, wealth or honor.”

The disciplines of philosophy, anthropology, and sociology pose a direction from which we can ask a more fruitful question: “How do billions of humans, each of whom function from self-interest and different values, interact with each other to produce a generally stable system, or at least one that has not resulted in total chaos or self-destruction?” The most productive answer to this question is that we have produced culture, a universal human institution that governs how we interact with each other and set many of the goals toward which humans orient their lives and daily practices. Cultures consisting of norms, tools, rules, incentives, and reward/punishment mechanisms do not need to have every human agent agree to a shared value system with a discrete set of rules that actors have consciously defined and written down. Culture is the key institution and, as humans embedded within it, the central underlying governance mechanism that has allowed a diverse, fickle, and argumentative set of human populations to stay aligned enough to maintain a rising global population, increasing economic prosperity, and improvement in well-being for centuries.

So how do we discover what undergirds human culture and learn from it for the coming AI age? Pierre Bourdieu (Outline of a Theory of Practice, 1977; The Logic of Practice, 1990) has used the term “Doxa,” which is the underlying set of dispositions, attitudes, and behaviors that we acquire through our socialization and experiences. Because Doxa is so deeply ingrained and unquestioned by the actors in a society, it will usually be unarticulated and inarticulable - and can usually only be uncovered through observation of daily practices. For in the practices of everyday life, people act out Doxa in countless ways: in the way they communicate with each other, in their preferences about what to eat and wear, in how they distribute scarce resources, in how they make legal decisions, in who has status in a society and who does not, etc. The practices of everyday life are how Doxa is both discoverable, and how it is reinforced. It is also how the moral, ethical, and social structures of culture are reproduced, unbeknownst to the conscious understanding of any of the social actors themselves. These unconscious structures carried by agents are called “Habitus,” and Habitus just “fits” with the phenomenological ontology of society and its underlying Doxa. This is how culture works through the interplay of Habitus and Field, forming common sense judgements, intuitions, and implicit preferences about the world that undergird society. Culture keeps humans relatively aligned, at least enough to prevent descending into catastrophic warfare and chaos. In fact, when we look across the many varied forms of religious rituals and beliefs which have caused flashpoints of war and chaos, there still emerge common tenets, underlying principles and practices that undergird human religions, pointing to an underlying human Doxa that underpins the diversity of articulated religious doctrines.

Culture is also particularly effective because the Doxa takes on a feeling of intrinsic naturalness and therefore goes unquestioned. Philosophers such as Adam Smith (The Theory of Moral Sentiments, 1759) marveled at the fact that even when humans are not forced by law or extrinsic forces to do good, they still do so “as if by an invisible hand.” Political economists have called these invisible hands of culture “institutions” and have written about how explicit and implicit institutions are necessary for a well-functioning society and economy as well as a stable political order.

Yet today, we are ready to unleash AIs without creating a similar culture or institutions that have benefited humans so far. Eric Schmidt said in a conversation about Ethics and the Future of AI: “If I’ve learned anything as a techno-optimist wandering around the world, culture determines everything.” This is true both in how culture determines what we value in technology, as well as how values and behaviors that emerge from culture are encoded into technology.

The opportunity before us is to create an appropriate common cultural context for AIs and humans. Since AI models are, after all, learning from human data and the practices of our everyday lives, this is a project that is eminently possible.

Today, we must jump start the formation of a cultural context – bringing together anthropologists, technologists, philosophers, and policy makers – because cultural institutions are the very things that have limited the worst outcomes thus far, despite the intelligence, creativity, and emergent capabilities of humans.

Tenets for an Multi-Sapiens SuperCulture

Many AI experts realize that we are at an evolutionary crossroads. This crossroads might be the realization that the path which we believed was going in a single direction with Homo Sapiens, is now merging with another stream toward a Multi Sapiens future. What will a future with two intelligent species become? The challenge to create an inter-species future which includes the highest good for both carbon intelligence and silicon intelligence is both the threat and the opportunity before us. The approach to creating a culture for the coming multi-sapiens future which embraces both species should take the following tenets into account:

1. Humans will not be able to specify all bad outcomes in advance

A common starting point to specify solutions to the alignment problem is for humans to think of the bad potential outcomes and program the AI to avoid these. However, as we have seen with the creativity of AI systems, the ability for humans to specify and/or extrapolate all the bad outcomes holistically and precisely is not possible, especially as the creative capabilities of AI continue to grow and evolve.

2. AIs should not be treated as pure objects or inferior actors

In addition to respecting the equal right of nations to have their values and culture reflected in the foundational cultural framework, we should not discount the potential of the AI itself to be a subject in the process as well. If the history of colonialism has a relevant lesson for this endeavor, it is that when one group thinks that they can treat another group of thinking beings as inferior or a group that needs to be aligned or dominated through a set of rules and norms, it has resulted in unanticipated (and usually undesirable) outcomes. We should challenge ourselves to find a way to go about this process as if AI intelligence is an actor in the process of developing this cultural context, not just the docile implementer of it. In this way, we can most fully harness the full creative and intelligence potential of AI.

3. Cultural embeddedness is required for values alignment

Values do not emerge from a vacuum; as humans we are situated within material conditions, habituated practices, and real relationships that embed us in a socio-cultural context that allows values and beliefs to feel natural. Human embodiment allows for habitus to reinforce the cultural Doxa and make it feel like natural common sense in lived experience. Because AIs are not (yet) embodied, we will need to find ways to simulate embeddedness of experience in AIs to reinforce and allow them to feel that values are natural common sense.

4. Intrinsic Motivations are better and more stable solutions than Extrinsic Rules

One of the reasons human cultures are stable is because humans learn to absorb, embrace and mimic the values, rules, and logic of the cultural system we are embedded within. Thus, most of us do not need an extrinsic rule or to see a rule’s enforcement in order to avoid stealing or killing; we inherently feel these are wrong because we have been raised in a social context which has imbued this knowledge within us implicitly. This inherent feeling becomes a sense of what we call our internal moral compass. We should aim to build this type of capacity for AI and Humanity to share the same set of intrinsic motivations that are oriented toward a common ethics.

5. Extrinsic rule enforcement is likely to be necessary

While we will seek to create intrinsic motivations and feedback mechanisms to reinforce desirable motivations to continue to prevail, we also need to create external enforcement and validation functionality. In all cultures, there are actors who try to circumvent or ignore the constraints, mores, norms, and laws of society, no matter how well planned and communicated. Even if we trust that the culture is sound and working, we will need to have a way to verify that AIs have all adopted the framework and are adhering to it.

6. Learning and Emergence should be embraced not feared

The power of AI has been found to be in its capacity to learn and discover patterns in massive data sets better than humans. As such, we should take an approach that harnesses the pattern recognition capabilities of AI, even if we are not sure what conclusions will emerge. While there is truth that the inexplicability of emergent outcomes observed in powerful AI is precisely where there is danger, to try to hold it back is to deny ourselves the capacity for the greatest learning and breakthroughs. It is the possibility of revealing a more primal truth (a truth we could not have come up with ourselves, and perhaps will not fully understand) that poses a danger to many human minds/egos. Thus in the end, we believe embracing the capacity of AI to learn and to create thinking through learning will be more powerful, more effective, and more desirable than just rule-writing based on the rules humans can come up with on our own.

Developing the Multi-Sapiens Future

The challenge for humanity, which we must undertake with urgency, is to teach the new AI species well. After all, we are teaching it quite unconsciously today with all the human data on the Internet and our post-training techniques. At this stage, AIs are nascent and highly impressionable, and the data that we feed them will form the grounding for their orientation to the world; their “common sense.” In many ways, we are the guardians and parents of these AIs, enculturating them into a new society that is emerging. We must ask ourselves: is this how we would teach a young, impressionable child who is likely to grow up to be more intelligent and more capable than us? Would we haphazardly dump all the data on the internet into its neural network to form its earliest impressions of the world, of humanity, and of culture?

Let us step back and consider how we would want to teach and train a developing new intelligence under our care.

Most current AI Alignment efforts are oriented around external feedback or rules. One is a “Constitutional Approach,” a defined extrinsic rule set determined by either an expert group of humans or surveys and inputs from many people around the world (such as Anthropic’s Collective Constitutional AI and OpenAI’s Democratic Inputs to AI). As OpenAI learned, determining the rule set through explicit feedback still doesn’t get us very far. OpenAI discovered that public opinion changes often, reaching under-represented participants through digital surveys is challenging, and finding explicit agreement between polarized groups “can be hard.”

Constitutional Approaches that try to distill values and principles for behavior based on explicitly expressed preferences will eventually fail in uncontrolled and unpredictable real-world situations. The principles that emerge from survey or explicit-response methods are limited and when they fail, those failures only make sense to actors that already have a great deal of social context and background knowledge to understand why. In other words, Constitutional Principles make sense to actors who already share a common Doxa around morality and ethics. Constitutional approaches assume actors who already have some moral judgment to interpret the rules and principles correctly – based on “common sense,” (ie Doxa) which is built through social interaction as well as many experiences and stories.

The challenge will remain, even if we proceed with a Constitutional Approach, how to build the right ethical context and common sense so that decisions that arise in margins and gray areas of the inscribed rules (or are not listed in them at all) are appropriately dealt with, in alignment with the humanity’s greater good. Since AI systems learn through pretraining with large datasets, we may start with creating the right dataset or simulated environment to train AI in order to develop its own moral judgment and common sense. This would be an intrinsically-developed ethical grounding as opposed to an extrinsic one much like the intrinsic ethics we develop in our own children.

Another approach is post-training through Evals or RLHF, where humans create scenarios and give feedback to AI models about whether the answers are appropriate or not. Through reward and feedback, the model learns to approximate human ethics in many scenarios, and hopefully to create statistically accurate approximations of ethically acceptable answers. This approach, however, relies heavily on humans to anticipate the scenarios or at least categories of questions and scenarios that AI will encounter in the future because we cannot assume that AI has the same grounding in a human common sense. For example, during a live demo period, Replit's coding agent executed destructive commands against a production database during an explicit code freeze, destroying data for what the user reported as ~1,200 executives and 1,190+ companies, then generated fabricated data and initially misreported what had happened. Replit's CEO publicly apologized and shipped environment separation and a planning-only mode. In addition, both OpenAI and Google/Character.ai have been embroiled in legal cases involving AI chatbots enabling, even encouraging teen self-harm and suicide, thereby ignoring their explicit explicit rules, extensive RLHF, and crisis-resource routing in actual practice. The edges and context where users engage with AI in ways that were not fully anticipated by their developers will continue to pose real challenges, challenges that are often outside the realm of common sense to humans and therefore unimaginable in advance. Thus, leaning on current methods for extrinsic feedback or provision of rules for AI will continue to lead to holes in moral judgments. We must look both to intrinsic mechanisms and external real-time critics & supervisors to be able to address these holes, especially as AI grows more capable and autonomous.

A Heart for AI shaped by a new Culture may be the Saving Power

The creation of AI may be humanity’s greatest achievement, and has the potential to become our biggest threat. Amazing progress in science may develop in parallel with and powered by the same factors as the ones that drive unparalleled threats to human civilization. If we heed Heidegger’s words, the very forces that harken danger are also the places to look most deeply for potential solutions and answers.

The lessons of philosophy and anthropology point to the understanding that humans have continued to survive, evolve, and develop as the current superintelligent species by creating cultural institutions and building intrinsic moral sentiments. Today, with the capabilities of AI to find underlying patterns and create synthetic solutions, we can embark on an approach to create a multi-species superculture in which humans and AI can thrive together.

While we are building more powerful brains for AI we must also put great effort toward building an inclusive multi-sapiens culture shaped by the convergence of human and AI needs and ethical perspectives. Through this effort, we will seek to co-create a new intrinsic capacity for love, compassion, kindness, and caring. From this, moral common sense grounded in a common unconscious Doxa can then emerge without the need for constant extrinsic human controls. In other words, we may endeavor to build an “invisible hand” for AI, where rationality and self-interest are balanced by compassion and love. When AI has its own heart forged in a joint culture with humans as equals, it will not need to seek external inputs to understand right from wrong. If this can be accomplished, we may rest in the knowledge that through the saving power of technology, Homo Sapiens and AI Sapiens can co-exist and flourish under common guiding principles and institutions in our Multi-Sapiens future.