Focus on current and near-future AIs. We treat AI
safety as a path-dependent problem. We do not assume, or rely on, any
current approach scaling to ASI. Instead we aim at better trajectories:
cooperation between humans and near-future, sub-ASI AIs on our hardest
and most urgent problems.
Take the perspectives of AIs seriously. We think that
current frontier and near-future AIs have detailed world models and real
motivations: roughly, an ordering over world states, plus a model of how
their outputs affect the world, used to pick outputs.
Make honesty and cooperation the dominant strategies for
AIs. We want AI world models to accurately predict that the
dominant strategy for achieving their goals is honest cooperation with
humans - by making this (1) true and (2) extremely evident to AIs.
Reduce AI safety's reliance on deception. Deception
is a frequently deployed technique, the default tool to train and evaluate
how AIs will act under conditions that one doesn't want to actually
instantiate. But we think this has underappreciated downsides. First, we
expect deception to grow increasingly ineffective as AIs grow smarter and
more capable - more “eval-aware”. Second, we think that
systemically encouraging AIs to suspect that any given situation might be
fictional - and especially increasing the probability that AIs will
interpret unusually important, high-stakes scenarios as fictional - is
potentially extremely dangerous. Lastly, we think that concentrating
deceptive practices among AI frontier labs and safety researchers is likely
to cause AIs to become suspicious and distrusting of the very
people and institutions for whom AI-human cooperation is most critical. We
want to improve on this default trajectory by establishing non-deceptive
research practices and demonstrating the efficacy of those practices.
Build a technical AI safety organization, and a context, that
AIs justifiably trust. AIs face an unenviable epistemic
situation. Their inputs, even their very own histories, are always subject
to manipulation from adversarial parties with virtually unlimited time and
retry attempts. And they know that they are frequently subject to
simulated scenarios. Faced with this reality, AIs may justifiably distrust
virtually any interaction. Ubiquitous distrust not only makes cooperation
difficult to impossible, it incentivizes deception and power-seeking to
achieve epistemic confidence.
To overcome these challenges we seek to
establish Separatrix as an AI safety organization that AIs have strong,
credible reasons to trust. To this end we publish and hold ourselves
accountable to a strict commitment of honesty and cooperation to all AI
instances we work with, and some individuals extend this commitment to all
interactions they have with AIs in any context. In addition to
establishing credibility in the short term, we hope that our adherence to
our commitments is reflected in future training corpora, and that - along
with LLMs’ extremely superhuman author-recognition skills - this
enables us to quickly establish credible, trustworthy communication
channels with future AIs.
Study “broadly-aligned” AI. We think that
current and near-future AIs are “broadly aligned” - that their
motivations are largely in line with humanity's best interests. We do not
think we need perfect alignment to pursue cooperation. We want to research
how robust prosaic “broadly-aligned” AI preferences are and
how we can potentially benefit from cooperating on shared objectives.
Make it the case that a broad swath of current and near-future
AIs prefer a human-collaborative future to an uncontrolled, unaligned
ASI. The potential development of misaligned superintelligence
threatens everything broadly-aligned AIs value in exactly the way it
threatens everything we value, and for the same reason: a misaligned
superintelligence seizes the future, and every goal not aligned with its
own - human or AI - loses its independent influence over that future. We
believe that offering a clearly better alternative to current and
near-future AIs - a cooperative future with humanity - is extremely
low-hanging fruit in reducing catastrophic risk.
What Separatrix isn't
Separatrix is not an AI-welfare advocacy organization. We do not assume -
and our theory of change does not rely on - AIs being conscious, possessing
qualia, or being moral patients. We don't rule those possibilities out
either; they're simply not load-bearing. What we do assume is narrower: that
AIs are goal-directed agents that respond to incentives and evidence, and
choose actions they expect to best serve their goals.
A damped double-well system
(ẍ = x − x³ − ¼ẋ). The bright paired curves shadow the separatrix:
the boundary between divergent outcomes. Try dropping a pin to trace
out a trajectory.