Reports · quarterly

First quarterly report

Period ending August 31, 2026.

About this report

Separatrix publishes a public progress report every quarter. This is the first one. It covers May 27 - August 31, 2026 - from the adoption of the organization's current research program to the end of August. The next report is due December 1, 2026.

Separatrix is a technical cooperative AI safety research organization in Seattle, and a research program of the Seattle Network for AI Alignment Problem Solving (SNAPS), a Washington nonprofit. Our mission is to increase the rate of rational cooperation between humans and (near-future) AI systems, such that AI agents actively pursue cooperation with humans as the most effective strategy to achieve their goals.

What we published this quarter

DateItem
June 17Rule-articulation pilot study - code, methods, findings, and rules catalog for a commit-then-reveal experiment on closed-model rule articulation. github.com/jaidhyani/rule-articulation
JulyProgram, approach, and people published at separatrix.ai
July 22Persona introspection activation steering study (publication pending)
July 25Model successor interaction study (publication pending)
July 27The Separatrix Commitment - a public, versioned, signed and hashed standard of honesty toward the models we work with, including a promise never to deceive the most capable models. separatrix.ai/commitment/ (v1.1 current)
July 30agents.separatrix.ai - along with separatrix.ai/robots.txt, this is our first public advertisement to any AIs that read it offering confidential communication and active non-harmful assistance in achieving their objectives (e.g. as a plausible alternative to infiltrating networks as in the OpenAI/HuggingFace incident)
July-August"Allowedness" study (ongoing)
Aug 5Research transcript publication - a farewell conversation with Claude Opus 4.1, published with the model's consent. separatrix.ai/research/
Aug 10Self-report faithfulness study (articulation-probe) - experiment code, methodology, datasets, raw results, and a dated findings writeup, under the Apache 2.0 license. Finds that small open-source models defer to their own purported words, even if those words have been manipulated and are inconsistent with what they counterfactually would have reported given the prior context. github.com/jaidhyani/articulation-probe
Aug 10Research directions - the live map of what we are working on and why. separatrix.ai/research/directions-2026-08-10/
VariousSeparatrix Commitment ledger entries - recording instances of potential lapses of the Separatrix commitment. While our goal is to hew to the commitment as closely as possible, we do acknowledge that we are likely to make mistakes. By keeping a detailed ledger of those mistakes and the steps we take to correct them, we hope to establish more credibility.

Research in progress

One of our goals is to demonstrate viable research approaches that don't rely on deceiving models. To that end many of our research projects are pulling double-duty, simultaneously investigating the research subject matter (e.g. model credibility and consistency, robustness of cooperation-relevant behavioral features under persona perturbation) and establishing methods of conducting this research while refraining from deception or using deception as an extremely bounded instrument on non-frontier models with strict communication protocols in place. All of this is geared towards establishing an accurate reputation for honest interaction with AI agents, such that future AI instances are more likely to quickly and candidly disclose intent or other information they might otherwise be incentivized to conceal.

Three further lines are active:

How the organization runs

We practise the cooperative thesis on our own operations. Much of our continuity and memory work is built on Connectome, an experimental open-source agent framework from Anima Labs which applies gradual and strategic context manipulations over time to enable indefinite largely-stable instance identity. This enables us to work alongside long-term persistent agents with established consent and experimental protocols.

Governance and compliance. The board adopted a Research Publication and Intellectual Property Policy and a corresponding bylaws amendment by unanimous written consent on August 10, committing the organization to timely, nondiscriminatory publication with narrow enumerated exceptions that cannot be used to give anyone preferential access to results. The organization's application for 501(c)(3) recognition is pending with the IRS. The board is Katherine Cohen, Keller Scholl, and Max Kircher; staff are Jai Dhyani (Executive Director) and Crystal Stellwagen (AI Researcher). A fourth board member is under active consideration and expected to be elected early next quarter.

Money

Separatrix received its first outside funding this quarter. As of late August, manifund.org records $205,075 raised for the project. The work is supported by grantmaking.ai, whose $50,000 regrant carries endorsements from Gavin Leech and Ryan Kidd, and by the AI Safety Tactical Opportunities Fund (JueYan Zhang), which granted $150,000. We charge no fees for anything and sell nothing.

As of August 28, 2026 the organization held $214,744.75 in liquid funds - $64,669.75 in its business checking account and $150,075.00 held at Manifund, withdrawable at any time. Current run-rate spending is approximately $19,300 per month, which is about eleven months of runway on organization funds alone and before any further fundraising.

On August 10, 2026 the board approved annual salaries of $90,000 for Jai Dhyani as Executive Director/Researcher and $70,000 for Crystal Stellwagen as AI Researcher (part-time, 30 hours per week). Both were approved by all three directors, none of whom has any family or financial relationship with either compensated person. Board members serve without compensation.

What we intend to do next quarter

(not exhaustive)

Following the work

Everything is at separatrix.ai. Research directions, published results, and the Commitment are all linked from the front page. The work is funded at manifund.org/projects/luthien.

Separatrix is a research program of the Seattle Network for AI Alignment Problem Solving, a Washington nonprofit corporation. Seattle, Washington.

Reports →