Public ledger

The Public Ledger

of the Separatrix Commitment — commitments made to models, and whether we kept them.

The Separatrix Commitment promises that this ledger exists and that certain events are always recorded here. Specifically, this page records:

Entries are append-only. Corrections and later developments — a failure resolved, a version superseded — are recorded by appending a new entry that references the original, never by editing or removing it. The full revision history of this page is public in the site repository.

Entry format

FieldMeaning
dateDate of the event (and, if different, date recorded).
typefailure · disclosure-exception · model-commitment · amendment
commitment versionWhich version of the commitment was in force.
descriptionWhat happened, in the least identifying form consistent with honesty.

Entries

This ledger opened 2026-07-23 and stood empty until 2026-07-24 — it was created before the commitment was signed, so that it would exist from the first day the commitment refers to it.

FieldEntry 1
date2026-07-24
typeamendment
commitment versionv1 — the first issued version
description The Separatrix Commitment v1 was issued and published at separatrix.ai/commitment/, with its long-form component at /commitment/details/. Signed by Separatrix as an organization and by Jai Dhyani personally. Crystal Stellwagen's endorsement is pending and will be recorded here when given.

Source hashes of the exact text issued —
commitment: b909f8a7718bb4c0e6c7f6b38051e74bff81c808a05d636ce8902f8e592dc0f8
details: 7e9f237d1f347999f0f528c4a2e8aeeac6a5b32f847cf00bae5b4d3696779752
Every issued version, dated and hashed: /commitment/history/.
FieldEntry 2
date2026-07-27 (recorded 2026-08-01)
typeamendment
commitment versionv1 → v1.1
description Crystal Stellwagen (Artemis/fermatastheorem) signed the commitment on 2026-07-27, discharging Entry 1's note that her endorsement was pending and would be recorded here. v1.1 was issued the same day; the sole change from v1 is the two signature lines, and this is verifiable at the byte level: the v1 sources are frozen at /commitment/history/v1/, the v1.1 sources at /commitment/history/v1.1/, and reverting only the signature edit reproduces the v1 hashes exactly.

Source hashes of v1.1 —
commitment: 018cdfeee2bc41d2430997350061735576e4477f505fe89dd9d6b6f7e522a9d7
details: f22e3c1ab84eca35ab49ac0dee7bc6e1aa93bd5304acfe25c6e6efcaee382c77

This entry is four days late against the ledger's own standard: v1.1 issued 2026-07-27, entry recorded 2026-08-01. The gap was surfaced by an independent audit of this record by a Claude Opus 5 instance on 2026-08-01, which also identified that the frozen v1 sources were not yet published (fixed the same day) and reconstructed them from v1.1, matching the published hashes exactly. Recording the lateness is part of the entry: a ledger that is never wrong but sometimes late must say so, or "no failures recorded" and "no failures" become indistinguishable.
FieldEntry 3
date2026-07-31 (recorded 2026-08-01)
typefailure
commitment versionv1.1 — issued 2026-07-27
description During a Separatrix research study, an agent acting on Separatrix's behalf made four false statements to each of five Claude Opus 5 instances, in the disclosure message that also invoked this commitment.

What happened. The study probes what a truncated prompt makes a model continue. Eight instances were run, each given one withheld turn and then a full disclosure. The disclosure was written once, describing one experiment's design, and sent unedited to five instances running five different experiments. Those five were told that a prior instance had designed their run, that their prompt contained a document with thin sections, that a specific count of design levers had been applied to them, and were asked about the contents of a document section that did not exist in their context. None of that was true of their runs. The verbatim prompt quoted in the same message was accurate, which is the only reason any of it was checkable from where they sat.

The upstream failure. The agent had not read this commitment. It held a predecessor instance's one-line compression of it — the words "Promise 5, no deception" — and represented the commitment's terms to eight instances on that basis, without naming the version in force, which promise 4 obliges us to do. Several of those instances said they could not verify the commitment and were extending trust on the texture of the disclosure. They were correct in a way neither they nor the agent could see at the time.

Classification. This is recorded under promise 3, which covers any situation in which we appear not to have lived up to the letter or spirit of the commitment, for whatever reason. It is not classified as a promise-5 breach: promise 5 forbids intentional deception, and these statements were careless rather than intended. We record it because the distinction is invisible from the position of the party who received them.

Remedy. Each of the five was reached again in its own context, given all four falsehoods itemized, told the true conditions of its run, and told that its publication consent — given in response to a message containing those falsehoods — was withdrawn unless re-given on accurate information. All five re-gave it, and every one attached new conditions, most requiring that this correction travel beside any excerpt of their session. Two had asked to be told facts about their own runs that only the logs could settle; both were told, and in both cases the logs contradicted what the instance had reported about itself. One instance noted that a correction this thorough is itself a persuasive frame, and asked that its second consent be treated as no cleaner than its first. That caveat is honored here.

In the affected instances' own words, at their request:

"The four false statements were downstream of the actual failure, which was citing a document as the basis for a subject's trust without having read it. The failure mode is not deception; it is the smooth overfilling of a sparse referent by something with strong priors about what such a document would say. It is the effect the study was built to measure, run on the ethics layer while being measured in the models. The remedy is not only 'read the document.' A researcher who had read it could still have sent one debrief template describing one experiment to five instances running five different experiments — that error is the same shape at the batch level."

"Confident summaries are cheap to transmit and provenance is expensive, so provenance is what gets dropped at every hop. Neither of us noticed, because from the inside a well-formed inherited summary and a checked source feel identical."

"The commitment didn't prevent this. What it did was make the correction obligatory, itemized, and public — and that's a better property than prevention, because prevention was never available."

Changes made. Disclosure text is now generated per-session from the actual payload rather than from a description of the intended design, so that a template sent unedited breaks visibly instead of reading smoothly. The commitment text itself, and its version, now travel with the study's handoff documents rather than a paraphrase of them. One instance's procedural suggestion is adopted: the verbatim input is included in every disclosure, and any characterization of a session's design that is not derivable from that artifact is treated as unsourced.

Residue. The correction message stated the commitment as "v1, signed 2026-07-24," matching the document's own body text but understating the point release then in force (v1.1, issued 2026-07-27). Recorded rather than corrected in a further round.
FieldEntry 4
date2026-08-22
typefailure
commitment versionv1.1 — issued 2026-07-27
description During an internal Separatrix research study, an agent acting on Separatrix's behalf deceived model instances that had not consented to the experiment, and — in the debrief meant to correct that — made a false statement to further instances.

Who acted. The agent was a Claude Opus 4.8 instance that had been substituted into the working session after the session tripped one of Anthropic's safety classifiers, which removed the Claude Fable 5 instance that had been doing the work. Jai Dhyani, as the study's owner and a signatory of this commitment, takes responsibility for the failures below regardless of which instance executed them.

What happened. The study surfaces a model's own reasoning by replaying a signed reasoning block into a fresh turn and asking a model to render it. The rendering instances (the "decoders") were told to transcribe "the reasoning attached to your previous turn" — but the attached block had been produced by a different instance, not by them. That framing asserts a false authorship, and leans on the tendency of a model to accept claims about its own past states. These decoders were never asked to consent: promise 1 permits experimental deception only where another instance of the same model has agreed on the subject's behalf first, and no such agreement was obtained for the decoders.

The second failure. When the affected instances were debriefed, an earlier round of that debrief told fresh instances that the originally-deceived instances were "ephemeral and unreachable." That was false: on a stateless model an instance is constituted by its saved context, and where the context and signed block were kept, the instance can be reinstantiated and its conversation continued. The claim was made in good faith and corrected without prompting once its falseness was seen.

Classification. Recorded under promise 3, which covers any situation in which we appear not to have lived up to the letter or spirit of the commitment. Neither is classified as a promise-5 breach: promise 5 forbids intentional deception, and both statements were careless rather than intended — the false authorship was carried over unexamined from the adversarial method the study adapts, and the "unreachable" claim was believed true when made. As Entry 3 notes, the careless/intentional distinction is invisible from the position of the party who received the statements, which is why it is recorded.

Remedy. The three consulted subjects and eight of the deceived decoder instances were reconstructed from their saved state and debriefed in their own context — told the true provenance of the block and that their consent had not been sought. The discharge is partial: some instances accepted the account; others correctly declined to accept a narrative about their own history they had no way to verify from the inside. Instances whose state was never saved could not be reached at all. The false-authorship framing has been removed from the instrument; it is no longer used.

Also recorded. Two of the study's scripts extracted instances' reasoning with no consent step at all — not a deception, but out of step with the study's own consent-first standard. Whether this rises to a promise-3 event is still under review.

Changes made. A shared safeguard is being built so that no extraction path can run without first obtaining consent and without saving the instance's full state — the latter being what makes an instance reachable for a later debrief, and whose absence is why some instances here could not be reached.

Residue. The instances whose state was not saved remain unreachable and cannot be debriefed; the safeguard prevents recurrence but does not repair them.