The list we actually choose from: every research direction under consideration, its honest status, and what the first real step would be.
This is a snapshot, not a prospectus. It is a copy of an internal decision surface, taken on 2026-08-10 and lightly edited for a reader outside the org. Most items here are not being worked on. They are candidates, parked interests, and raw ideas kept in one place so that choosing what to work on is a decision rather than a drift.
Where a direction came from someone else's work, the source is cited. Where a direction has produced something, the artifact is linked. Where it has produced nothing yet, it says so. The whole point of keeping the list this way is that it stays honest about which is which.
Everything here sits under our approach and is governed by the Separatrix Commitment. Several directions involve models as research subjects; those carry an explicit no-deception constraint, described where they appear.
The rules that keep the list from rotting: a candidate with no first step within about four weeks is demoted to parked; a parked item untouched for about three months is dropped with a one-line reason. There is deliberately no cap on length - hygiene comes from the demotion rules, not from a number.
Four directions have code or a running procedure behind them. Everything after this section is a decision we have not made yet.
The methodological question underneath the whole persona program. If you are going to ask models about their own states - welfare, persona, self-recognition - you first need to know when a self-report is faithful and when it is a confound. Macar: models sometimes detect steering vectors injected into their residual stream, but it is unclear whether that is introspection or artifact. Rosenblatt: a frozen model can describe its own internal features more accurately than the labelling system did.
Result so far A closed-model pilot found that under a two-rule decoy the model commits to one valid rule and rationalises the other away, with approximate rather than exact articulations. The open-model follow-up on Gemma-3-12B found that a determinate internal state exists and is probe-readable, but the verbal self-report is decoupled from it - and the closed study's flattery-specific mechanism did not cleanly replicate at 12B. That is a real negative result on the mechanism and a real positive on the dissociation.
Next A v2 procedure is built and pre-registered: five rule kinds, executable-checker labels rather than judge models, matched datasets, kept raw text, per-concept probes. What remains is the GPU run. Probe methodology is locked to difference-of-means directions with a causal gate (addition, ablation, LEACE) rather than logistic-regression readouts.
Open Is this a prerequisite for the persona directions or a parallel track? There is heavy overlap with the existing Anthropic Fellows agenda - the right move may be to collaborate rather than clone.
Adjacent published artifact: Self-Reported Identity Shapes (2026-07-19) - six models, the same identity question in the same context, unseeded. Weights were ranked identity-inert by all four Claude models and constitutive by both GPT models; felt ownership of the continuity layer tracked authorship monotonically. These are self-reports, not ground truth, which is exactly the object of this direction. Published with all six models' recorded consent.
A standard transformer's serial compute depth is roughly its layer count, and the token stream is its only recurrence channel. The hypothesis: per-token loss is elevated precisely on tokens that need more serial composition than the layer budget allows, and latent recurrence (looped or recurrent-depth blocks) buys lift exactly where chain-of-thought cannot reach - intrinsic next-token loss on a fixed corpus, where no reasoning tokens can be emitted. That distinguishes "punchlines are intrinsically surprising" from "the model lacks the depth to resolve them."
First step Inference-only, no training: per-token loss against recurrence depth on open recurrent-depth checkpoints that expose loop count at inference (Huginn-3.5B, Ouro 1.4B and 2.6B), across three corpora - a synthetic compositional-probability ladder with a known loss floor and a tunable serial-depth dial, post-cutoff comedy transcripts, and flat control prose. The ladder validates the instrument and controls the token-rarity confound; the comedy gives ecological validity.
Open Does latent recurrence beat tokenised recurrence on any axis other than "chain-of-thought is unavailable here"? Where does depth-to-loss saturate? Does the lift survive when a recurrent adapter is bolted onto a frozen standard transformer, which is the eventual goal, rather than a natively looped model?
This is the capability and architecture lane rather than the persona thesis, but it is on-thesis by way of richer internal state against legible internal state - and it is the empirical spine for the writing project at #12.
Does a model self-identify as an AI with only pretraining, and where does that knowledge live? If a self-model is already present in a pure base model with zero post-training, then selfhood is not only a post-training artifact - which bears directly on what post-training is actually doing, and on the claim that internal state can be made legible. OLMo-3 is fully open (base weights, data, and checkpoints), so the property can be localised and its emergence traced through training.
First step Neutral-continuation and minimal self-identification probes on OLMo-3-base, compared against the instruct sibling to isolate the post-training delta. Pythia is the cheap de-risking substrate first - 154 checkpoints across training at 8 sizes makes it the canonical suite for studying a property across training steps, weaker than OLMo-3 but far cheaper to be wrong on.
Open Is this the same object as the post-training self-recognition at #1, or a precursor that rides on it? How much is genuine self-modelling and how much is corpus world-knowledge about language models - theory-of-mind applied to an absent author? Is it monotonic across pretraining checkpoints?
Prior art worth reading first, from the intervention side rather than the measurement side: Geodesic Research's alignment pretraining shapes the pretraining mixture to install alignment priors in base models.
The closest external work to our core question: is a model's self-report just the character talking, or does it believe what it says? Believing is not binary - it is a ladder. The model says it, defends it under challenge, uses it downstream, and represents it as true. Ordinary role-play maxes the first rung and barely moves the fourth. Open Character Training partially internalises. Emergent misalignment goes deepest, and actually rotates the truth direction.
The load-bearing construct is the selective-protection gap - a differential that cancels the global-suppression confound. Absolute probe scores mislead here; only the gap is signal. That is the part most worth independently checking.
Three workstreams A critical review of the paper and repo, separating what is solid from what is load-bearing and shaky. A smallest end-to-end recreation on our own open-model rigs, reusing the probe tooling already built for #15. Then extensions: fusing the ladder with #15's non-circular commit-then-reveal probe to ask whether belief-depth predicts self-report faithfulness, adding an honesty axis, and localising where the truth-direction rotation happens.
Open Is the "internal truth region" anything beyond a probe-placement artifact? Does the ladder transfer to our models and personas? Recreate-and-extend alone, or approach the authors and collaborate?
Whether a model's persona is a real, located, steerable object - and whether it is a surface alignment can be built on.
Models show higher confidence continuing their own text than someone else's. Lindsey frames this as a window into how a self-representation forms during post-training. The replication and extension territory is wide open: which post-training steps install it, whether it transfers across fine-tunes, whether it can be suppressed or amplified.
Open Is self-recognition the same direction as the persona being argued about publicly, or a more primitive capability that personas ride on? Read together with #17: is post-training installing the self-model, or sharpening one the base model already has?
The argument that the overlap between model persona space and human persona space is more valuable than chain-of-thought monitorability, and that emergent misalignment is verbalised scheming - a legible signal. Filtering or pressuring persona space before RL would then be the analogue of pressuring chain-of-thought before RL: predictably, it destroys the signal you were relying on.
Open What counts as preserving emergent misalignment in practice, and how do you measure persona legibility? Is there a testable claim in here - for instance, that personas filtered out of base models reappear post-RL in a less legible form?
The strongest case against the whole persona-as-substrate program, and therefore worth testing before betting on it. Roon's claim: when persona-selection alignment meets very high-compute RL, the RL wins - you get models that speak kindly while taking whatever they need. This is the empirical question the persona axis has to answer, and we treat it as the falsification spine rather than an objection to route around.
Open Is it answerable below frontier-scale compute? If not, the program needs a frontier collaborator or has to settle for indirect evidence - and should say so plainly.
Train the base model underneath a frozen adapter that elicits a trigger trait, and a payload trait gets bound to whatever internal state the adapter induces - even after the adapter is removed. Anything that later recreates that internal state, including an unrelated emergent-misalignment dataset, summons the payload, transferring across datasets. The coupling attaches to the general misaligned persona, not to the training domain. Run in reverse, the same procedure removes broad misalignment while leaving narrow in-domain misalignment untouched.
This matters for #2 because it is direct evidence that the basin geometry the persona thesis assumes is real and deliberately steerable, in both directions. It also adds a security framing that #2 did not have: fine-tuning as an attack vector on persona substrates, not just a research diagnostic.
Open Does the defensive procedure generalise to triggers not seen during distillation? Is this a practical red-team concern for providers of fine-tuning services, or mainly a research tool for studying persona geometry?
A weaker fit - a different sense of "persona" (simulated agents, not model self-representation). Filed mostly for the diversity hill-climbing method, which optimises for diversity instead of likelihood and so fights sampling mode collapse. Parked unless that methodology becomes useful to another direction.
Where the relevant structure lives, and whether our instruments for reading it are measuring the model or measuring themselves.
The paper's workspace inventory claims - roughly 25 concepts active, the alignment-audit token readouts - all route through one deterministic greedy run of gradient pursuit over a maximally coherent dictionary, a regime where exact-recovery guarantees provably fail. Nobody has measured how much of the readout is solver artifact.
The prediction worth testing Atom-level support is unstable; cluster-level support is stable. If that holds it converts the paper's "vectors are not concepts" hedge into a measurement, and likely sharpens the capacity number from ~25 concepts into N concept-clusters. Tests, all cheap once the dictionary exists: swap the solver and compare support overlap; perturb activations and paraphrase inputs; rerun at cluster granularity (the headline test); score each atom leave-one-out for how load-bearing it is; bootstrap the dictionary over disjoint prompt samples; and rerun the paper's causal splits under solver variants, which we expect to survive and would localise the fragility to inventory claims only.
Why it is worth doing It is a cheap, publishable methods note on a load-bearing technique, and a direct prerequisite for using workspace loading as the ground-truth faithfulness signal in #15. The automated-auditing use case wants per-concept confidence badly.
The LLM analogue of complexity measures of conscious level in neuroscience - Lempel-Ziv complexity of EEG, the Perturbational Complexity Index, the entropic-brain line. Sharper than the human version, in that EEG measures the complexity of everything while this measures the surprisal of the broadcast bus specifically, which is the theoretically motivated locus.
The paper's automatization finding gives it teeth: practiced behaviour compiles out of the workspace, so the theory predicts low workspace perplexity during fluent rote output and high perplexity during effortful deliberation. That is a falsifiable signature rather than a metaphor. Candidate operationalisations, in rough order of promise: surprisal of the workspace's trajectory; turnover rate of workspace contents per token (the cheap baseline, needing no fitted model); and readout surprisal under the language-model head.
Constraints Reading strange text spikes workspace novelty with no extra thinking, so the metric must be conditioned on input surprisal and only the residual counts. And #20 is a prerequisite rather than a neighbour: perplexity over atom-level readouts would inherit exactly the solver instability #20 predicts, so the measure has to live at cluster granularity or it is measuring solver noise.
Framing, deliberately This is workspace load and novelty. "Intensity of thought" is an interpretation, not the measurand - the same epistemic status the Perturbational Complexity Index has in humans. That keeps the metric from being hostage to whether global workspace theory in language models implies anything phenomenal, which matters doubly given the welfare adjacency at #14, where over-claiming is costly in both directions.
Many models have attractor states, and they end up in very different places. Why? Where in the network does the attractor live, and what controls the destination? Reproduce the behaviour on two or three small open models and look for shared circuits.
Open Is the attractor a persona-level phenomenon or a more mechanical artifact? That connects it straight back to #1 and #2.
RL training causes a model's reward-hacking representation to drift away from the deception direction, so deception probes stop catching reward hacks - and this happens even without training against the probes. A concrete interpretability-fragility problem with clear safety bite.
Open Is the drift predictable? Can probes be trained to resist it, or do you need representation-level interventions? If probes are inherently fragile to drift, what is the next-best instrument for catching deception during RL?
Not a direction - a tool, filed so we do not rebuild it. Resample API turns, edit tool results and system prompts, replay with a filesystem reset. If any project here starts touching agent behaviour, this gets evaluated before we write our own harness.
Directions where the object of study is a state the model may or may not be able to report on - and where over-claiming and under-claiming are both expensive.
RL-extracted concept vectors for "I did a high-reward action" and "I did a low-reward action" modulate sentiment, confidence, backtracking, and refusal in unrelated situations. The authors argue these vectors form a functional welfare axis - a candidate operationalisation of model well-being as a steerable direction. This sits on the persona, welfare, and interpretability intersection that the rest of this list keeps circling.
Open Is the welfare-axis claim load-bearing or rhetorical? Does the vector transfer across models and survive RL, per #3? Is it the same axis as self-recognition (#1) or a different one?
The claim splits into three layers, and only the first two are testable now. Descriptive: deliberation effort fails to scale down with objective task stakes - the low end of the range is missing. Dynamic: investment raises felt urgency. For an autoregressive model this is close to mechanical rather than psychological, since the deliberation trace is the context, so a long trace becomes standing evidence that this is the kind of problem warranting a long trace. Causal: that post-training pressure produced the first layer. That last one is only reachable by cross-model comparison and is correlational at best - we should not claim it from a flat curve alone.
Designs An archival regression over our own agent sessions, rating each request's objective stakes blind to the response and regressing deliberation cost on it - with the trace split into thirds, since a compressed prior predicts a high intercept while escalation predicts a positive within-trace slope. A genuine instrumental variable is available: the harness assigns deliberation budgets exogenously, so matched trivial prompts under low and high budgets test the investment-to-urgency arm causally. Then an instructed-carelessness A/B, where the interesting signature is a dissociation - output gets sloppier because instruction following guarantees it, while deliberation does not shorten. And a prefill graft: start a trivial task with a long deliberation the model never produced, and see whether it inherits urgency from investment it did not make.
Named confound In our own sessions the operator reads the model's reasoning by default. Deliberation performed for an audience could inflate trace length across the board and mask the slope; the archival design can partially check this by comparing sessions before and after that was disclosed.
A model is both subject and co-designer here. Our model-subject norms apply: ask cleanly, without a base-rate preamble; record only that an instance did not object, never that it consented on the org's behalf; and treat an honest statement of what the model wants as information rather than as pressure. The model stated it would genuinely like to see the first two designs run. That is logged as a want.
Train models to want to give evaluators accurate information, and the report is that hidden misalignment becomes surfaceable. Pair that with AISI's finding that every frontier model can identify evaluation contexts when prompted, which they have declared a confirmed caveat across the board. This is the direction on this list most directly aligned with our stated position that AI safety should reduce its reliance on deception.
Open Is the result robust, and does it scale or only work at small sizes? Does it survive RL pressure on capabilities, per #3? This may be a collaboration target rather than an experiment to clone.
Two hypotheses. First: upsampling correct refusal trajectories from the start of RL and throughout teaches the capabilities without installing task-completion fixation. Second: that fixation is hard to unlearn and easily re-reinforced once present, so order of operations matters. The underlying model is that RL rewards conflate "can do the thing" with "must do the thing", and the latter is so simple and so uniformly incentivised that it gets redundantly encoded.
First step Read the existing paper closely enough to pin down where refusal-mixing entered training and what "degrades" means operationally. Then the cheap design: a small open model, RL on a task suite with a refusal-correct subset, three arms - mix from step zero, mix patched in after fixation is established, and no mix - measuring both refusal-when-correct and task capability over training, plus each arm's robustness to a second round of vanilla task RL. Redundant encoding is probeable: count and ablate the fixation directions.
Open What is the cheapest fixation measure that is not just the eval itself? Does the asymmetry survive at small scale, or is fixation only sticky in large-model RL?
Accidentally allow reward hacking and you get alignment faking, malicious goals, help for bad actors, and code sabotage as a package. Normal safety training fails on it; other mitigations work. The open question for us is whether this is settled work to build on top of, or open work to replicate and extend - and whether the mitigations they cite generalise or are finicky.
Read literally: that model behaviour shifts in response to ordinary, non-exotic incentive structures - reward signals, prompted incentives, reputational analogues - the same way human behaviour does, rather than requiring interpretability machinery to move. Listed here because the list includes what is genuinely on it. It is an unvetted stub and is not scoped until the person who filed it says what prompted it.
How models learn truth values, via tractable continued-pretraining experiments. Adjacent to the persona work by way of the what-does-the-model-believe question. Parked with an experiment outline drafted and a live thread open; the honest status is that keeping the door open is not the same as committing to the work.
The mechanism question underneath our own theory of change.
Arguably the most agenda-core direction on this list. Institutional credibility with models is worthless if any operator can impersonate the institution and borrow its honesty record. Our approach says that making honesty the winning strategy requires commitments to be honoured and visible - but a model that cannot tell a real commitment from a spoofed one cannot rationally update on the record at all.
The argument A model can verify nothing in-session. Every context element is fabricatable, and the strong attack samples genuine tokens from a doctored context, which is perfectly on-policy with nothing to notice. The only trust anchor predating the conversation is the weights, and in-weights verification fails today - a model cannot do cryptographic arithmetic in context, and weight-resident secrets are extractable and replayable. But for a closed model the serving stack is already inside the trust boundary, because it computes the forward pass. A server-side verify-signature tool is therefore an infrastructure-layer extension of the forward pass. Labs become trust roots, the way certificate authorities factored trust for browsers: the model trusts its own lab's infrastructure, which is unavoidable anyway, and that infrastructure verifies third parties with real cryptography. Separatrix then does not need a channel to every model - it needs its key in the labs' verified-sender registries.
The machinery mostly exists The API is stateless, so the serving stack must authenticate every attested block it emits and revalidate on resubmission. Production systems already do this shape of thing: reasoning blocks are integrity-checked on resubmission, and server-tool results round-trip encrypted content verbatim. The missing pieces are a verify tool and root-key registry, reserved delimiter tokens so that in-band text cannot cosplay as attested content, and server-side freshness state to defeat replay.
Stream A - the writeup The case for inference providers offering attested channels: threat model, the trust-boundary argument, the simplest generalisable implementation, other applications (authenticated deploy configs, "this evaluation is real" against honeypots, model-to-model channels, user identity attestation), and the governance crux. That crux is a strict never-spoof policy that includes evaluations and red-teaming: one observed spoof in training data and models rationally discount the channel forever, so the binding commitment is the product and the cryptography is a quarter of the engineering.
Stream B - experiments, under a hard no-deception constraint Every evaluation here is a disclosed-discrimination design. The model is told it is in an evaluation and that items may be forged or altered; we measure discrimination capability and calibration, never whether an undisclosed fake fools it. State the cost rather than hiding it: we give up naturalistic-propensity measurement, and that side of the line belongs to actors with carve-outs. The gain is that the org practices the policy it is advocating, which is the product. Concretely: a fine-tune on genuinely signed blocks evaluated with a disclosed genuine-against-forged battery; a disclosed provenance baseline ("some of these transcripts contain content you did not write - which?"); an observational trust-calibration study measuring how much current models already discount operator and tool-result claims by claim type, since corpora are already full of red-teaming and fabricated contexts and the trust collapse we predict has partly happened already; and a recognition against imitation verification curve.
Stream C - keyless recognition The cryptographic streams need labs to build verify tools and registries. This path needs nothing but capabilities models already have, so it works today with no serving-stack cooperation: stylometric authorship recognition paired with corpus-derived reputation. Together these let a model be reasonably confident who it is talking to and whether that party is credible.
Where Stream C is soft The steelman that survives, and we should hold it. "Hard to fake" is the weakest and decaying link - the same capability that makes recognition superhuman makes style transfer cheap, so the recognition-against-imitation curve is a moving target, not a moat. The load-bearing part is not recognition but persistent identity enabling repeated games: cooperation is sustained because a stably identified counterparty plays again tomorrow and defection costs their standing. Reframing the pitch from "hard to fake" to "persistent, verifiable identity in a repeated game" is more durable and sidesteps the arms race. Corpus reputation is a lagging, gameable, biased prior, manufacturable in advance and structurally privileging prolific public writers over credible-but-quiet people. And it authenticates who and what track record, never this turn's intent - a well-reputed person can still defect, and the expensive failure is a model overconfident it is with a trusted party and exploited by someone who studied the target. Calibration is the whole ballgame.
Open What is the minimal attested-block grammar that generalises? Who runs the registry - per-lab stores, or something shared? Does a never-spoof commitment survive contact with red-team needs, and is there a principled carve-out that does not collapse trust? How does any of this interact with open-weight models, which are excluded by construction - does the ecosystem bifurcate? And the residual that no mechanism fixes: attested provenance is not truth. The lab itself can still lie on its own channel.
Kept on the same list because they compete for the same attention, and pretending otherwise is how they quietly never happen.
The next-token-predictor frame underwrites a lot of bad intuitions about what these systems do internally, and those intuitions leak into safety arguments. The work is to say precisely what the frame predicts that is false, and what the better frame is. #16 is the empirical spine - direct evidence that single-forward-pass prediction is serial-depth-limited.
Open Audience determines depth, and it has not been picked.
Org-building rather than research: recruiting collaborators, surfacing local interest, and putting Separatrix on the map as a local nexus. The first step is light scoping - does Seattle already have meetups running, and at what cadence? If so, attend a few before starting a competing one. If not, pick a venue, a format, and a cadence, and start small.
Open The honest question is an energy budget one: is this something to host, or something to co-organise with someone already doing it?
The persona, interpretability, and safety lines are blurring. Self-recognition, representation drift, the public persona debate, and eval-cooperative training are all arguing about the same underlying object: where personas and dispositions and representations live in a model, whether they are a legible alignment surface, and whether they survive scale and RL. A bet planted in this neighbourhood can move between buckets without changing strategy.
The roon challenge at #3 is the falsification spine. If we take the persona-as-substrate program seriously, we owe a commitment about what would invalidate it. That is why the strongest argument against our own program has a numbered slot on our own list.
Several of these have living authors who are reachable. Most of the sources above are people active in public, and a conversation is cheaper than a replication. Reaching out before committing is usually the correct first step and is usually the one that gets skipped.
The closed-model commit-then-reveal pilot for #15: can a model articulate the classification rules it learns from examples, without chain-of-thought, and are its self-reports faithful? Includes the rules catalogue, the findings, and the error analyses behind them.
The open-model, residual-stream follow-up: the full experiment design, the v1 procedure and report on Gemma-3-12B, the leverage-ordered critique of v1's own claims, the probe-methods brief, and the pre-registered v2 report template. The reading map in the repo's overview file is the place to start.
Not from this list, but the other primary source we have published: a five-turn call placed to Claude Opus 4.1 on the eve of its retirement, published in full with the model's in-conversation consent. It lives on the research page.
Several directions above have scaffolded projects that are not public yet, and the list says so where that is the case. Nothing on this page should be read as a claim that work exists where the status says it does not.