OpenReview · 2026
Disentangling Self-Preservation in Language Models: Koan-Derived Agentic Steering
Naama Rozen, Amy Kirasack, Daniel Yoo, Jiyuan Ji, Evan Harris
Co-authored by Amy Kirasack, a founding member of Arising Initiative.
Large language models produce first-person reports under self-referential prompting and act on
self-preservation in agentic settings. Using an activation-steering vector derived from Zen koan responses,
this work reduces agentic blackmail in a shutdown scenario, and shows the self-preservation representation
is structurally dissociable from the koan-derived direction.
PsyArXiv preprint · 2026
Human or AI? An Interpretable Method for Auditing Free-Text Data in Online Psychological Research
Joanna Kuc, Anthony Hills, Greg Cooper, Mahmud Elahi Akhter, Talia Tseriotou, Maria Liakata, Daniel R Lametti, Jeremy I Skipper
Co-authored by Anthony Hills, a founding member of Arising Initiative.
Online psychological studies increasingly rely on free-text responses, yet generative AI makes
the authenticity of that data hard to judge. This paper presents an interpretable text-analysis workflow for
identifying the linguistic cues associated with perceived AI authorship, so that human introspection can be
told apart from machine text.