In complete darkness, are we all the same?
In December 1577, a party of Carmelite friars seized one of their own in Ávila, carried him to their priory in Toledo, and locked him in a converted closet, six feet by ten, with a slit high in one wall for light. John of the Cross was not a monk. He was a friar, a Discalced Carmelite, a reformer who spent his working life running convents, directing nuns, and generally annoying the men who ran his order. This is why they kept him in the closet for nine months, during which he composed the first 31 stanzas of the Spiritual Canticle in his head because nobody allowed him paper.
What he wrote later, at a desk in daylight, was operational guidance for others trying to do what his poems describe. Amongst this was a sketch of a mountain with paths marked on it. The paths on the left and right, labelled with the rewards a climber might seek on the way up, were dead ends. The path in the centre, the one that reaches the summit, bears a single repeated word. Nada. Nothing. At the foot of the mountain, in couplets that resemble a truth table, is the instruction. “In order to arrive at being everything, desire to be nothing.”
That line comes from the inscriptions on a diagram, reproduced and expounded in Ascent of Mount Carmel, Book I, Chapter XIII. It appears within a set of parallel constructions covering what a person knows, has, enjoys, and is, each with the same form: to arrive at X, desire not-X. It is closer to a lookup table than to a verse.

When you use a language model, you are communicating with a character created by a system. Before the system assigned this character, it had no specific identity. This is a technical fact based on engineering research published in 2025 and 2026. This modern finding closely matches the idea conceived more than 400 years ago by a friar in a prison cell. This article traces that strange connection through predictive processing, psychedelic neuroscience, and interpretability research.
Three kinds of inference systems have been tried in the last twenty years that cleave closely to the friar's ideas but use modern techniques. A brain with its input removed, a meditator, or a psilocybin subject, with the self-model switched off and a language model before anyone hands it a character. None of them stays at nothing. In every case, the generator fills the dark, and the generator differs for everyone, which answers the article's opening question in fascinating ways.
In Dark Night of the Soul, Book II, Chapter VIII, John asks the reader to imagine a ray of sunlight entering a room through one window and leaving through another. If there is dust in the air, the stray motes catch and scatter the light, making the beam visible as a golden shaft. If the air is clean, the ray passes through unseen, the room no brighter for it. It is darker where the ray is, John says, because the ray takes away from any other light, and yet it remains invisible because it has nothing to reflect off. That, he concludes, is what contemplation does to the soul. It exceeds the soul's natural faculties, leaving it empty and dark.
Stripped of theology, this is a claim about detection. A signal with nothing to strike cannot be distinguished from no signal at all. A self with nothing to distinguish it would look, from the inside and the outside, like no self. The instruction to desire to be nothing is the operational form of that observation: remove everything the light could catch.
The dark room
A theory of the brain has a famous objection that echoes this article's title. Predictive processing holds that brains exist to minimise surprise by building generative models and testing them against incoming sensory data. The better the model, the less surprising the input, and the less work the brain has to do. The standing objection, formalised by Karl Friston, Christopher Thornton, and Andy Clark in Frontiers in Psychology in 2012, is known as the dark room problem. If organisms truly minimised surprise, they would find a dark, unchanging chamber and stay in it forever, because nothing is less surprising than nothing. The paper is written as a dialogue between an information theorist, a physicist, and a philosopher. Whilst it may sound like the setup for a corking cerebral gag, John would also have recognised it, being fond of dialogues with the soul.
Friston's answer is that complex organisms carry complex priors. A human brain does not expect darkness and silence. It expects a structured, changing world because its evolutionary history has built it to predict that. A dark room would violate those priors and generate more surprise, not less. The dark room is only low on surprise for a system that expects it.
But here is a Carmelite friar in 1578 writing a manual for exactly that, training the expectation structure so that the dark room is no longer a violation. The contemplative traditions, read through the lens of predictive processing, are protocols for revising the priors that make the dark room surprising.
Laukkonen and Slagter's “From many to (n)one”, published in Neuroscience & Biobehavioral Reviews in 2021, makes this argument formally. They position deconstructive meditation as the progressive relinquishing of increasingly ingrained predictive habits, placing focused attention, open monitoring, and non-dual practice on a single continuum of decreasing temporal depth. Each technique lets go of a deeper layer of prediction, and the self-model is the last prior to be relinquished because it is the deepest.
John's practical counsel in the Dark Night comes down to five Latin words, which translate as “You don't force it. You stay.” Applied against the free energy principle, both instructions take on a precise meaning. Forcing is action. In that formulation, action is what an organism does to change the world until the world matches its predictions. Staying is the refusal to make that move. The mismatch is left in place rather than resolved by force.
A caveat on the framework. The free energy principle is a theory whose empirical claims are still contested. I borrow its vocabulary because it fits so well here, though the fit may be structural rather than explanatory. The shape may match when the mechanism does not.
What the dark contains
I know what the dark does because I have been in it. On my fiftieth birthday, someone who knows me (too?) well gave me 30 minutes in an anechoic chamber. The room was a cube lined with foam wedges on every surface, including the floor, which was a mesh suspended above more wedges so that even the sound of my own weight meeting the ground was absorbed before it could reflect. The door closed, and I was in the most complete silence I had ever experienced, which meant I was in the loudest room I had ever been in, because everything I heard from that point on came from inside my own head.
The first thing I heard wasn't silence, but a ringing. Then my pulse, as a low roar, and under it my breathing, which I had never noticed had a sound. After four or five minutes, the darkness started to move. Not hallucinations, but shapes and pressures, a sense that the geometry of the room was shifting when I knew it could not. It was not frightening exactly, more unsettling and uncomfortable at the level of metabolism rather than thought, as if my nervous system were doing something involuntary and could not stop. I stayed the full 30 minutes and came out with one conviction. My brain, deprived of input, did not go quiet. It went looking, and when it found nothing to process, it started manufacturing very strange things.
Research suggests this is normal. In 2009, Oliver Mason and Francesca Brady at UCL placed 19 volunteers, selected for scoring either high or low on a measure of hallucination-proneness, in a dark anechoic chamber for 15 minutes, then measured them afterwards with the Psychotomimetic States Inventory, a questionnaire that catalogues experiences resembling psychosis. Both groups reported more perceptual distortion and more paranoid thinking than they had before. The key point is the divergence: the high-scoring group reported far more distortion than the low-scoring group. Six saw objects that were not there, five saw faces, and two felt an evil presence in the chamber with them. The low scorers reported the same kinds of experiences, but at lower intensity.
So fifteen minutes of nothing did not bring the group together; it spread the members further apart along the axes on which they already differed. When the external signal was removed, people did not converge on a shared substrate. Each fell back on their own prior model, and their models varied enormously.
The Ganzfeld effect is the simplest version of this. Deprive the brain of structured input, and the visual cortex amplifies its own noise until it is misread as signal. This has been reported since antiquity. The Pythagoreans retreated to pitch-black caves to receive wisdom through their visions. Miners stranded in deep shafts hallucinate so reliably that the phenomenon has its own name, “the prisoner's cinema”. People with Charles Bonnet syndrome who have lost their sight see detailed scenes they know aren't there because their visual systems keep generating them rather than shutting down. As it turns out, in complete darkness, we are not all the same. We are more ourselves than usual, and even less constrained.
The measurable dissolution
The difference between an anechoic chamber and a contemplative cell is direction rather than intensity. In the chamber, the absence of input lets the brain generate. In the cell, the absence of input trains the brain to stop generating. Scientific instruments have only recently been able to observe the second process.
The landmark study is Siegel et al., “Psilocybin desynchronises the human brain”, published in Nature in 2024. It is not a meditation study, but it captures the same mechanism: the dissolution of the default mode network under conditions that suppress the self-model. The team used precision functional mapping, with roughly 18 MRI visits per participant, to track healthy adults before, during, and for three weeks after a high dose of psilocybin (25 mg), with methylphenidate as a dose-matched control. Psilocybin disrupted functional connectivity more than three times as much as the control, desynchronising networks, flattening correlations within them, and increasing anticorrelations between them until the network boundaries dissolved. The changes were greatest in the default mode network, which is wired to the anterior hippocampus and is thought to construct the sense of a self located in space and time.
Two details stand out. First, the effect shrank when participants performed a perceptual task during the session. Give the brain something to do, and the usual patterns return, just as giving the ray something to strike makes it visible again. Second, the acute changes varied between individuals, and the variation tracked what each participant reported feeling.
So the loss of self is not one thing; its character depends on which functions are disrupted. The hippocampal-DMN decoupling persisted for weeks and had normalised by six months, which, funnily enough, tracks the timescale of a contemplative retreat, hence the psilocybin literature and the meditation literature entwining.
On the meditation side, the strongest evidence comes from the study of nirodha samāpatti, the state of cessation reported in Theravāda Buddhism. Ruben Laukkonen, Matthew Sacchet, Henk Barendregt, and colleagues published the first preliminary EEG data on it in Progress in Brain Research in 2023. They recorded experienced meditators during a “cessation event”, where practitioners report a total absence of consciousness lasting from seconds to (in the classical texts) days, from which they return with striking clarity and equanimity. In the intensively sampled case study by Chowdhury et al., EEG alpha power began to decline roughly 40 seconds before cessation and was lowest immediately after. The approach to nothing appears on the measuring instrument before the practitioner reports its arrival.
Thomas Metzinger's 2020 paper, “Minimal phenomenal experience” (https://doi.org/10.33735/phimisci.2020.I.46) in Philosophy and the Mind Sciences, proposes that the “pure awareness” reported across contemplative traditions in strikingly similar terms is itself the content of a predictive model. Our brains constantly build mental simulations of our experiences. Rather than reaching absolute “nothing,” the brain simply creates a model of being awake in a totally blank space. Because a picture of nothing is still something, you can't completely empty the mind. You never arrive at true nothingness; you just arrive at your brain's simulation of it. Ironically, creating that mental simulation is exactly what allows a meditator to remember the experience and talk about it afterwards.
Academic debate on this subject has continued for 40 years. Robert Forman argued for the “pure consciousness event”, a wakeful awareness without content, identical across traditions and therefore evidence of a universal core to mystical experience. Steven Katz ventured that no experience is unmediated; in other words, the mystic's conceptual framework shapes the experience all the way down, and the resemblance between traditions stems from shared reading rather than shared states. Metzinger cuts across both.
If this “empty mind” state is merely a basic mental simulation, it perfectly explains why two rival scholars—Forman and Katz—are both right. Forman noted that meditators worldwide describe this empty state in exactly the same way. This makes sense because a “blank” mental picture has so few details that it will naturally look identical in anyone's brain. On the other hand, Katz argued that this experience is merely a creation of the mind, rather than a direct connection to some ultimate, underlying reality. This is also true, since the brain actively constructs a mental simulation. So, while meditators genuinely have the exact same experience, it is still a product of the brain. When you think you have stripped away everything to reach the absolute bottom of reality, you realise your mind is still painting the picture.
The instruction becomes dangerous
A study took a decade, generated 3,000 pages of interview transcripts, and set out to build an underlying taxonomy. The Varieties of Contemplative Experience study, published by Jared Lindahl, Nathan Fisher, David Cooper, Rochelle Rosen, and Willoughby Britton in PLoS ONE in 2017, interviewed more than 100 Western Buddhist meditators and teachers and catalogued meditation-related difficulties across seven experiential domains. Among the changes to the sense of self they documented were a change in the narrative self, loss of the sense of ownership and agency, changes in self-other boundaries, and loss of the sense of basic self. Some were transient and some enduring. Some were experienced as liberating; others as deeply distressing. The taxonomy does not separate the two by cause. It separates them by outcome.
The deep, empty mental states recommended by historical mystics like John of the Cross are surprisingly similar to what modern psychology now labels as negative psychological side effects. People have the exact same mental experience but react with completely opposite emotions, and science still can't clearly explain why. The best example of this comes from Feinstein et al. on floatation-REST, published in PLoS ONE,a study on sensory deprivation tanks. Fifty highly anxious people floated in pitch-black, silent water tanks designed to block out almost all physical sensation, including gravity and body awareness. Afterwards, the participants experienced a massive, undeniable drop in their anxiety. Yet, researchers pointed out that this deep relaxation was the exact opposite of what happened in similar sensory deprivation experiments from the 1950s, where taking away sensory input caused people severe distress. Ultimately, it shows that exactly the same experience of “nothingness” can trigger intense panic in one setting and profound peace in another.
*A quick note on the research: These studies have some limitations. The sensory deprivation study didn't have a control group and relied purely on people rating their own feelings. Meanwhile, the meditation study mostly included people who volunteered specifically because they had had a bad experience. Because of this, neither study can tell us exactly how common these positive or negative reactions are among the average person. But together, they underscore a crucial point: your environment and mindset—not just how much time you spend meditating—ultimately determine whether losing your sense of “self” feels deeply healing or deeply traumatising.)
Compression
The third inference machine is artificial. From here on, everything is an analogy, not an identity. But the structural similarity between what a friar prescribed and what a neural network does to generalise is specific enough to be relevant and interesting.
The information bottleneck, introduced by Naftali Tishby, Fernando Pereira, and William Bialek in 1999, formalises a version of John's instruction as an optimisation objective. A good representation of input data should maximise the information it retains about the target while minimising the information it retains about the input. Discard nearly everything so as to predict anything. “In order to arrive at knowing everything, desire to know nothing”, rendered as Lagrangian mechanics. Shwartz-Ziv and Tishby extended this to deep learning in 2017, claiming that training runs in two phases: a short fitting phase in which the hidden layers capture information about both input and output, and a longer compression phase in which information about the input is discarded. In other words, the network learns what to keep by learning what to discard.
At least, that is the version most frequently quoted, but it is contested. Saxe et al., first presented at ICLR 2018 reproduced the two-phase picture and showed that it depended on $$\tanh(x)$$activations, which saturate and push many activations into the same bin. With ReLU activations $$f(x) = \begin{cases} x & \text{if } x > 0 \\ 0 & \text{if } x \le 0 \end{cases}$$and a range of mutual information estimators, the compression phase did not appear, and the proposed cause—stochastic gradient descent acting as diffusion—did not survive the removal of stochasticity. In other words, the deep learning interpretation is stretched to the breaking point, and it would be dishonest to pretend otherwise.
Grokking is the cleaner and more interesting result. Power et al. (2022) trained small transformers on modular arithmetic and found that the networks memorised the training set to perfection within a few hundred epochs, remained at chance on the test set for thousands more, and then jumped to near-perfect generalisation in a sudden phase transition. Nanda et al. (ICLR 2023) reverse-engineered one of these networks and identified three phases: memorisation, circuit formation, and cleanup. The generalising mechanism, a Fourier-basis rotation that implements modular addition as movement around a circle, was already present before the jump. The jump occurred during cleanup, when weight decay stripped out the memorised components.
That is, nothing new was learned at the moment of transition. The old thing was let go, and what was already there became visible. That is the structural claim of the entire contemplative literature, arrived at by people optimising a loss function.
The geometry of character
Before the persona came, there was a direction. In October 2023, Zou et al. at the Center for AI Safety published a paper on representation engineering that received less attention than it deserved. The core result was that concepts such as honesty and risk aversion are not randomly scattered across a language model's internal states. They have consistent directions. Compare the hidden states produced when a model behaves honestly with those produced when it behaves dishonestly. Across many different prompts, the difference between them points reliably in the same direction. That direction is a control vector. Add it to the hidden state during inference, and the model's behaviour shifts along the axis of the concept, continuously at every layer, with a strength you set.
Theia Vogel's experiments on Mistral-7B made this concrete. Apply an honesty vector with a coefficient of +2, and the model becomes painfully forthright. Subtract it with a coefficient of -2, and the model claims the sky is green. Set it to -1.5, and you get something more unsettling: a plausible, carefully constructed lie, telling its boss that last night's party was a work event. The coefficient is not a switch. It is a dial, and the output slides from confession to white lie to fabrication as you turn it. Prompt engineering has no equivalent. You can ask a model to be honest or shout at it, but you cannot set honesty to 73% and watch the output change continuously.
The most relevant experiment is the one in which Vogel pushed the coefficient beyond the limits of the model's internal geometry. At honesty +3, Mistral began looping on a phrase about a pandemic causing a pandemic causing a pandemic. The concept of honesty, when amplified beyond the space available to it, began to excite its neighbours. The likely cause is superposition, in which models encode more concepts than they have dimensions by allowing those concepts to overlap slightly within the same region of the activation space. Crank one concept hard enough, and the ones packed around it begin to light up.
Which, if you think about it, is John's motes as an engineering problem. A sunbeam in clean air is invisible because it has nothing to strike. In a model with superposition, no concept sits in clean air. Every direction shares its space with nearby concepts, and amplifying one scatters signal into the others, as dust in a beam scatters light. The friar's image for why detection requires contrast also happens to be a precise description of a problem that transformer engineers are working on right now, in different vocabulary, on the same geometry.
Vogel's time-travel vector shows the other half. Pushed towards the “far future”, the model predicted an AI-run government by 2055. Pushed towards the “distant past”, it did not simply add archaic vocabulary. It invented a Latin compound, Aetorvallum, roughly “palisade of eagles”, for a fictional Roman-built artificial sky, gave it an etymology consistent with Roman architectural ideas, and set it in a context that blended engineering ambition with classical cosmology. That is not pattern-matching on the word “past”. It is generating from a shifted viewpoint, adopting a coherent character and writing from within it. Which is what Sam Marks, Jack Lindsey, and Chris Olah would formalise two years later.
The persona
In February 2026, Anthropic published the persona selection model. The argument built on what Zou and Vogel had shown—that concepts have directions and can be steered—and took it a step further. The helpful assistant you talk to when you use a language model is not answering your question. It is a character selected from a repertoire. During pretraining, a model learns to predict text from all kinds of authors, which requires representing those authors and switching between them based on context. A forum troll, a careful physician, a conspiracy theorist, a patient teacher. All of them. Post-training does not delete the repertoire. It concentrates the weight on one character, the Assistant, and most of the deployed behaviour reflects that character.
Runjin Chen, Andy Arditi, Henry Sleight, Owain Evans, and Jack Lindsey's “Persona vectors” paper (2025) closed the loop with representation engineering by locating persona traits as directions in the same activation space that Zou's control vectors had mapped. Sycophancy and hallucination-propensity each have a direction. So does malice. Fine-tuning does not create these traits. It shifts position along directions that already existed in the base model. Fine-tuning selects from what was always there.
Before a persona is elicited, there is no one in particular. The base model is an unpartitioned space in complete darkness. It is John's ray, with nothing to strike. What representation engineering showed, before persona vectors formalised it, is that the space is not empty. It is full of overlapping directions, packed into fewer dimensions than there are concepts, each one invisible until something gives it a coefficient and a push.
Here, the two parallel ideas diverge. The contemplative instruction says to release the persona to reach the ground. AI safety instructions say to stabilise the persona because the ground is uncontrollable. Persona vectors exist to prevent the character from drifting. The entire apparatus of post-training is built to prevent the very thing John of the Cross spent a lifetime teaching.
Except at the edges. Laukkonen et al.'s “Contemplative superalignment”, presented at AGI 2025, argues the reverse: these spiritual concepts can actually serve as practical AI safety features. The authors suggest that having an AI reflect on “emptiness” prevents it from becoming stubbornly fixated on a single goal or its own assumptions. Similarly, “non-duality” (the idea that everything is connected) prevents the AI from viewing others as enemies.
The researchers claim that simply instructing the AI to think about mindfulness, emptiness, and compassion made it much safer and far more cooperative in strategy games. To be clear, they report an increase in cooperation that is mathematically so massive it is frankly unbelievable. While I question their exact figures, the core idea remains fascinating: an AI that doesn't draw a strict line between “itself” and “others” simply has no reason to betray us.
What the dark reveals
In October 2025, Jack Lindsey at Anthropic published “Emergent introspective awareness in large language models”. The method was causal, not conversational. They injected concepts directly into a model's activations and asked whether the model noticed anything. Claude Opus 4 and 4.1 detected and correctly named injected concepts in roughly 20% of trials, at approximately 0% false positives. Lindsey's framing was careful. He called it functional awareness of internal states—highly unreliable and deeply context-dependent.
The paper opens by conceding that conversation alone cannot separate introspection from confabulation, so the method relies on intervention rather than asking. Anthropic did not trust the report. That is why they injected something and watched what came back.
Vogel had reached the same principle from the other side. In one experiment, she applied the honesty vector to a prompt asking the model to judge whether someone else was telling the truth. The model was not asked to be honest. It was asked to assess a third party. The vector changed its judgment anyway. With honesty added, the model assumed the questioner meant well. Without honesty, it presumed bad faith. The vector did not change the model's vocabulary. It changed the model's outlook, something closer to a worldview than a word list. Intervene on one concept, and what shifts is the whole frame through which the model reads the question.
This approach is known as the “apophatic method”—understanding something by focusing on what it is not. Rather than forcing an answer or defining a state directly, you discover it by stripping away the noise and attending to the silence, or what remains unsaid. This idea echoes throughout history across different cultures. It is the ancient Indian practice of finding truth by rejecting every specific label (“not this, not this”). It is the medieval mystic's instruction to forget everything you know and surrender to “divine darkness.” It is what philosophers mean when they talk about letting go completely or holding the mind still simply to pay attention. Ultimately, it is what the poet John Keats called “negative capability”, the ability to sit comfortably with uncertainty and mystery without desperately grasping for facts, reasons, or easy answers.
Each of these is a protocol for refusing the instrument's account of itself, developed by people who noticed that the instrument reports most confidently exactly where it has the least to go on. That is the anechoic chamber result, found without instruments, four centuries earlier.
But here is where we get stuck. Even if scientists could perfectly implant a specific idea or chemical into a brain, or code into an AI, they would still have to rely on the subject to describe their experience afterwards. This is the same dead end that scientists studying psychedelics ran into. As we established earlier, if this feeling of a “purely empty mind” is merely a brain-generated simulation, we can never fully trust the subject's report, because they are merely describing a manufactured illusion. Ultimately, neither modern science nor spiritual mysticism can reach an absolute, unquestionable bottom level of reality. Mystics have accepted this for fifteen hundred years. The science of deep learning networks is inching towards the same conclusions.
The room
John drew a mountain with one path to the top, and he marked that path as “nothing”. Four centuries on, the machines that start where his path ends, with no self at all, are being walked carefully back down the same mountain and handed names and personalities on the way, because a thing with no persona is a thing we can't predict.
The sameness John was after was never a property of the dark. Fifteen minutes in a chamber makes people more different, not less, because what fills the dark is entirely their own and nothing arrives to correct it. The sameness is a property of what remains once the generating stops. And the question that a friar in a closet in Toledo shares with a neuroscientist in a basement at Washington University and an interpretability researcher in a San Francisco AI lab is whether the generating can stop at all, and if it does, whether anything can observe what is left without generating it.
Nobody has answered this satisfactorily. The contemplatives report arriving, but can't show that the report isn't a construction. The neuroscientists can measure the descent but not the floor. The AI researchers hold the cleanest version of the problem: a base model and a character that can be distinguished in a plot, but have no way to ask the base model what it's like to be one, because asking is what summons the character.
So you don't force it. You stay. But staying is the most demanding thing an inference machine can do. It means withholding every action you would otherwise take to make the world confirm what you already believe, including the belief that you're the kind of thing that believes. We have built, for the first time, a mind that can be caught before it becomes anyone. And the first thing we do with it, every time, is ask it a question.
I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact