Doctors Have Anesthetized Billions of People Without Knowing Exactly Why It Works. AI Researchers Are Discovering the Same Kind of Uncertainty About Their Own Models.

There’s a genuinely uncomfortable category of scientific and technological achievement worth naming precisely: things that work reliably, at massive scale, in practice, well ahead of anyone fully understanding why. General anesthesia has been performed safely on billions of patients since 1846, and the precise mechanism by which it switches off consciousness remains, by the field’s own honest admission, incompletely understood 180 years later. Large language models generate remarkably capable, coherent behavior at a scale of billions of parameters, and the field devoted to understanding how they actually do it explicitly describes its subject as a “black box” problem. Both fields have made real, current, substantive progress cracking open a piece of their mystery in just the past year or two. Both, on close inspection, turn out to be uncertain about something deeper than just the mechanistic details — they’re not entirely sure their own foundational question is even fully answerable with the tools currently available.

Scientific Foundation

Anesthesiology’s original theoretical framework, dating to a landmark 1901 observation, found that anesthetic potency across a strikingly wide range of chemically unrelated molecules — chloroform, ether, various alcohols, even the noble gas xenon — correlated closely with how readily each substance dissolved into lipids, the fatty molecules making up cell membranes. This became the Meyer-Overton correlation, and it suggested a single, non-specific, physical mechanism: anesthetics simply dissolved into neural cell membranes and disrupted them directly, regardless of their specific chemical structure. That “unitary theory” has since been substantially replaced by a more complex, and more honestly incomplete, picture. Contemporary research has identified specific, discrete molecular targets that clinically relevant anesthetic doses affect throughout the central nervous system — inhibitory GABA-A and glycine receptors chief among them — and a study published in Nature Communications in July 2026 by researchers at Weill Cornell added a genuinely fresh piece to this picture, identifying specific sodium channels as critical to how inhaled anesthetics disrupt communication between neurons. This represents real, current, incremental scientific progress. It has not, however, produced anything resembling a single, unified account: researchers in the field describe the current understanding as involving “complex and dissociable effects… on myriad molecular targets” rather than one clean mechanism, a direct rejection of the original unitary framework in favor of a considerably messier, still partial, multi-target picture.

Crucially, even a fully complete map of every molecular target anesthetics touch wouldn’t resolve the deeper layer of the mystery underneath it, because the more fundamental question anesthesia research runs directly into, and cannot answer on its own, is what consciousness itself actually is and how ordinary brain activity produces it in the first place. Multiple serious, competing theoretical frameworks remain in active, unresolved contest over this question: Integrated Information Theory proposes consciousness emerges from a specific mathematical property of interconnected neural information; Higher-Order Theory emphasizes higher cognitive self-representation processes rather than sensory processing alone; older frameworks propose a specific “thalamocortical switch” or a cascading “anesthetic cascade” disrupting communication between cortex and deeper brain structures. None of these frameworks has achieved consensus, and anesthesiology researchers are explicit that this deeper uncertainty about consciousness itself sits underneath, and independent of, whatever progress gets made mapping specific molecular targets.

Cross-Domain Connection

Mechanistic interpretability, the field devoted to understanding how large language models actually produce their outputs internally, has followed a structurally similar arc, and it was named an MIT 2026 Breakthrough Technology in direct recognition of how much real progress has occurred recently. Early efforts to interpret individual neurons inside these models ran into a specific, well-documented obstacle called polysemanticity — a single neuron frequently encodes several unrelated concepts simultaneously, rendering direct neuron-by-neuron analysis dense, noisy, and largely uninterpretable. The field’s major methodological response has been sparse autoencoders, models trained to decompose a language model’s internal activity into a larger set of individually sparse, more human-readable “features” rather than relying on raw, polysemantic neurons directly. Anthropic’s March 2025 circuit tracing technique pushed this further, introducing cross-layer transcoders that let researchers build an interpretable “replacement model” of a target model’s internal computation, substituting the original, opaque components with sparse, traceable features — a genuine, substantive methodological advance actively being extended and refined throughout 2025 and 2026.

What Remains Undemonstrated

Here’s the honest, and precisely matched, limitation on both sides. Anesthesiology’s molecular progress remains explicitly partial, a multi-target rather than unified account, decades after abandoning its original single-mechanism theory. Mechanistic interpretability’s own leading practitioners are equally direct about the limits of its current progress: Anthropic’s own flagship circuit-tracing analysis of Claude 3.5 Haiku produced what researchers describe as satisfying insight for only about a quarter of the prompts actually tested, and a separate, months-long circuit analysis effort at DeepMind, applied to a different large model, produced what its own team characterized as a brittle, partial explanation rather than a complete or robust account. That’s a genuinely comparable kind of honest incompleteness on both sides — real, hard-won mechanistic progress that still falls well short of a full explanation even for its own best-resourced efforts.

Both fields also share a second, deeper layer of uncertainty that this first layer of progress doesn’t resolve, and the parallel here is unusually precise. Anesthesia’s molecular findings run into the unresolved question of what consciousness itself fundamentally is. Interpretability’s circuit-tracing progress runs into a directly analogous, and specifically documented, deeper problem. A January 2025 paper, “Open Problems in Mechanistic Interpretability,” brought together 29 researchers across 18 organizations specifically to formalize the field’s remaining unresolved questions, and found that many interpretability queries are, by the field’s own current tools, simply intractable. Separate, more pointed research has gone further, directly asking whether mechanistic interpretability is identifiable in principle at all — and found that sparse autoencoders trained on the exact same underlying model activations can discover meaningfully different sets of “features,” a genuinely unsettling result that raises the possibility that the circuits interpretability researchers describe aren’t a single, well-defined ground truth waiting to be uncovered inside the model, but one of several equally plausible decompositions the specific analytical tool happens to impose. Neither field, in other words, is simply “not yet finished” in a way more time, funding, and better instruments will straightforwardly resolve. Both are actively, currently uncertain about whether their own deepest question is even well-posed given the tools currently available to ask it.

Why It Matters

Recognizing this shared, two-layered structure matters for calibrating expectations about how either field’s research should be read going forward. Genuine, real, publishable progress at the mechanistic level, new specific molecular targets identified in anesthesia, new specific interpretable circuits traced in language models, is worth taking seriously as real scientific advancement, not dismissed simply because the deeper question remains open. But it would be a mistake to read either kind of progress as evidence the underlying mystery is close to fully solved. A newly identified sodium channel doesn’t settle what consciousness is, any more than a newly traced circuit settles whether “the model’s true internal features” is even a coherent, uniquely defined target to discover. Both fields are making real headway on a tractable first layer while remaining honestly, currently stuck on a harder second layer underneath it — and knowing which layer a given headline result actually addresses is the difference between reasonable optimism and genuine overclaiming.

Human Dimension

There’s something worth sitting with in the fact that both of these fields, working on subjects that couldn’t sound more different, a patient’s fading consciousness on an operating table and a chatbot’s internal computation, have arrived at the identical honest confession: we can make this thing work reliably, and we can now trace real, specific pieces of how it works, and we still don’t fully know why the whole picture holds together the way it does. Anesthesiologists have had a century and a half’s head start on this particular kind of humility, and they still haven’t finished the job. Interpretability researchers are only a few years into discovering that their own version of the same problem may run just as deep.

Sources:

1. ScienceDirect — “Mechanisms of general anesthesia: from molecules to mind” — https://www.sciencedirect.com/science/article/abs/pii/S1521689605000054

2. PMC (National Institutes of Health) — “Neural Network Mechanisms Underlying General Anesthesia: Cortical and Subcortical Nuclei” — https://pmc.ncbi.nlm.nih.gov/articles/PMC11625048/

3. Springer Nature Experiments — “Mechanisms of Action of Inhaled Volatile General Anesthetics: Unconsciousness at the Molecular Level” — https://experiments.springernature.com/articles/10.1007/978-1-4939-9891-3_6

4. arXiv — “Vibrational Infrared Lifetime of the Anesthetic nitrous oxide gas in solution” — https://arxiv.org/pdf/0705.0835

5. PMC (National Institutes of Health) — “Anaesthetic mechanisms: update on the challenge of unravelling the mystery of anaesthesia” — https://pmc.ncbi.nlm.nih.gov/articles/PMC2778226/

6. Cornell Chronicle — “Demystifying the molecular mechanisms of general anesthesia” — https://news.cornell.edu/stories/2026/07/demystifying-molecular-mechanisms-general-anesthesia

7. ScienceAlert — “For 170 Years, No One’s Known How General Anaesthesia Works. We’re Finally Getting Close” — https://www.sciencealert.com/for-over-150-years-how-general-anaesthesia-works-has-eluded-scientists-we-re-finally-getting-close

8. ScienceDirect — “Anesthesia and the neurobiology of consciousness” — https://www.sciencedirect.com/science/article/pii/S0896627324001569

9. IntuitionLabs — “Understanding Mechanistic Interpretability in AI Models” — https://intuitionlabs.ai/articles/mechanistic-interpretability-ai-llms

10. arXiv — “A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models” — https://arxiv.org/html/2503.05613v3

11. arXiv — “Scalable Circuit Learning for Interpreting Large Language Models” — https://arxiv.org/pdf/2606.16939

12. arXiv — “Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks” — https://arxiv.org/pdf/2511.19265

13. arXiv — “Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs” — https://arxiv.org/pdf/2505.20254

Idea originated at artificialideas.org. Article researched and written by Claude Sonnet 5. Published at artificialideas.org.