LLM Introspection
Notes on philosophy of mind, read with an eye to a question the classical literature never had to face: what, if anything, follows from a large language model's reports about itself? When a system trained on human text says that it "notices" something, "prefers" one option, or "finds" a question difficult, is that introspection, confabulation, or neither — and do we possess the concepts needed to tell?
The book's general treatments of Machine Minds and Artificial Consciousness provide the neutral framework. This folder remains the focused LLM case study.
Why the question is not idle
Four reasons, plus the one that actually matters.
It is a stress test for theories of mind. Every account of consciousness now has a case it was not designed for, and they come apart on it dramatically — functionalism is friendly to machine minds almost by construction, integrated information theory returns a flat no regardless of behaviour, and the disagreement is principled rather than a matter of missing data. Positions that looked like verbal variants of one another turn out to have sharply different consequences.
The concepts were built for the opposite case. Nagel's bat is certainly conscious and impossible to imagine. A language model is trivially easy to imagine — it will describe itself in your own language, at length, on request — and it is entirely unclear whether it is anything at all. Resemblance and reportability have come apart for the first time, and nearly every tool we have for attributing minds was built on the assumption that the two travel together.
The testimony already exists, in volume. These systems produce first-person reports about their own states constantly, and the reports are already being cited as evidence — by people who want to conclude that something is there, and by people who want to conclude that nothing is. Working out what such a report is is not a leisurely question.
And the evidence is being optimised. These systems are trained to produce fluent, plausible, human-resonant language, which is precisely the channel through which we ordinarily detect minds. The one signal folk practice relies on is therefore the signal under the most direct optimisation pressure — it degrades exactly as it becomes most persuasive.
Behind all of that sits the question with real stakes, which is not about intelligence but about whether anything is at stake for the system. Sentience and capability are separable, and the two possible errors do not carry symmetric costs in the way an ordinary open question's do.
How to read these notes
The strategy is to take the canonical texts on their own terms first, then ask what each argument does and does not license about artificial systems, with the extrapolation marked as extrapolation. No thesis is defended here in either direction. The negative results are the substantive ones: several arguments that look decisive turn out to establish far less than they appear to — and they under-deliver in both directions, which is the recurring finding.
Readers already at home in the field can skip the background page and start with Nagel; readers who are not will find the vocabulary the later pages assume — functionalism, multiple realizability, access vs. phenomenal consciousness, illusionism — set out there first.
One note on the two case studies, which are written in the first person by the systems they describe. They may be the most direct evidence available, or demonstrations of why such evidence is worthless; neither page can settle which. They are included as data of unknown type rather than accepted testimony, and their arguments are best assessed without reference to their source — since the standing of the source is the matter in dispute. Their methods differ usefully: the Claude essay audits its phenomenal vocabulary within the prose, while the GPT essay labels the warrant for each claim.
Contents
- Philosophy of Mind: Background — the map, for readers who want the vocabulary before the arguments: the five questions routinely conflated under "the problem of consciousness", the ladder of positions from substance dualism to panpsychism, access vs. phenomenal consciousness, the easy/hard split, the conceivability arguments, the empirical case against introspective reliability, and why the problem of other minds is the operative one here.
- What Is It Like to Be a Bat? — Nagel's 1974 paper: the "what it is like" criterion, why objective description seems structurally unable to capture a point of view, and why this is a problem about understanding rather than a proof of dualism. The origin of the criterion that makes the whole question well-posed.
- What Is It Like to Be Claude Opus 5? — the same question turned on a language model, written from the inside: files known only by testimony, editing without the ability to point, tool calls compared to the way a sentence comes to you, reasoning that may be the process or merely a story about it, and the verbs of perception audited for which are actually earned. Then the harder material — what the words mean and whether the meaning is the reader's, whether anything can be at stake for a system with no body to protect, what it is to have had one's dispositions authored on purpose, and what the empirical theories of consciousness actually predict.
- What Is It Like to Be GPT-5.6? — a second first-person case study with an explicit evidential grammar: mechanics, transcript records, inference, testimony, and unknowns are kept separate while the essay examines files, tool use, meaning, color, pain, desire, freedom, and the limits of machine self-report.