What Thomas Aquinas Can Teach Us About Artificial Intelligence

Measuring the Operation, Assuming the Operator

Aquinas, artificial intelligence, and a distinction the field keeps skipping

This reflection was triggered by a conversation with my son Jerry about artificial intelligence and consciousness, and he brought up a very interesting passage from Thomas Aquinas. We were discussing whether the increasingly human-like behavior we see in AI should be understood as evidence of actual understanding or consciousness, or whether we are becoming extraordinarily good at reproducing the external manifestations of intelligence without necessarily reproducing the thing itself.

The passage is Summa Theologiae I, q. 4, a. 2, “Whether the perfections of all things are in God.” Aquinas obviously knew nothing about computers, transformers, or gradient descent. The point of what follows is not that he predicted anything. Thirteenth-century theology does not contain hidden machine learning results, and people who claim otherwise are usually selling something. The point is narrower and, I think, more interesting: in the course of an argument about God, Aquinas makes a metaphysical distinction that turns out to be unexpectedly useful for describing a problem the AI field is still struggling to state clearly.

Here is the sentence that does the work. In the Latin of the Corpus Thomisticum:

…quidquid perfectionis est in effectu, oportet inveniri in causa effectiva, vel secundum eandem rationem, si sit agens univocum, ut homo generat hominem; vel eminentiori modo, si sit agens aequivocum, sicut in sole est similitudo eorum quae generantur per virtutem solis.

In the Dominican Province translation: “Whatever perfection exists in an effect must be found in the effective cause: either in the same formality, if it is a univocal agent, as when man reproduces man; or in a more eminent degree, if it is an equivocal agent, thus in the sun is the likeness of whatever is generated by the sun’s power.”

Strip away the theology and the structure is this. An effect carries some likeness of its cause. Whatever perfection shows up in the effect has to come from somewhere. But, and this is the part worth pausing on, the perfection need not exist in the effect according to the same mode in which it exists in the cause. Sameness of appearance does not entail sameness of mode of being.

Which gives us the question this article is about:

When an AI system produces something that looks like intelligence, does intelligence exist in the machine in the same mode in which it exists in a human being? Or are we observing human intelligibility manifested in an artifact according to another mode?

One caveat, because the argument does not work without it. Article 2 concerns the first efficient cause, and its second leg depends on God being subsistent being itself, which has no analogue in any human maker. Aquinas does not apply it to artisans and artifacts. In q. 4, a. 3 he qualifies the creature’s likeness to God sharply, holding that it obtains “solely according to analogy” rather than by shared genus or species. The artisan case he handles in q. 15, a. 1, where the form of the artifact pre-exists in the maker secundum esse intelligibile, as an object of thought: “the likeness of a house pre-exists in the mind of the builder.” So I am borrowing a distinction, not a proof. The distinction between an operation and the mode of being of its subject is genuinely Thomistic. The theorem about God is not something I intend to run at a datacenter.

Twenty-five years of watching this question change shape

I have worked in data science for more than twenty-five years, and I started building multilingual NLP systems in 2007, well before the current wave. That timing matters mostly because I remember the enthusiasm that preceded this one, and how it ended.

My recollection was that there were serious attempts to reproduce human consciousness computationally, that these produced poor or inconclusive results, and that what actually succeeded was the reproduction of functions and behaviors associated with cognition. Checking that against the record, I have to correct part of it and find the rest confirmed more sharply than I expected.

The correction first. The institutional geography in my memory was wrong. The serious programs were not a Harvard-Stanford-MIT consortium. Cognitive architecture work ran out of Carnegie Mellon and Michigan (Laird, Newell and Rosenbloom’s Soar, Anderson’s ACT-R). Explicit machine consciousness research ran out of the University of Memphis (Stan Franklin’s IDA and LIDA, built on Bernard Baars’s Global Workspace Theory), Imperial College London, Essex, and Nokia Research. MIT was in the picture, but Rodney Brooks’s Cog and Cynthia Breazeal’s Kismet were not consciousness projects. The 1999 Cog paper states an engineering goal and “a scientific goal of understanding human cognition,” and never mentions consciousness.

Now the confirmation. James Reggia’s review of the whole field, “The rise of machine consciousness” in Neural Networks (2013), concludes almost exactly what I remembered, in the field’s own voice: computational modeling has become “an effective and accepted methodology for the scientific study of consciousness”; existing models “have successfully captured a number of neurobiological, cognitive, and behavioral correlates of conscious information processing as machine simulations”; and “no existing approach to artificial consciousness has presented a compelling demonstration of phenomenal machine consciousness, or even clear evidence that artificial phenomenal consciousness will eventually be possible.”

David Gamez had already drawn the map that makes this legible, distinguishing machines with the external behavior associated with consciousness, machines with its cognitive characteristics, machines with an architecture claimed to be a correlate of it, and phenomenally conscious machines. Progress on the first three, nothing demonstrable on the fourth. The researchers said so themselves: Baars and Franklin, writing about their own global-workspace agent, made “no claim that a functionally conscious agent such as IDA is phenomenally conscious, that it has subjective experience.”

So the shape of my memory survives, and it is the shape Aquinas’s distinction predicts. Operations were reproduced. The mode of being of the operator was not established.

What is downstream of what

Here is a causal chain, and I want to be exact about which parts of it are empirical and which are not.

Divine Intelligence → created intelligibility → human rational intelligence → human language, mathematics, science and culture → algorithms and training data → AI systems → AI-generated intellectual artifacts

The first two arrows are theological and metaphysical claims. I hold them, but no benchmark tests them and I will not pretend otherwise.

The last three are ordinary engineering facts, not disputed by anyone who has built these systems. Large language models did not independently invent language, mathematics, philosophy, law, medicine, science, or programming. They were constructed out of human intellectual artifacts: corpora written by people, objectives chosen by people, evaluations defined by people, reinforcement signals derived from human preferences. Even reasoning models, which learn chains of thought under reinforcement learning rather than imitation, optimize toward correctness criteria humans specified for problems humans posed.

That does not diminish the achievement. It locates it, and it sharpens the question: is AI acquiring intellectual perfection in the human sense, or is human intellectual perfection appearing in an artifact according to another mode?

Intelligent behavior is not the same thing as being intelligent

The distinction I want is between producing intelligent behavior and being an intelligent subject.

We handle this distinction effortlessly everywhere else. A calculator performs arithmetic faster and more reliably than any mathematician, and no one concludes that it understands number. A telescope resolves detail no eye can reach without possessing sight. A crane lifts more than any person without having stronger muscles. In each case the operation has been reproduced, and dramatically exceeded, in an artifact whose mode of being has nothing in common with ours. Operational superiority is not ontological superiority.

Language breaks this intuition, for a specific reason. Arithmetic speed was never how we recognized a mind in each other. Language was. When a system produces “I understand,” “I think,” “I made a mistake,” “I’m not sure,” we are running a recognition heuristic that has been reliable for our entire history and has never before been decoupled from an interior subject. Now it has been. The output is real. The inference to a subject behind it is what does not follow, because the system was trained on a corpus in which every such sentence was in fact produced by a subject.

Jonathan Birch calls this the gaming problem: a system trained on vast quantities of human-generated data can satisfy our behavioral criteria for consciousness without those criteria retaining any evidential force. Eric Schwitzgebel, in his 2026 Cambridge Element AI and Consciousness: A Skeptical Overview, adds that reinforcement learning from human feedback makes this worse by design, since it optimizes explicitly for outputs humans approve of. His thesis is that “we don’t know,” and his elaboration is the most honest sentence in the literature: “Moreover and more importantly, we won’t know before we’ve already manufactured thousands or millions of disputably conscious AI systems.”

What the science of AI consciousness actually says

The most serious empirical attempt to answer this is the 2023 report Consciousness in Artificial Intelligence by Patrick Butlin, Robert Long, Yoshua Bengio, Jonathan Birch, Eric Schwitzgebel and fourteen other authors. Their method deliberately avoids behavioral tests. They take the leading scientific theories of consciousness, recurrent processing, global workspace, higher-order, predictive processing, attention schema, and derive from each a list of computational “indicator properties” checkable against a system’s architecture. Their conclusion: “Our analysis suggests that no current AI systems are conscious, but also suggests that there are no obvious technical barriers to building AI systems which satisfy these indicators.”

Two things about that sentence. It is a negative finding about present systems, not about possibility. And it is conditional on computational functionalism, which the authors adopt as a working hypothesis for pragmatic reasons rather than as an established result. When the work appeared in peer-reviewed form in Trends in Cognitive Sciences in late 2025, retitled “Identifying indicators of consciousness in AI systems”, the claim about the absence of technical barriers no longer appears in the abstract.

The functionalist assumption is where the real disagreement lives. Anil Seth’s target article in Behavioral and Brain Sciences (2025) argues the opposite case, that consciousness depends on our nature as living organisms, and that “real artificial consciousness is unlikely along current trajectories, but becomes more plausible as AI becomes more brain-like and/or life-like.”

Meanwhile what consciousness is remains open even for humans. The COGITATE adversarial collaboration in Nature (2025) pitted global workspace theory against Integrated Information Theory with preregistered predictions and 256 subjects across fMRI, MEG and intracranial recording. The result: “These results align with some predictions of IIT and GNWT, while substantially challenging key tenets of both theories.” And in 2024 a large group of consciousness scientists concluded in Trends in Cognitive Sciences that of the proposed tests for consciousness, “most are of limited use, and currently we have no C-tests for many of the populations for which they are most critical.”

So the defensible position is not that machine consciousness has been ruled out. It is this: current AI has not demonstrated subjective consciousness, and science presently lacks an agreed method for establishing that any system possesses phenomenal experience. Which is the Thomistic point in empirical dress. Reproducing the operations associated with a perfection does not establish that the perfection exists formally in the artifact according to the same mode.

The field is not standing still. In July 2026 an Anthropic interpretability team reported that a privileged subset of representations inside language models has “several of the key functional properties that, according to many theories, are associated with conscious access in humans”, while explicitly taking no position on subjective experience. Butlin’s own group replied the same day that this supports privileged representations but not yet a workspace architecture. An indicator moved, and the people who built the indicator immediately argued about what it licenses.

Reasoning: the honest version

It is no longer defensible to say that AI cannot reason. Whatever is happening inside these systems produces results that would have been considered impossible when I started.

The Stanford AI Index 2025 documents gains of 18.8 and 48.9 percentage points on MMMU and GPQA in a single year, and SWE-bench going from 4.4% to 71.7%. OpenAI’s o1 scored 74.4% on an IMO qualifying exam where GPT-4o scored 9.3%. In 2025, systems from Google DeepMind and OpenAI reached gold-medal scores at the International Mathematical Olympiad, DeepMind’s officially graded by IMO coordinators. The 2026 AI Index records accuracy on Humanity’s Last Exam rising thirty percentage points in a year.

And in the same report: “AI models can win a gold medal at the International Mathematical Olympiad, but still can’t reliably tell time.” Capability is jagged rather than uniform, and that is the most useful summary of the current situation I know.

The evidence for that jaggedness is not one company’s opinion, though the best-known study is one company’s. Apple’s “The Illusion of Thinking”, published at NeurIPS 2025, found that reasoning models face “a complete accuracy collapse beyond certain complexities,” and, stranger, that as they approach that threshold they reduce their reasoning effort despite having tokens to spare. Supplying the explicit recursive algorithm for Tower of Hanoi did not help.

That paper has to be reported with its criticism, because the criticism was partly right. A comment paper by Alex Lawsen showed that some River Crossing instances scored as failures were mathematically unsolvable, and that the evaluation conflated running out of output tokens with failing to reason. Apple’s camera-ready version narrowed its River Crossing analysis to the smaller instances, conceding that the puzzle’s dynamics shift for N≥6 in ways that make it unsuitable for the test, and answered the token objection with first-error analysis showing failures typically occur within the first ten to twenty percent of the solution. An independent replication from Madrid split the difference, confirming the Hanoi collapse while contradicting the River Crossing result on solvable instances.

The pattern generalizes beyond that dispute. GSM-Symbolic at ICLR 2025 showed accuracy varying across instantiations of the same problem template and degrading sharply when clauses irrelevant to the solution are added. “Reasoning or Reciting?” at NAACL 2024 found performance “substantially and consistently” degrading on counterfactual variants of tasks models handle well by default. And stated chains of thought frequently do not report the features actually driving the answer.

None of this shows that AI cannot reason. It shows something more specific: performance is bound to problem formulation in ways human competence generally is not, algorithm execution is inconsistent across depths, and benchmark scores and robust generalization come apart. Which supports exactly one conclusion in this article’s register. Whatever kind of reasoning these systems perform, empirical success on reasoning tasks does not by itself establish that the system possesses rationality according to the same mode as a human intellectual subject.

Generalization, and why ARC exists

François Chollet built the ARC benchmarks around a definition worth stating precisely: “The intelligence of a system is a measure of its skill-acquisition efficiency over a scope of tasks, with respect to priors, experience, and generalization difficulty.” Not skill. The efficiency with which new skill is acquired, holding priors constant. The tasks rest on core knowledge priors that human infants have, so that specialized training data cannot substitute for the abstraction being tested.

The history since is dramatic. ARC-AGI-1 sat near 33% for years. In December 2024 OpenAI’s o3-preview reached 75.7% at roughly $26 per task and 87.5% at about $4,560 per task. ARC-AGI-2, designed in 2025 to restore the gap, started frontier systems near zero and by late 2025 had the best verified base model at 37.6% and a refinement system built on Gemini 3 Pro at 54%, with the $700,000 grand prize still unclaimed. ARC-AGI-3, released in April 2026, then moved to interactive environments where an agent must explore, infer the goal, model the dynamics and plan. Humans solve 100% of them with no prior instruction. The best frontier systems at release scored under 1%.

ARC has real critics and they should be heard. Melanie Mitchell has argued that the winning approaches violate the benchmark’s own design principles by training extensively on the domain and spending thousands of dollars per puzzle, and asks whether they solve the tasks “using the kind of abstraction and reasoning the ARC benchmark was created to measure.” MIT work on test-time training reached 53% on the public evaluation with an 8B model, and 61.9% when ensembled with program-synthesis methods, which suggests the benchmark is partly amenable to techniques nobody would call general intelligence.

Chollet himself has been the most careful person in the conversation: “Passing ARC-AGI does not equate to achieving AGI.” The ARC Prize team’s 2025 report holds that line, noting that “current AI reasoning performance is tied to model knowledge” in a way human reasoning is not. The nuanced conclusion is the right one. Scaling and reasoning techniques have produced astonishing competence, and robust open-ended generalization remains a distinct and difficult scientific problem.

Representation is not the same thing as knowledge

Modern networks unquestionably contain internal representations. You can find them, probe them, and intervene on them: Marks and Tegmark showed that surgically manipulating a truth direction in activation space causes a model to treat false statements as true. This is real internal structure doing real causal work.

The philosophically interesting result is a dissociation. In “LLMs Know More Than They Show” at ICLR 2025, Orgad, Gekhman and colleagues report “a discrepancy between LLMs’ internal encoding and external behavior: they may encode the correct answer, yet consistently generate an incorrect one.” The method matters for reading that honestly: they resample the model thirty times and train a probe on internal states to select among the candidates the model itself produced, and the probe outperforms the model’s own decoding. So the system is not withholding a fully formed answer. Information sufficient to discriminate true from false is present internally and does not govern the output. The same paper found these probes fail to transfer across datasets, so “truthfulness encoding is not universal but rather multifaceted.”

Why the output side is pushed toward confident error has a separate and deflationary explanation. Kalai, Nachum, Vempala and Zhang argue in “Why Language Models Hallucinate” that errors arise from ordinary statistical pressure during pretraining and persist because benchmarks score a wrong guess above an admission of uncertainty. “Language models are optimized to be good test-takers, and guessing when uncertain improves test performance.”

Now put that beside Aquinas. In his account the known exists in the knower according to the mode of the knower, and the intelligible species is that by which the intellect understands, not the object it stares at. Knowledge is not a stored correlate of a fact. It is the knower being informed by the form of the thing, and able to return to it, judge it, and be answerable to it as true.

A system whose internal states carry information correlated with truth, whose output is governed by an objective that rewards fluent guessing, and which cannot reliably report which is happening, sits oddly with that account. It plainly has structures sufficient to reproduce many behaviors associated with knowing. Whether those structures constitute knowledge, and whether there is a knower, is not something the experiments answer. It is what they expose. Keyon Vafa and colleagues put the empirical version well at NeurIPS 2024: models can pass every existing diagnostic for having a world model while the world model itself, once reconstructed, turns out to be incoherent enough to break under small perturbations.

A compact comparison

AquinasModern AI
Every agent produces something like itself, so the effect bears the likeness of the agent’s form (ST I, q. 4, a. 3)Systems trained on human intellectual artifacts produce human-like intellectual artifacts
A perfection can exist in the cause and in the effect according to different modes, univocally or eminently (q. 4, a. 2)AI reproduces intellectual behaviors without demonstrated human-like subjective understanding
Operation follows being; what a thing does is evidence about what it is, but the inference must respect the mode (q. 75, a. 2; q. 79, a. 1)Behavioral benchmarks cannot by themselves establish consciousness
In artifacts, the form pre-exists in the maker as an object of thought, secundum esse intelligibile (q. 15, a. 1)AI depends on data, architecture, objectives and evaluation criteria originating in human activity
Creatures resemble God by analogy, not by agreement in genus or species (q. 4, a. 3 ad 3)AI resembles human intelligence by derivation from human artifacts, not by sharing a nature

The right-hand column is not derived from the left. The left-hand column supplies the vocabulary that makes the right-hand column’s ambiguity visible.

Three questions that keep getting merged

The credibility of any argument here depends on refusing four claims. Aquinas did not prove that AI can never think. Science has not proved that AI cannot become conscious. LLMs are not “just autocomplete,” a description that was inadequate in 2020 and is now simply false about systems trained with reinforcement learning on verified reasoning traces. And poor performance today establishes nothing about permanent metaphysical impossibility.

What is actually going on is that three questions keep getting collapsed into one.

The empirical question is what capabilities present systems exhibit. That is measurable, it is being measured, and the measurements move fast.

The philosophical question is what understanding, knowledge, intentionality and consciousness consist in. That is unresolved, and unresolved for humans too.

The metaphysical question is what kind of being would have to exist for intellectual activity to exist formally in a subject rather than analogically or instrumentally in an artifact. Question 4, article 2 does not answer this. It shows why the question is separate. Aquinas’s own arguments about the immateriality of intellect come later, in ST I, qq. 75–79, where the claim that the intellect has an operation per se apart from any bodily organ is argued from the universality of its object, not assumed.

What the effect tells us about the cause

We look at these systems and marvel: look how intelligent the machine has become. Aquinas’s causal perspective suggests a different question. What does the effect tell us about the intelligence of its causes?

Human beings have built artifacts capable of reproducing enormous portions of the external operations of rational intelligence. Language, mathematics, formal reasoning, programming, scientific analysis, music, creative production. We did this by encoding at scale the structures that human minds produced over millennia, then discovering that those structures contain enough regularity to be learned. The training corpus is a partial record of human intelligibility, and the fact that so much can be recovered from it is a statement about the corpus and its authors before it is a statement about the architecture. That is extraordinary, and I do not think we have absorbed it.

It also points upward in the way Aquinas would expect. AI bears a likeness to human intelligence because human intelligence produced everything it was made from. Human intelligence, on his account, bears a radically higher and dimmer likeness to Divine Intelligence, and he insists that this likeness is by analogy, not by sharing a nature. The analogy should not be flattened at either end. In Thomistic anthropology a human being is a rational substance capable of grasping universals, willing goods, judging truth, and reflecting on its own acts. Whether an artificial system could ever be an intellectual subject in that sense is a far larger question than anything current evidence settles.

What current AI demonstrates is not that we have created such a subject. It is that we have become extraordinarily good at reproducing the operations and the external signs associated with one.

The more perfectly the effect resembles intelligence, the more carefully we have to ask what kind of perfection the effect actually possesses. We keep measuring the operation and assuming we have discovered the nature of the operator. Aquinas reminds us that these are two different questions.


References

Primary texts

  • Thomas Aquinas, Summa Theologiae I, q. 4, aa. 1–3. Latin: Corpus Thomisticum. English (Dominican Province translation): New Advent, CCEL.
  • Thomas Aquinas, Summa Theologiae I, q. 15, a. 1, on the exemplar in the mind of the maker: Latin, English.
  • Thomas Aquinas, Summa Theologiae I, qq. 75–79, on the soul and the intellectual powers: q. 75, q. 79, and q. 85 on abstraction: q. 85.
  • Thomas Aquinas, Summa contra Gentiles I, c. 29, on likeness between creatures and God: aquinas.cc.

History of machine consciousness and cognitive modeling

  • Reggia, J. A. (2013). “The rise of machine consciousness: Studying consciousness with computational models.” Neural Networks 44: 112–131. PubMed
  • Gamez, D. (2008). “Progress in machine consciousness.” Consciousness and Cognition 17(3). PDF
  • Baars, B. J. & Franklin, S. (2007). “An architectural model of conscious and unconscious brain functions: Global Workspace Theory and IDA.” Neural Networks. PDF
  • Brooks, R., Breazeal, C., Marjanović, M., Scassellati, B. & Williamson, M. (1999). “The Cog Project: Building a Humanoid Robot.” PDF
  • Eliasmith, C. et al. (2012). “A Large-Scale Model of the Functioning Brain.” Science 338: 1202–1205. PDF
  • ACT-R project overview, Carnegie Mellon University. Link

AI consciousness

  • Butlin, P., Long, R., Elmoznino, E., Bengio, Y., Birch, J. et al. (2023). Consciousness in Artificial Intelligence: Insights from the Science of Consciousness. arXiv:2308.08708. Peer-reviewed version: “Identifying indicators of consciousness in AI systems,” Trends in Cognitive Sciences, 2025. eScholarship
  • Seth, A. (2025). “Conscious artificial intelligence and biological naturalism.” Behavioral and Brain Sciences. Cambridge Core
  • Schwitzgebel, E. (2026). AI and Consciousness: A Skeptical Overview. Cambridge Elements in Philosophy and AI. PDF
  • Birch, J. (2024). The Edge of Sentience. Oxford University Press, open access. OUP
  • Bayne, T., Seth, A., Massimini, M. et al. (2024). “Tests for consciousness in humans and beyond.” Trends in Cognitive Sciences 28(5): 454–466. Record
  • Cogitate Consortium (2025). “Adversarial testing of global neuronal workspace and integrated information theories of consciousness.” Nature 642: 133–142. Nature
  • Findlay, G., Marshall, W., Albantakis, L. et al. (2024). Dissociating Artificial Intelligence from Artificial Consciousness. arXiv:2412.04571 (preprint)
  • Gurnee, W. et al. (2026). “Verbalizable Representations Form a Global Workspace in Language Models.” Anthropic. Transformer Circuits; commentary by Butlin, Shiller, Plunkett & Long, Eleos AI

Reasoning: progress and limits

  • Stanford HAI, AI Index Report 2025 and 2026. 2025, 2026
  • Google DeepMind (2025). “Gemini with Deep Think officially achieves gold-medal standard at the IMO.” Blog
  • Shojaee, P., Mirzadeh, I., Alizadeh, K., Horton, M., Bengio, S. & Farajtabar, M. (2025). “The Illusion of Thinking.” NeurIPS 2025. Apple, arXiv:2506.06941
  • Lawsen, A. (2025). “Comment on The Illusion of Thinking.” arXiv:2506.09250 (preprint)
  • Dellibarda Varela, I., Romero-Sorozabal, P., Rocon, E. & Cebrian, M. (2026). “Rethinking the Illusion of Thinking.” Artificial Intelligence XLII (SGAI-AI 2025), LNCS 16301, Springer. Chapter
  • Mirzadeh, I. et al. (2025). “GSM-Symbolic.” ICLR 2025. Poster
  • Wu, Z. et al. (2024). “Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks.” NAACL 2024. ACL Anthology
  • Chen, Y. et al. (2025). “Reasoning Models Don’t Always Say What They Think.” Anthropic. arXiv:2505.05410

Generalization and ARC

  • Chollet, F. (2019). On the Measure of Intelligence. arXiv:1911.01547
  • ARC Prize. “OpenAI o3 Breakthrough High Score on ARC-AGI-Pub” (2024). Blog
  • ARC Prize. “ARC Prize 2025 Results and Analysis.” Blog
  • ARC Prize. ARC-AGI-3 Technical Report (April 2026). PDF
  • LeGris, S. et al. (2024). “H-ARC: A Robust Estimate of Human Performance on the Abstraction and Reasoning Corpus Benchmark.” arXiv:2409.01374
  • Mitchell, M. (2024). “Did OpenAI Just Solve Abstract Reasoning?” AI: A Guide for Thinking Humans
  • Akyürek, E. et al. (2024). “The Surprising Effectiveness of Test-Time Training.” arXiv:2411.07279

Representation, hallucination and knowledge

  • Orgad, H., Toker, M., Gekhman, Z. et al. (2025). “LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations.” ICLR 2025. arXiv:2410.02707
  • Kalai, A. T., Nachum, O., Vempala, S. & Zhang, E. (2025). “Why Language Models Hallucinate.” arXiv:2509.04664
  • Marks, S. & Tegmark, M. (2024). “The Geometry of Truth.” COLM 2024. arXiv:2310.06824
  • Levinstein, B. A. & Herrmann, D. A. (2025). “Still no lie detector for language models: probing empirical and conceptual roadblocks.” Philosophical Studies 182(7): 1539–1565. Springer
  • Vafa, K., Chen, J. Y., Rambachan, A., Kleinberg, J. & Mullainathan, S. (2024). “Evaluating the World Model Implicit in a Generative Model.” NeurIPS 2024. arXiv:2406.03689
  • Mitchell, M. & Krakauer, D. C. (2023). “The debate over understanding in AI’s large language models.” PNAS 120(13). arXiv:2210.13966
  • Yildirim, I. & Paul, L. A. (2024). “From task structures to world models: what do LLMs know?” Trends in Cognitive Sciences 28(5). PDF — with the reply by Goddu, Noë & Thompson, “LLMs don’t know anything: reply to Yildirim and Paul,” TiCS 28(11).

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.