Virtual Bookshelf

I always relish the opportunity to look at someone’s bookshelf. What books are so important to someone (or so central to how they want to be seen by visitors) that they keep them in their living room for all to see? Even better if they can tell me something about what that book means to them.

The following are a few of my ‘keepers’: books, articles, medieval commentaries, blog posts, etc. that I read a long time ago (at very least a few months) and still think about frequently in my academic life. It doubles as an index of the high-level topics that interest me currently.

Language, Culture, and Concepts

The Origin of Concepts (Susan Carey, 2009). I started reading this book immediately after I listened to Tom Griffiths’ interview with Carey. In the interview she talks about her cognitive-historical work studying conceptions of thermal phenomena prior to the distinction between heat and temperature. This is exactly the sort of philosophy/cognitive science/intellectual history mashup that hooks me every time. The first part of the book deals with innateness. Prior to reading Carey, I was a pretty committed empiricist; I saw no reason why anything beyond some perceptual primitives (edges, faces) and extremely general learning paradigms (e.g. predictive coding; see below) couldn’t be learned during early development. I had never engaged with the question deeply before, but the best arguments I had heard for innate core cognition were of the form ‘six-month-olds can make moral judgments, so moral reasoning must be innate.’ This is pretty weak - six months is a long time, and you can’t make a priori assumptions about what’s hard or easy to learn. Now after a few hundred pages of Carey, I’m convinced. Carey is not just extremely rigorous, she is also refreshingly clear about precisely what ‘innate’ means and what the various philosophical positions on the topic are (or were, as of a couple decades ago). Regarding incommensurability - the sort of qualitative / structural difference between e.g. the way a toddler understands ‘number’ and the way you do - I was roughly familiar with the concept from Thomas Kuhn and from observations of religious people with wildly different definitions of ‘God’ or ‘authenticity’ (a healthy dose of Alison Gopnik over the last few years didn’t hurt either). But Carey really drove it home for me how many processes of persuasion or opinion change are not matters pushing a sliding scale from “slightly agree” to “strongly disagree” or from 23% probability to 36%. The beliefs that admit to degrees like this are themselves concepts, and the way we draw boundaries around concepts can change dramatically within individuals, not to mention between them. This is a real problem for the parts of my research that treat beliefs as purely a matter of degree, and it bothers me.

Meaning Making Under the Sacred Canopy: The Role of Orthodox Jewish Marriage Guidebooks (Novis-Deutsch & Engelberg, 2012) and The Fourth Book of Maccabees in a Multi-Cultural City (Tessa Rajak, 2016). As an undergrad, these two examples helped me see a pattern: ideologues declare a culture war while being so deep in their opponents’ culture that they can’t even see the irony. IV Maccabees was written by a Greek-speaking Jew in the Eastern Roman Empire (perhaps Antioch, Syria). It is about how the traditional Jewish way of life is much better than anything those Greek philosophers could dream up, because… wait for it… the Torah is designed to foster the [extremely Stoic] virtues of self-discipline and control over the emotions. Greeks in the book are portrayed as womanly and emotional (the author clearly thought this was a bad thing) whereas Jews are, well, Stoic. In ~2,000 year retrospect, this is hilarious. Meanwhile, traditionalist Jews (i.e. my people) are still at it, composing marriage guidebooks that deplore the shallowness of modern Western culture and proclaim that traditional Jewish marital laws are much better because… wait for it… the Torah designed to foster the [extremely modern Western] virtues of emotional depth, mutual respect, and intellectual partnership in romantic relationships. I don’t mean to disparage these ironic apologetics—I am fully in favor of reframing old material in new terms, and cultural exchange is great. And none of this is unique to Jews of course. The lesson here is that ‘culture wars’ are often waged between groups that mostly share the same culture.

Language as social action (Thomas Holtgraves, 2001). I read this textbook in my last year of undergrad at the recommendation of my mentor, Bruno Galantucci. Holtgraves in turn lead me to read Searle’s Speech Acts and to think seriously about the relationship between language, culture, and cognition. I’m indebted to Holtgraves for teaching me that modern textbooks can be well-written, and (along with Dr. Galantucci) for introducing me to the intersection of philosophy, psychology, cognitive science, linguistics, and sociology that has continued to draw me ever since.

Word Embeddings. word2vec, GloVe, fastText. Each of these algorithms is a bit hacky when you get into the details, but I just can’t get over word embeddings. Such beautiful little models that can nevertheless capture so much of the complexity of language. A big part of the allure is just the vision of meaning arranged as a point cloud in high dimensional space — credit for that should probably go to Roger Shepard. But the way word embeddings encode distributional statistics so directly deserves more attention; I think it’s a shame that the NLP community abandoned these simple models so quickly after the rise of transformers. My two favorite papers on the theory of word embeddings are Improving Distributional Similarity with Lessons Learned from Word Embeddings (Levy et al., 2015) and Symmetry in language statistics shapes the geometry of model representations (Karkada et al., 2026). Regarding the latter, J.J. Gibson should probably get a shout-out here: Learning systems are always reflections of their training data/environmental affordances. LLM internal states are mostly just compressed language statistics, European colonization is in some sense compressed biogeography (along with a fair bit of chance; see below on Diamond), and probably most cognitive biases are calibrated to environmental patterns (see Sam Gershman’s artful Bayesian brain apologetics).

Course in General Linguistics (Ferdinand de Saussure, 1916). I arrived at Saussure by way of Chandler’s primer on Semiotics. A lot of the literature associated with Saussure and semiotics is quite silly (especially post-Derrida) but Saussure himself was astoundingly sensible while generally remaining vague enough to comfortably accommodate new paradigms like my own beloved semantic embedding spaces. Actually, apart from his bigoted view of text as subordinate to spoken language, I suspect that Saussure is a better theoretical fit to modern work on the geometry of token embeddings even than Zellig Harris, the structuralist who is canonically cited as the father of distributional semantics (but I’ll reserve full judgment for after I get around to reading more Harris). My favorite Saussurean precepts: language belongs to the speech community, sense is mutually dependent on all other words in the lexicon, the paradigmatic axis is real (and now with LLMs, we can actually quantify it).

Responsum of R.Avraham ben David (a.k.a. Ra’avad, 12th century Provence) as reported by R. Asher ben Yeẖiel (a.k.a. Rosh, 13-14th century Rhineland/Castile) Yoma 8:14, R. Shlomo ibn Adret (a.k.a. Rashba, 13-14th century Catalonia) Responsa Vol. 1, §689 and subsequent legal scholars. A particular example of a ubiquitous phenomenon in any information ecosystem: conceptual shapeshifting. This example is the medieval prehistory of the now-classic debate about whether a life-threatening situation suspends the prohibitions of Shabbat (i.e. they do not apply to such a situation in the first place) or merely overrides them (i.e. they apply but preserving life takes precedence). Studying this literature lead me to an intimate appreciation for how two people/communities can talk about something in similar terms but have in mind an entirely different conceptual framework. The suspend/override dichotomy certainly did not exist at all for Ra’avad, whereas Rashba seems incapable of reading the former as talking about anything other than this dichotomy. The modern literature is filled with people like this, who try to force each medieval or ancient authority into one or the other camp despite the fact that most of these authorities would certainly have had no idea what you were talking about if you could ask them about it. Even if they did understand you, they might have a very different conceptual framework in mind; R. Yisha’aya di Trani (a.k.a. Rid, 12-13th century Italy), who as far as I can tell is the originator of the dichotomy’s characteristic terminology, clearly did not consider it to have the same implications as later scholars did, since he explicitly claims that the prohibitions are ‘suspended’ but then makes a ruling that would canonically be characteristic of the ‘override’ position. I’m not complaining—one man’s misunderstanding is another’s intellectual innovation. The point here is that ideas are slippery things, especially when they are expressed in succinct language by people inhabiting different speech communities or schools of thought. It is not uncommon that an idea flows from one community to another and morphs incommensurably in the process, even when the communities involved are explicitly aware of the intellectual influence. The suspend/override dichotomy happens to be the most salient example for me, but the history of philosophy is likewise full of this, and the way mainstream news is often shared to promote conspiracy theories reeks of this kind of conceptual shapeshifting.

Causality, Explanation, and How to do Science

Statistical Rethinking (Richard McElreath, 2015) and accompanying lectures. McElreath’s course is more than just analysis skills - it’s a full epistemology and a manifesto about how science should be done. I worked through this course, along with Solomon Kurz’s translation into tidyverse + brms, in the year before I started my master’s degree. The likes of Dustin Fife had already opened my mind to the general(ized) (mixed) linear model. But McElreath taught me to think of models as pure manifestations of theoretical assumptions (“golems”), made me a Bayesian, and inducted me into the extended Judea Pearl fanclub: DAGs, Bayes networks, do calculus, etc. With such eloquent coverage of hierarchical Bayesian methods, it was a short leap from here to the predictive processing and Bayesian brain literature (see below) which I started reading around the same time.

Other amazing educators on Youtube: 3Blue1Brown (math), AlphaPhoenix (physics/material science), Andrej Karpathy (large language models), and a few others. I’ve been an avid consumer of educational YouTube since my early highschool days. Watching the very best of this genre along with the scores of wannabes and copycats (not going to name any names here) has led me to a firm belief in clear explanations as the central driver of scientific progress. Developing an intuitive explanation is not just a matter of production value or pedagogic skill, it is real scientific work. Thomas Kuhn talked about the ‘mopping up’ process of normal science as primarily about data collection and adding new domains of application to the dominant paradigm, but I think the slow slog of distilling clear explanations — the kind of work that happens on YouTube and in undergraduate survey courses — is at least as important. I thought about this a lot while writing my first textbook; as Feynman supposedly put it, “I couldn’t reduce it to the freshman level. That means we don’t really understand it.” The Feynman quote is spot on but it leaves out a critical part of how we reduce things to the freshman level: we infuse them into the broader culture (viz. YouTube) so that the freshmen know the basics before they even set foot on campus. This is the essence of science at the societal level. This is how we ease the burden of knowledge: we compress our best theories into ever simpler and more intuitive explanations so that scientists don’t have to spend their whole lives learning the background.

Vision: a computational investigation into the human representation and processing of visual information, Chapter 1 (David Marr, 1982) and 138 Openings of Wisdom (Qlaẖ Pitẖei H̱okhma; R. Moshe Chaim Luzzatto a.k.a. Ramẖal, 1785). The realization that a given algorithm can be instantiated by many different physical substrates and vice versa is a central one for modern cognitive (neuro)science. Marr’s particular division into three levels can be reasonably questioned, but the basic insight is timeless: Truth exists at multiple levels of abstraction, and the different levels need not seem compatible at first glace. As Hofstadter wrote in his Ant Fugue (see below), ants are communists, but ant colonies are libertarians. Ramẖal’s philosophical mysticism talks about this sort of thing all the time. So a realist metaphysics which sees causes/abstractions as real and manifestations/sensory signals as irrelevant is something more like ‘God’s perspective’, and a nominalist view which sees sensory signals as the only real thing and abstractions as useful fictions is called something more like ‘our perspective’, but both are legitimate in their domain (is the input layer more real than the last hidden layer of a deep neural network?) Inasmuch as most “things” are in the middle of the spectrum between abstract/causal/latent and concrete/physical/manifested, and since our mortal reasoning is capacity-limited and not infinitely deep, in practice everybody lives somewhere in between the two perspectives - disregarding certain things as too physical to be relevant (e.g. electrons, neurotransmitters) and certain things to be too abstract to be relevant (e.g. philosophical musings like these). Similarly, according to Ramẖal, organizing divine attributes by causal structure results in conceptual opposites being grouped together, while organizing them by conceptual similarity results in an incoherent jumble with regard to efficient causation - both perspectives (and many in between) are legitimate in their domain, but the theorist has to be careful to distinguish them. Also, a core tenet of both cognitive science and philosophical kabbalah: Many things are metaphors for (i.e. approximately isomorphic to) many other things.

Models of my Life (Herbert Simon, 1991). Simon was super cool and way smarter than I could ever hope to be.

4 3 2 1 (Paul Auster, 2017) and Guns, Germs, and Steel (Jared Diamond, 1997). I read Diamond early in high school, much to the surprise of my great uncle when he tried to lecture me about it at a family picnic. The particular thesis of the book was less important than the way of thinking: it was an introduction to the way real processes of causation can be extremely indirect, systemic, or circumstantial; the reason Europeans were able to colonize the Americas so quickly was about e.g. the geographic distribution of domesticable species, or perhaps the particular structure of political unrest in central America at the time. There is no better demonstration of this mess of causation than Auster’s semi-fictional autobiography in which four initially-identical versions of himself gradually drift apart due to tiny differences in their upbringing. Those of us who deal with statistical models know that various sources of ‘random noise’ always have to be accounted for, but it’s important to be reminded of the complexity of causation that might lie behind that noise.

Reinventing Knowledge: From Alexandria to the Internet (Ian McNeely and Lisa Wolverton, 2008). Read this in undergrad. Deserves credit as the first book that pushed me to think about the institutional structure of modern academia and the impact it has on the conceptual structure of scientific knowledge. Many of my thoughts on what universities might turn into in the post-AI era are in some way derived from this.

Gödel, Escher, Bach: an Eternal Golden Braid (Douglas Hofstadter, 1979). It feels cliché to put this on the list; I only picked up this book because I had heard so many scholars cite it as a formative influence on them. But it is popular for good reason. Hofstadter’s ability to explain esoteric subjects in a way that is both clear and entertaining - and without demeaning the reader - is something to be admired and emulated.

Predictive Processing, Bayesian Cognition, and Cybernetics

Resource-rational analysis: Understanding human cognition as the optimal use of limited computational resources (Lieder & Griffiths, 2019). Read this in undergrad at the recommendation of my mentor, Bruno Galantucci. It was my first introduction to why “are human’s rational” is a silly question but also a fruitful one. After reading this I started down the rabbit hole of approximate Bayesian reasoning, etc., and haven’t left it yet.

Surfing Uncertainty (Andy Clark, 2016). The best primer on predictive processing I’ve read so far, which is frankly a pretty low bar. See above on how if we can’t explain it clearly, we don’t understand it well enough yet. Hierarchical predictive processing (not to mention the Free Energy Principle) is sublimely parsimonious but infuriatingly underspecified - much like the theory of evolution by natural selection.

Active Inference (Thomas Parr, Giovanni Pezzulo, and Karl Friston, 2022) and related writings by Friston. Plus credit to Darwin. See my essay on Friston and Epicurus.

Behavior: The Control of Perception (William Powers, 1973). Read this on the recommendation of Paul Cisek. I highly recommend reading Chapters 1 and 4. The approach is very similar to Fristonian hierarchical predictive processing, except that 1. the writing is clearer (especially about the implications for the way psychology should be taught), 2. “purpose” or “reference point” are used instead of “prediction” (which makes more intuitive sense, I think), 3. Powers doesn’t deal much with uncertainty, and 4. it was written in the `70s.

Trapped Priors As A Basic Problem Of Rationality (Scott Alexander, 2021). Definitely reflects a partial misunderstanding of (hierarchical) Bayesian inference, but I think there’s something to this nevertheless. More pointedly: what kind of deviation from perfect Bayesian inference might result in patterns like these? This article has been in the back of my mind for a couple years now.