Attention as the Core of Cultural Machines.
From Structuralism to Transformers
Much of this serves as an early thought, in a bid to spend more time reading and detailing things I find fascinating. Naturally, I have come across the 2017 paper, “Attention Is All You Need,” by Vaswani et al. This paper argues that paying attention to what's relevant is more important than looking at the full picture. Sometimes we don't really know who the audience is or where to file something until we come across it.
I almost did not write this. Part of spending so much time online and being a keen observer of culture, internet culture and cycles, is that I have also taken time to see how blog culture, and online spaces have changed, some for good but rather spaces that are meant to be inherently for sharing thoughts ideas and even teachings to some extent have become an open marketplace that demands a certain style of format to allow those thats use them to derive some form of monetary value, but what does that say where one becomes spread thin against multiple platforms trying to find a home for a few fleeting thoughts on a sunday afternoon. In addition to this, I may have spent a good 40ish minutes of my time with lo-fi beats playing in the background, promising to improve my concentration whilst I thought of the perfect opening for this essay, and by the time I finish, I'll figure out that it's probably in the middle of the text.
Back to not getting off the subject, yes, the concept of attention seems the natural progression of things. Why should one spend countless hours pondering every single detail when we could just hyper-focus on the ones that lead to the desired outcome?
The paper itself is deceptively modest in its ambitions. Vaswani and colleagues were solving a specific problem: how to process sequence language, primarily, without the bottlenecks that plagued earlier architectures. Recurrent networks read words one at a time, like someone mouthing a sentence slowly, unable to hold the beginning clearly by the time they reach the end. The Transformer abolished that. Instead of a sequence, it proposed a relation. Every word, simultaneously, in computing its weighted relevance to every other word in the sentence. Not reading left to right, but reading the whole thing at once, dynamically, attention is distributed and redistributed across the entire field.
They called the mechanism self-attention, and they titled the paper as though making a small technical claim. Attention is all you need. It reads like an engineer’s shorthand. It landed, eventually, like a philosophical proposition. To me, it showed how things have always worked. The structuralists, for example, Saussure argued at the turn of the twentieth century that language cannot be viewed as a mere collection of words with fixed inner essences; rather, we should view language as a system of differences, somewhat of a network where each element derives its meaning from its relation to every other element. Not “what a word contains” but “how a word sits in relation.” Meaning as position, not possession. Self-attention is, in a sense, that same operation, just written down as an equation.
This then led me to pull out my copy of Leif Weatherby’s Language Machines: Cultural AI and the End of Remainder Humanism. The book was published in 2025, and for what I'm trying to discuss, it has given me a sort of loose framework. Weatherby’s argument is that LLMs aren’t simulating human minds and never were. Instead, the book introduces the idea that LLm’s are producing culture, and that is not hard to deny, perhaps more now, the internet and its spaces seem to have been invaded by carefully worded clones. What the book brings forward is that LLMs have allowed for the capture of the generative, relational, formal side of languages, these parts that have always operated beneath authorial intent, individual expression and now for running it forward at scale. The debate about AI and human cognition that followed has been loud and largely circular. I myself, despite having spent considerable time studying transformers and language models, still find myself circling the question: can it really think? Does it truly understand? Where is the inner life? Weatherby sidesteps all of it, and honestly, that is the more interesting move. Is language uniquely human, and if so, what has it done for us as a species? Can machines truly think the way we do?
Post-structuralism spent decades trying to say this. Barthes declared the death of the author in 1967, not as provocation for its own sake but as an attempt to locate meaning where it actually lives, which is in the text’s relations, its echoes, its conversations with everything written before it. Derrida followed the same thread; there was no hors-texte, no outside-the-text, and meaning was endlessly deferred through a network of differences that never fully resolved. Foucault asked not who wrote something, but what function the author-figure served, what it organised and contained.
None of these ideas felt comfortable when they arrived, and like most new concepts, such as AI is going to take all of our jobs, to many, it felt like an assault on the human. What the Transformer architecture did (in Vaswani’s paper), without meaning to, was demonstrate them as engineering requirements.
What self-attention actually does, stripped of the notation, is almost embarrassingly intuitive. Every word in a sentence turns to every other word and asks: how much do you matter to what I’m trying to mean right now? The answers come back weighted, with some connections strong and some negligible, and meaning assembles itself from that negotiation rather than from any fixed definition. The word “bank” doesn’t know what it is until it looks around at its neighbours. River, or money. The sentence decides.
I guess what I'm trying to get at with all this, and what has become a long way of saying it, is it was how familiar it felt. This is what reading does. What conversation does. What culture does. We’ve always been in this negotiation, weighing relevance, adjusting meaning in relation, debating, assessing and deciding; we just never had to write it down as an equation before.
This is probably a good place to stop, not every thought needs a home, a format, JSON input or a monetisation strategy. Some of them just need a Sunday afternoon, a paper you weren’t sure you’d understand, and the particular pleasure of finding that one thing connects to another in a way nobody asked you to notice, Saussure via a machine learning paper, cultural theory via engineering shorthand. There is something almost too neat about writing an essay on relational meaning that only found its shape in relation to each idea pulling the next into focus, the conclusion somewhere in the middle, as I half-knew it would be. So here it is. Filed under: things I found fascinating and couldn’t help myself.

