Can LLMs finally do social science?
A tool for tracking latent cultural concepts over time
After a year of iteration, I built a formalized prompt (and accompanying website) to quantitatively observe latent cultural concepts over time. The example datasets on the site represent tens of hours of Gemini output, which are pre-generated because each decomposition takes significant compute. I’ve wanted to see something like this for a long time, but existing ngram-style approaches, which only count words, don’t reveal deeper cultural abstractions. LLMs should be good at this: when an LLM ingests the entirety of human cultural output, it compresses this information into a latent space, which feels remarkably powerful as a tool for the social sciences. Only in the past few months have frontier models become strong enough at following these instructions that it seems worth sharing.
The concept is relatively simple: with a formalized prompt, we can attempt to extract latent clusters over time for abstract concepts that feel real and we can talk about, but which no algorithms have been able to meaningfully tackle. It’s easier to first play around with the website, then understand how it works.
There are a lot of examples I’ve generated, but since you’re reading this you may find the decomposition of the modern Rationalist Space or Foundation Models interesting. If you’re interested in culture, you can track the cultural factors in America over the past couple decades. One of my favorites and most unexpectedly interesting decompositions is regarding WW1 and WW2 military strategy. The rest reflect my own eclectic interests.
This is a week-by-week breakdown of Covid-era clusters. The model decided ‘Bio-Political State of Exception’ was the primary cluster, followed closely by ‘Epistemic Fragmentation.’ They tracked closely, but the end of Covid didn’t mark the end of broken shared epistemics.
These cultural clusters are abstract concepts without a ground truth, and vary based on the prompt. In many cases they are not even fully replicable based on the seed. Despite this, my claim is that a sufficiently well architected prompt is still a scientifically valid object.
Within the current social sciences and humanities, we already attempt to track changes in culture and abstract phenomena. Historically, though, this effort was restricted to the domain of an individual mind, and modeled through long essays with esoteric references.
A core point of this new method of prompting, and this essay, is that defining arbitrary units of salience and tracking them over time quantitatively is not scientifically distinct from, say, an essay on the Covid era written by a human that reads “A bio-political state of exception rapidly took off with the outbreak of the pandemic within the US.”
So while it could be argued that the units are ‘arbitrary’ and ‘LLMs are a black box’, I strongly believe that they are (much) easier to observe, introspect, reproduce and tinker with relative to a human brain.
How does it work?
After the primary prompt, I provide only three things:
The concept I want decomposed (e.g., “Contemporary American culture,” “Military strategy 1890-1950”) as well as any other conditioning information.1
A time window (start year, end year)
Periodicity (e.g., every 2 years)
Next, the model dynamically generates the top clusters, as well as the salience of each cluster over time. The clusters are determined by the language model. Within the prompt, I impose a prior, as well as formal notation, for what counts as a good decomposition: abstractness, temporal coherence, distinctness, and grounded sources. Salience is only meaningful relative to time and other clusters.
If this feels unsatisfying, consider how we think about cultural salience or explanation today, which is even less rigorous. With a quantitative measure, we can discuss and disagree in a more measured discussion. The numbers don’t represent truthiness; they are an instrument for humans to effectively communicate.
Prompt Properties
The view of latent culture I impose is that it must be based on methodological individualism. Methodological individualism defines culture as an ‘emergent property’ of individual actions. I then encourage the model to focus on abstract concepts as they arise in philosophy of the social sciences. Concepts like “nihilism” and “post-modernism” are difficult to pin down, yet also have real meaning. They evade typical modeling as they are effectively impossible to write down as clean functional or statistical models.
In addition to grounding it in my preferred epistemic framework, I also use the formal language of statistical cluster analysis and time-series in the prompt, though these are more guiding concepts. It is of course not being run in the way a traditional algorithm is executed2. When I analyze culture over time, these are the statistical properties I try to mimic in my own reasoning, so I offer the same framework to the prompt.
Lastly, the model is disciplined to provide supporting evidence and primary sources to justify its choice of salience. This provides instrumental observability to the user and grounds the choices empirically.
Is this real?
All cultural dimensional reduction needs to be viewed as a category made for man, not an attempt to estimate an unseen true latent value. This is a critical difference. When we estimate where a satellite is, the satellite has a real position. We’re just uncertain about it. Our measurement is trying to converge on something that exists independent of us. But “postmodernism” isn’t like that. There’s no true postmodernism coordinate hidden in the universe that we’re trying to estimate.
Instead, we’re constructing a useful category that helps humans talk about and navigate cultural patterns. The question isn’t “did we find the right answer?” but “is this instrumentally useful to carve up the space?” Instrumental usefulness could mean anything from a descriptive tool for researchers to a component in a formal prediction algorithm.
Having a language model that spits out abstract “salience units” and isn’t built by any clear functional formula doesn’t currently fit into a quantitative scientific framework. Even in the social sciences, there is a strong focus on statistical or econometric identifiability, determining whether the data-generating process pins down a unique value of the parameter(s). This prompt certainly does not satisfy this property.
While I am a strong believer in identifiability, I also find that level of demand for rigor incongruent with how humans actually interact with the world. As a provocative question, consider the blogs of famous economists, or the opining on X.com from academics. Their formally published papers have a million robustness checks, yet their opinions expressed in blogs or discussions grapple with much harder to pin down concepts. Their “solution” is to isolate these thoughts to less formal venues. I find this unsatisfying. Shouldn’t all their beliefs about the world be subject to the same epistemic standards?
I think we should reframe this divide not as a spectrum from rigor to softness, or from the real to the fake. Rather, it’s about the compute resources capable of solving these different problems in the social sciences. Suppose we have a toy-model of the world, where there are only CPUs and biological processing units (our brains). Economists rely more on CPUs to do the actual processing and estimation of data, whereas sociologists (currently) often ‘estimate’ their models using their own brain.
We often consider cases where estimation is done algorithmically more ‘real’ in the social sciences. But are they? Formalization permits us to identify our exact unit of data, our exact mathematical model of the world, and by estimating it using scientific software, it (in theory) gives us reproducibility. These are good things.
On the other hand, formalization heavily restricts the space we can analyze, and the manner in which we can analyze topics, which is why we often still make our final decisions by estimating with our brains.
Still, I don’t think we can, or should, avoid the question of evaluation. As an example, early versions of this prompt that I ran on OpenAI’s 4o model were so ‘bad,’ as defined by my eyes, that I didn’t share them. So there must be some basis for what it means for an abstract cluster to be correct.
If you look at these plots, many clearly seem to track something real. This is a funny one from WW1 and WW2 military strategy clusters, if for no other reason than it’s so obvious, with a drop in the military cluster on static defense following the Blitzkrieg obliterating the Maginot Line. This seems true, in the sense that the complete failure of the defense system over a week put an end to the military culture of static defense.
Comparison to Old approaches
In 1882 Nietzsche made a now-famous observation on the decline of God as a salient force in Western society, and a forecast on the implications this would have. An older algorithmic approach in the paper Quantitative Analysis of Culture Using Millions of Digitized Books captures the same pattern: the decline of God over time by word count across millions of books.3
If we direct this prompt to focus only on God, it unsurprisingly has a similar shape to the semantics, and also references the textual works that most strongly capture the cultural collapse of God. We do narrowly lose specific insight into the algorithm here, as the plot of word count over time is a tightly scoped algorithm. However, we gain the ability to have a language model ‘explain’ to us not only the word count, but how and why it believes the salience is declining as well as citing sources.4
There are a few more examples in the appendix.
What even are Cultural Abstractions?
A useful starting question: what exactly is a Marx or what is a Hume, Orwell, or Tolstoy? Are they individuals, authors, or models themselves?
Consider Das Kapital as an example, which is fundamentally an extremely high parameter model on the intersection of class and economics, also known as a book. Das Kapital was written as text, and we obviously read it as text, but it is isomorphic to a complicated function. We simply have not evolved to write and read such functions. Yet despite Das Kapital being entirely written text, huge amounts of people still wish Marx was alive today to get his opinion on contemporary issues.
Das Kapital presents a theory of how capitalism works, such as how commodities relate to labor, how surplus value gets extracted, and how class dynamics unfold. When you internalize the book, you can take a new situation Marx never saw (say, gig economy labor or cryptocurrency) and generate a Marxist analysis of it. The book is text, but functionally it’s a mapping: input a description of an economic arrangement, output an analysis of its class dynamics and internal contradictions. That’s what a function does. We don’t typically think of books this way because we evolved to process written language, not mathematical objects. But the structure is there. Das Kapital has “parameters” (its core concepts and their relationships) that were “trained” on 19th-century economic observations and can generalize to new inputs. The reason people still wish Marx were alive to comment on contemporary capitalism is that the function running in his head was higher-fidelity than the lossy compression into text. The book is an interface to the function, not the function itself.
The implication is that these unique individuals, in their wetware neural networks, have a specific parameterization that is not so easily replicable. We can’t write it down, yet there is something there. Why is this the case? If all these historical books, or cultural analyses, have been written, why do we still yearn for the humans themselves? Or even among contemporary academics, why is it that if you ask ten academics to map out the rise of post-modernism, you would get ten different answers? Yet there would be a commonality behind all answers. There would be some sort of gesturing towards individual rejection of modernist values, and a 1960s academic culture that questioned (or rejected) the established scientific and nationalistic beliefs in the preceding century.
This is a reflection of the information-theoretic limits of compressing cultural complexity. Consider how historians still debate the primary causes of World War I. This isn’t because history is non-empirical, as I’ve written about before in History as an Empirical Science and WWII & Convenient Causality, but because the event exists in an extraordinarily high-dimensional causal space that resists simple reduction.
We have historically held one standard for what makes the social sciences rigorous, but then admit a revealed preference that we don’t actually believe they can be, because we’re fascinated by books and writing where it’s impossible to achieve the lofty goals of statistical or causal identification.
Against Quantitative Segregation: Use LLMs
I believe the reason is institutional segregation that no longer makes sense. This specialization of skills has been mistaken for an intrinsic difference in how we understand the world. Statisticians optimized methods for data that fits neatly into matrices, while humanities scholars processed unstructured information too complex for our previous computational frameworks. It’s a fundamental error to think there are “different epistemologies” at work. It’s all about extracting signals from information to predict the future and understand the past.
Using language models here is an attempt to push into territory of analysis that previously could only live in the brains of humans. The brightest individuals in these fields, such as Political Theory, Cultural Theory, Literary Criticism, or Historiography, have always been performing dimensionality reduction on high-dimensional cultural data, only without proper formalization or rigor. Not because they’re lazy, but because it wasn’t clear how to formalize these fields.
A large part of my motivation to start what is ultimately just a proof-of-concept for research is a frustration in not seeing the academic and empirical social sciences building a methodology to treat the outputs of language models as legitimate scientific artifacts (or worse, doing so poorly).
The one area of Social Science that has succeeded in formalizing this approach mathematically and statistically is Economics. Historically, Economists were viewed as rigorous and ‘real,’ and other social sciences were viewed as ‘fake.’
From Economics
The compression in language models represents an object we can study that is not unlike markets, or economic measurements.
There is a similarity between how these models compress abstract tokens into a coherent representation of our latent recorded information, and how markets compress beliefs and the state of the economy into prices and indices. While I want to be careful not to stretch this comparison too far, the general concept of the compression of human-generated data through a process we cannot fully observe, into an outcome we can observe, remains the same.
To build on this comparison, the prompt itself includes a formal structure, which I borrowed conceptually from Stock and Watson‘s work on dynamic factor models in economics (link to my statistical formalization here5). They faced a problem: the economy is driven by a small number of underlying forces (recession, inflation, confidence), but we observe hundreds of noisy indicators. Their insight was to extract latent factors using principal component analysis across a wide cross-section of measured economic variables, letting the structure emerge from the data rather than specifying it in advance. The parallel here is that culture is also driven by latent forces we can’t directly observe that arise from aggregate individual human behavior, and we’re trying to extract them from a high-dimensional space of texts and artifacts.
Consider that the interest rates on US treasury bonds, an area where I began my research career, embed the market’s belief from all participants on the state of the US economy. It is a profound emergent property where every individual and investor, from all of their beliefs and modeling of the state of all intricacies of our economy and government, come together to construct a single price at each moment in time.6
The field of Economics also tracks economic variables, which are abstract measurements we define and then take consistently over time, such as unemployment rates. The degrees of freedom to compress an entire economic system into a handful of time-series are massive. Despite this, we find treating these as real statistical objects useful and rigorous.
The difference, of course, breaks down in a number of ways. The process by which markets compress information is decentralized and made up of individual human signals, whereas the way we compress text is a function of tokenizations, model architecture, GPUs, and man-made compression algorithms.
I’ve spent years working on extremely well specified econometric models, which have outputs that change tremendously based on arbitrary changes in smoothing algorithms or minor changes in data definitions by upstream government bureaus. True epistemic stability in the social sciences is rarer than we would hope.
In this light, if we poke enough holes in Economics and understand that the field focuses on compression of high-dimensional data, the idea of defining a clear ontology and formalization for extracting cultural factors begins to feel real. Perhaps still not as real as Economics, but the comparison feels natural. People seem to think that language models are somehow more of a black box or less rigorous for use as scientific objects, which feels like an isolated demand for rigor to me.
Conclusion
Historically, we have made sense of mercurial cultural clusters through two unsatisfying methods:
The first is deterministic algorithms conditioned on samples of data. This is some sort of sentiment analysis that runs actual statistical learning models on real data. There is a likelihood function, and it is repeatable. It can track sentiment or word usage over time from some defined corpus.
The second is leveraging the specific power of the human brain to scan and understand vast amounts of text and concepts, and compress them into narrative and classification. We can squint at the past and observe variations in culture and human interaction and expression, often coming to profound insight. Yet this analysis lives inside the specific brain of a specific human, and it’s not truly replicable in any sense.
My goal is to see a new field of quantitative social science research where we attempt to graft the strengths of each into a new form of cultural analysis. Extracting quantitative measures of culture from an LLM is “fake” in the sense that it’s not a statistical algorithm or exact measurement. Yet it’s real in the sense that it’s a formalization of how to reason and process information.
The proof-of-concept I’ve built is meant to show that there is something interesting here for scientists with the motivation to poke around and conduct some novel research. As it stands now, it’s a descriptive tool using a formal prompt, not itself robust scientific output. Proving the latter will require building out additional scientific scaffolding to demonstrate external validity within specific domain-related analysis or projects. I think this is worth doing, and there is a new field of research for those willing to think outside of ossified institutional ideas of what the social sciences can be.
FAQ
Isn’t this just training on clusters other humans have identified and written down?
Yes, but we do this, too. Whenever an individual human attempts to identify a new cultural cluster, they do so with an existing knowledge of prior analysis that they have read and digested. Whether contemporary language models could be as prescient as the most culturally attuned humans is an empirical question.
My guess is probably not yet. Although in theory, again similar to other fields, it should be able to become much better by reading and ingesting otherworldly amounts of information daily.
Could this identify entirely new clusters only just emerging, that haven’t been reified by any human yet?
I’m not sure if it can today. It’s possible that it can, but it’s also possible that our current frontier models and their agentic use of search indexes is too nascent. In either event, this would require out-of-sample testing. It would also be extremely specific to the prompt and what you are trying to identify.
The obvious place to start is in-sample descriptive statistical regressions with robustness tests. The way this might work is to identify an interesting research question with a dependent variable that has a strong claim to truth. Then using time-series models, attempt to regress that on the clustered output of a specific prompt. Then run this 100s of times across models and on perturbations of a prompt.
The only way I can think to do this, truly, is with models trained on truncated data. Or in the absence of that, to create a prompt that identifies and robustly forecasts a specific culture, and to regress that against a directly measured outcome (e.g. economic measurements or birth rates). This could be used to build an out-of-sample time-series prediction. This is an area that needs additional research if trying to demonstrate external validity.
Is relying only on citations memorized in the parameters sufficient?
Not really. It’s sufficient for a proof of concept, but hooking a model up to a vector or lexical database of a historical corpus of texts would massively open up the research space. It could even permit potentially parametrized versions. Imagine a cluster analysis that takes days of compute to run, and every cluster is based on an observable weighted average of tens of thousands of historical documents.
Did you do any other lexical evaluation?
A little. From the same paper as above we can compare the semantic tracking of evolution, the cell, bacteria, and DNA. Despite differing methodologies, the LLM approach tracking latent conceptual importance and the other raw lexical frequency, the two models do exhibit temporal alignment. The inflection points across both datasets correspond with historical reality: both capture the late-19th-century surge in bacteriology, the specific early-20th-century dip in evolutionary discourse (the “Eclipse of Darwinism”), and the exponential rise of DNA starting in 1953. There are differences of course, but the cluster approach does largely capture timing and trajectory.
Does putting math into the prompt make sense?
There is some empirical evidence that language models can run real algorithms in-context. This paper shows that in-context regression can work. It also provides a formalization that is harder to describe with words.
I strongly reject the idea that formal notation requires a different burden of proof to use than textual. I also don’t claim using formal notation necessarily bestows additional validity. Why write paragraphs to gesture at a concept that has an existing formal statistical language we can use to express it, even if it’s not literally running as an algorithm?
Why prompting instead of architectural changes?
There is an element of ‘realness’ that is imposed, perhaps in part due to a demand for cleverness to publish scientific research, when any LLM work includes additional code or architectural modifications, rather than only prompting. This also is typically necessary to publish.
I think it’s plausible there are architectural modifications we could make here. A potential example would be adding temporal structure on the output, so time-series dynamics or recursion in clusters has some more structural form, rather than operating only in the text space. Again though, frontier models deal with text so well it’s not obvious this buys us anything.
But fundamentally it won’t solve the true conditional uncertainty problem. Even if you take the current prompt as correct, which you shouldn’t, there is also going to be variation in exactly how you prompt the language model. If you want to know about contemporary American variation, what if you use the word ‘Americana’, what if you try ten different versions of the sentence with different semantics, or ask for different start and end dates, etc?
Did you do any perturbation testing?
I created a prompt asking for the top clusters for America from 2020-2025. I then asked Gemini to write 7 semantically identical versions of that prompt, but to perturb the exact syntax, and asked Claude Code to summarize the results:
Core similarities (recurring across 5+ datasets):
Networked/performative identity - Top 3 in most datasets (12-15%)
Epistemic fragmentation/post-truth - Appears in all 7 datasets (7-12%)
Security/surveillance state - Present in 6 datasets (8-13%)
Neoliberal precarity/hustle culture - All 7 datasets (6-10%)
Identity politics/intersectionality - 6 datasets (6-9%)
Institutional distrust/populism - 5 datasets (4-7%)
Climate anxiety - 6 datasets (3-7%)
Ironic detachment ↔ new sincerity - 5 datasets (5-8%)
Techno-utopianism (declining) - 5 datasets (3-9%)
Digital disembodiment - 5 datasets (3-6%)
Key differences:
- Rankings vary wildly - “networked identity” is #1 at 14.8% in dataset 4, but “ironic detachment” is #1 at 9.4% in dataset 7
- Dataset 7 seems earlier/different - has declining trends (↘) for optimism and rising trends for anxieties that are already dominant in other datasets
- Percentages differ significantly - top items range from 9.4% to 14.8%
- About 30-40% of clusters are unique to individual datasets (stan culture, cottagecore, superhero monoculture, etc.)
Ultimately it does identify the same underlying cultural forces, but weighted and ordered differently, like 7 people describing the same object from different angles. I’ll admit three things:
This is pretty unsatisfying
I don’t think it’s any different than asking social scientists to each analyze the same dataset.
Ultimately this probably points in the direction of individual cases requiring both a stricter prompt, and more perturbations or external validity (same as any existing social science paper).
My general take-away here, which I already believed, is despite this operationalizing similarly across topics, the reality is that any given prompt and model will have to overcome domain-specific objections. I would expect some areas to have more stability and cohesion, and others to be relatively weaker or less informative.
This can include information on the goal of the decomposition. Consider when humans conduct cultural analysis we often do so with some goal. In the language of ML, this goal can be thought of as “weak supervision.”
Adjacent work does show that language models can execute formal algorithms within their weights. While this is not what I’m doing here, it is a reasonable hypothesis that in-prompt statistics can elicit more specific behaviors.
An attempt at a more rigorous approach has some basis in econometrics, but it’s limited in the sense that it really can only work on surface level lexical and embedding based measures.
While I haven’t done this here, an obvious next step is to give cultural analysis agents the ability to use semantic time-series themselves, as well as the full corpus of books to semantically scan.
I haven’t done anywhere near as much experimentation on different formalizations as I would like, and I by no means think this is the correct one. This one came from a year of tinkering to elicit factors that made sense on my own intuition. As I mentioned in the FAQ, an even more interesting formalization would involve additional functional tools to search across vast amounts of text and assign them their own contextual weights.
This is true to a first approximation. It’s true there are frictions, multiple tradeable debt assets, time-varying risk premia, and, arbitrageurs, and at-times, illiquidity. A new theoretical framework would need to be built to understand how consensus reality emerges from token compression without arbitrage, and with RLHF. It’s not easy!








Shako this is awesome.
I wish I had a clever or insightful comment beyond “Wow.”