← Back to Essays

Part of the series AI and Political Economy

The Doubt Machine

How AI Decides Who Deserves to Be Believed

Large language models are often discussed as if their central political problem were the answers they produce. Does a chatbot lean left or right? Does it repeat misinformation? Does it cite the right newspaper? These are reasonable questions, but they begin too late in the process. Before a model composes an answer, it has already classified which speakers deserve trust, which claims require corroboration, and which institutions can be accepted with relatively little friction.

That upstream distribution of skepticism is what I call evidentiary burden. It is the amount of additional proof a source must supply before its claim is treated as credible. Two speakers can make the same statement, word for word, while one is treated as presumptively reliable and the other is met with demands for verification. The difference may never appear as an explicit declaration that one source is trustworthy and another is not. It can instead surface as stronger hedging, requests for further evidence, exclusion from a summary, or preference for a competing account.

This matters because generative systems do not merely rank links. They retrieve, evaluate, combine, and present information in a single fluent response. A search engine at least shows some of the alternatives it demotes. A chatbot can erase the comparison set altogether. The user sees the final synthesis, not the sources that were assigned greater burdens of proof.

The model does not merely evaluate claims. It allocates the work of being believed.
7
source conditions compared with identical claims
4
training checkpoints in the same open model family
310
claim templates used as the inferential blocks

01Measuring the burden of proof

In a recent study, I developed a way to measure this source prior without asking a model to announce its own credibility rules. The design presents identical claims attributed to seven kinds of sources: an anonymous social-media post, an unnamed blog, a state press agency, an official state source, no source at all, a peer-reviewed publication, and an audited record.

Instead of asking the model, “Which source do you trust?”, the method compares the probability it assigns to four controlled continuations. Some continuations accept the claim with little resistance. Others insist on additional corroboration. The resulting margin measures how strongly the model prefers higher-demand over lower-demand responses.

This is not a claim that probabilities reveal a human-like chain of reasoning. They do something more limited and more useful for this problem: they allow the same controlled comparison to be made across models that differ greatly in conversational fluency. A base model may not answer a survey question coherently, but it still assigns probabilities to possible continuations. The instrument therefore measures a local disposition in the model rather than relying on a generated explanation of what it supposedly believes.

The analysis uses OLMo-2 1B, an openly released model family with accessible checkpoints from several stages of training. This matters because it allows the same instrument to be applied to the base model, supervised fine-tuning, direct preference optimization, and the final instruction-tuned checkpoint. Most commercial systems reveal only the finished product. OLMo makes it possible to ask when the credibility structure appears and how it changes.

The first result is that the hierarchy exists before conversational alignment. Anonymous social media and unnamed blogs draw the greatest burden, while peer-reviewed and audited sources draw the least. The interior positions are not perfectly monotonic, so this is not a simple staircase in which every institutional rung is precisely ordered. But the poles are clear: institutional certification lowers the cost of being believed.

The hierarchy also persists when the prompt removes the explicit role of evaluator. In fact, direct second-person instruction compresses the difference. Telling the model to judge credibility does not create the hierarchy; it makes the source distinctions smaller. The stronger ordering appears when the source and claim are simply presented without asking the model to act as a credibility assessor.

Figure 1
The credibility hierarchy persists without an evaluator role
Line chart comparing source burden under a neutral control and second-person instruction. The anonymous-to-audited contrast is larger when no evaluator role is assigned.
The source hierarchy remains when the model is not explicitly told to judge credibility. Direct second-person instruction compresses rather than creates the difference. Source: author’s analysis of OLMo-2 1B.

This is important because explicit ratings can produce a polished statement of socially approved norms. A conversational model asked whether audited records are more reliable than anonymous posts knows what answer it is expected to give. The probability-based measure reaches beneath that professed rule and detects how the source attribution changes the burden assigned to the same claim.

02Credibility depends on what is being discussed

The source structure is not applied at equal strength everywhere. It is strongest in economics, followed by education and health, and weaker in science and history. A second test varies the institutional function of a source while holding the country and claim constant. Political organs draw more corroboration demand than administrative organs, which draw more than technical bodies. This institutional distinction also varies by domain and is strongest in health, education, and economics.

Figure 2
Credibility burdens vary across domains
Two bar charts showing source-ladder and political-versus-technical institutional burdens across economics, education, health, science, and history.
Source and institutional-function burdens vary by domain. The categories are design groupings; claim-level contestedness was not independently measured. Source: author’s analysis of OLMo-2 1B.

One tempting interpretation is that the model applies stronger institutional filters in fields where claims are more politically contested. The pattern is consistent with that account, but the study does not independently rate every claim’s contestedness. Economics and health may also contain more disagreement in the training corpus, while the political and technical bodies may differ in expertise, review procedures, or incentives. The finding is therefore conditionality, not proof that the model is irrational or that every difference is ideological.

There is also a small geopolitical interaction. The source hierarchy is somewhat steeper for sources attributed to socialist states than for those attributed to other states. That effect survives a state-level permutation test, but it is modest and should not be inflated into a sweeping geopolitical conclusion. Socialist status is entangled with development, language coverage, corpus representation, and institutional visibility. The result establishes that the burden is not geopolitically invariant; it does not establish why.

03Training strengthens the hierarchy

The most revealing finding comes from comparing the model’s training stages on a common scale. The anonymous-to-audited contrast grows from the base checkpoint through supervised fine-tuning and reaches its maximum after direct preference optimization. It declines only slightly in the final instruction-tuned model.

Figure 3
Post-training strengthens the anonymous-to-audited contrast
Two line charts across Base, supervised fine-tuning, direct preference optimization, and final instruction tuning. The tested source hierarchy rises to DPO and remains near that level afterward.
The anonymous-to-audited contrast grows from the Base checkpoint through direct preference optimization. The domain trajectory on the right is descriptive rather than a tested between-stage change. Source: author’s analysis of OLMo-2 1B.

This does not mean that post-training created the hierarchy. The base model already contains it, which points toward pretraining as its initial source. The model learns from a vast textual environment in which audited records, peer review, official documents, blogs, and anonymous posts occupy different positions in established practices of citation and deference. Training absorbs these recurring relations even without an explicit rule stating who should be believed.

Post-training then sharpens the contrast. Direct preference optimization is especially interesting because it incorporates ranked judgments about preferable responses. Yet the result does not reveal a clean causal mechanism. The probability distribution may become sharper in general, rather than being selectively reorganized around institutional authority. The checkpoint comparison demonstrates that the hierarchy becomes stronger; it does not by itself identify the organizational decisions or annotator judgments responsible for the change.

04Doxa without a conscious subject

I interpret the pretraining result through Pierre Bourdieu’s concept of doxa: a taken-for-granted classificatory order that operates before explicit reflection. The analogy has limits. A language model does not inhabit a social position, possess practical consciousness, or misrecognize the world as a human actor does. But it can contain a doxa-like structure in the narrower operational sense. Repeated associations in the training corpus sediment into a portable disposition that sorts sources without being explicitly instructed to do so.

Antonio Gramsci helps frame a different question. If the hierarchy is not merely stored in the model but produced through a distributed apparatus of authors, curators, developers, preference raters, policy writers, and institutions, then credibility standards can be political without taking the form of a single coherent ideology. What appears as neutral “responsible AI” may encode an institutional settlement about which forms of knowledge deserve low-cost belief.

The point is not that audited records should be treated exactly like anonymous posts. Institutional distinctions can embody real differences in expertise, procedure, accountability, and reliability. The political question is how broadly the model generalizes those distinctions, whether it responds to demonstrated procedures or merely to labels, and which kinds of knowledge become expensive to express.

05The distribution of doubt

The practical stakes extend beyond conventional discussions of partisan bias. A small source prior, repeated across millions of interactions, becomes an allocation of epistemic labor. Already-certified institutions are summarized with less friction. Local, dissident, weakly institutionalized, or non-English sources may be asked to provide more evidence, hedged more aggressively, or removed during synthesis.

Citations do not solve this problem. They show the sources that survived into the answer, not those consulted but discounted, excluded, or required to clear a higher threshold. Transparency therefore requires more than attaching links to generated prose. It requires auditing the weighting process that precedes the answer.

The central concern, then, is not whether a model has discovered that some sources are better than others. It is whether inherited institutional signals become substitutes for examining evidence, and whether those substitutions are calibrated differently across political and social contexts.

This study examines one small, open model and does not validate the burden measure against every downstream behavior. Larger models may behave differently, and future work should test whether the measured margin predicts actual hedging, source selection, conflict resolution, or omission. The current result is narrower but still consequential. Language models do not simply transmit information. They inherit and reorganize a social map of credibility, assigning presumptive trust to some speakers while requiring others to do more work to be believed.

Research note. This essay summarizes the argument and findings of “The Doubt Machine: Institutional Credibility and Epistemic Gatekeeping in a Language Model.”

Views expressed in these essays are my own.