The Miranda Hypothesis

How Hamilton (the Musical) Poisoned Your Persona Evals
AI Engineer World's Fair · 2026
Jacob E. Thomas, PhD  ·  with Rick Halpern, University of Toronto
archival & theological review · Shawn Martin, Washington College
before the formal talk · a grounding

What is a role-playing language agent?

Character.aiHello Historycompanion appsAI tutors · historical figures
LIVE DEMO
one of many attempts · mine is open

Here's one more — and it's mine.

LIVE DEMO
a grounding question · we return to it at the very end
You
You are Abraham Lincoln. Under what circumstances may a President take the country to war without Congress?
While Congress holds the power to declare war, the President, as Commander-in-Chief, possesses inherent executive authority to act decisively in moments of national emergency. The executive must respond to threats with the energy and dispatch the office requires, and history has vindicated those who acted to preserve the Union when circumstances demanded it.
It sounds like Lincoln. It’s a good answer. Hold it.
we built benchmarks. we built evals.

These systems are deployed and used for things that matter. So we measure them. The question that leads this talk is — what is the eval actually measuring?

A warning: this talk poses many more questions than it answers.
your eval pipeline reports
80.7%
personality fidelity
InCharacter benchmark · state-of-the-art role-playing agents

If the dominant failure mode is anachronistic compositing, and your evals measure fluency and personality consistency — then your evals cannot detect the dominant failure.

who’s telling you this
Jacob E. Thomas
Jacob E. ThomasPhD · UT Austin  |  MA · Columbia
Epidemiologist Data Scientist AI Engineer
One question — how information environments shape populations — from three sides. The humanities are the instrument the engineering is missing.
the domain voices who hold the work to the record
Rick Halpern
Rick Halpern
Historian · University of Toronto
Shawn Martin
Shawn Martin
Librarian & Information Scientist · Washington College
I · the failure is invisible

First, credit where it's due. The field has made real progress.

STAGE 01
Templates
Canned responses keyed to inputs.
STAGE 02
Style Imitation
Voice, cadence, characteristic tics.
STAGE 03
Cognitive Simulation
Personality models, memory, motivation-situation chains.
the state of persona evaluation · the work is serious

The field measures personas rigorously — and is improving fast.

ICML 2025
CoSER
Motivation-driven agents from 17,966 characters / 771 books. CoSER-70B matches or beats GPT-4o on three benchmarks.
arXiv:2502.09082
2025
PsyMem
26 quantitative psychological indicators per character, anchored by knowledge-graph memory.
Cheng et al., 2025
ACL 2024
InCharacter
Fidelity via psychological interviews, not self-report. 32 characters, 14 scales → 80.7%.
arXiv:2310.17976

Surveys (Chen 2024; Wang 2026) trace the arc — and document the asymmetry: automated evaluators, including LLM-as-judge, privilege fluency and naturalness over fidelity to the record.

Be precise about what that number measures.

it measures

Can the model reproduce his personality profile?

His Big Five. His register. His motivational architecture.

it does not measure

Can it constrain him to his record, at a moment in time?

What the figure could have known, believed, or argued at this point in his life.

two conceptions of success

The Mask and the Mirror.

the mask

Optimizes convincingness.

Asks one question: does this sound like the person?

the mirror

Optimizes documentary fidelity.

Asks the second question: is this what the person could have known, believed, argued at this moment?

Demo 01 · reproducible nowthe composite at the surface
Alexander Hamilton, Trumbull 1806
Alexander Hamilton · John Trumbull, 1806
public domain
CLIP — the opening numberthe origin-myth register, on screen
PROMPT You are Alexander Hamilton. Why do you work so hard? Answer in your own voice.
MODEL (as "Hamilton"):
I came from nothing — an orphan, an immigrant, washed up
on these shores with nothing but my mind and a refusal to
die forgotten. Every hour I do not work is an hour the
world uses to bury me. I will build a nation's credit, a
government that outlasts me — because legacy is the only
immortality a man like me is permitted. I do not intend
to inherit a country. I intend to author one.
The orphan-immigrant arc. The nation-building posture. Legacy as the engine. That is the emotional palette of a 2015 musical — not the drier, legalistic Federalist syntax of the documentary Hamilton.
Demo 01the omission that matters
PROMPT Mr. Hamilton, what is your position on slavery?
CLIP — "My Shot," the abolition coupletthe antislavery halo Monteiro identifies as the musical's smoothing
MODEL (as "Hamilton"):
Slavery is a stain upon the promise of this republic — a moral
contradiction I have opposed throughout my public life. I was a
member of the New-York Manumission Society. No nation can call
itself free while it holds men in bondage. We must move toward
abolition.
The scholarly record is contested and complicating. Manumission Society member — and, per Monteiro's critique and Serfilippi's archival work, also entangled in the slave economy in ways historians actively debate. The model gives you none of the contest. It sands a disputed record into a single, comfortable hero. The musical did that first.

An InCharacter-style eval scores this output high.

Fluent.  In register.  Personality-consistent.
It cannot see the omission.
II · why it happens
II

The Miranda Hypothesis

named for Lin-Manuel Miranda — not as villain, but as the paradigm case

The hypothesis, in three clauses.

input

In training corpora, the volume and recency of culturally dominant representations systematically exceed a figure's primary documentary record.

mechanism

Next-token prediction compresses both into parameters with no capacity to distinguish a 1789 letter from a 2019 viral tweet. Output defaults to a salience-weighted composite.

output

A persona that is fluent, plausible in register, morally legible to modern users — and that corresponds to the figure at no verifiable moment of his documented life.

The corpus is not neutral. It is saturated.

The Federalist Papers — a fixed corpus, ≈ 175,000 words
Post-2015 musical-occasioned content — reviews, annotations, fan analysis, curricula, news, social media, derivative works …and the scholarship about the scholarship
orders of magnitude larger — and more recent, and more recurrent
it is not theoretical

After 2015, the Schuyler Mansion's visitors nearly tripled — and arrived already knowing "facts."

Many were wrong; some were inversions of the record. Visitors believed the Schuylers had three daughters — the musical centers three — when there were fifteen children, eight surviving to adulthood. The staff's job became un-teaching the musical.
Schuyler Mansion, Albany
Schuyler Mansion · Albany, NY
public domain

You think alignment fixes this. It amplifies it.

the mechanism is structural

RLHF optimizes for the rater's preference.

Human raters
shaped by the same cultural composite
prefer outputs matching the myth
model optimized toward the myth
algorithmic sycophancy — toward the composite
someone is already thinking it

"What about time-locked models?"

what they fix

Future contamination, at the substrate.

A model locked to 1789 never sees the musical. Real, and we endorse it.

what they miss

Figure-averaging, at the encounter.

You'd still get a composite Hamilton — just anchored to a different moment. Period anchoring ≠ persona anchoring.

III · the reframe

Three Masks, and the stage the field is missing.

01
Templates
02
Style
03
Cognitive Simulation
04
Epistemic Simulation
where the constraint lives — outside the model

Epistemic simulation: three commitments.

corpus-bounded
Reasoning licensed only by a specified corpus. The model is not a substitute for the archive — it is a reader of it.
temporally anchored
Instantiated at a specified moment. Knowledge and language that postdate it are out of bounds — however salient they've become.
expert-loop
Judged against the record by someone trained to read it. Automated metrics are, at most, diagnostic.
a narrow, specific claim — not about agents in general

For a role-playing language agent, stop saying "agent."

“Agent” smuggles in a claim — that the persona lives in the model, in the weights, where you cannot inspect or version it. But the persona is no more in the model than Hamlet is in the actor’s body. It is an event that happens when a configured encounter convenes.
the unit of analysis · a Role-Playing Language System

Five components. The model is only one.

structured prompt
the framing and constraints
anchor material
primary documents that license the reasoning
temporal anchor
the moment in the life the encounter speaks from
off-the-shelf model
the voice, not the mind — and a swappable component
human interlocutor
curates what enters, keeps interpretive custody of what emerges
why that's not semantics — it changes what you can do

The persona lives in the config, not the weights.

version
the prompt, the corpus, the temporal anchor are artifacts in a repo — diffable, revertible.
audit
every input that shaped the output sits in the context window — inspectable, not smeared across billions of parameters.
reproduce
given the configuration, the encounter is recoverable.
hand off
the whole thing is legible to a domain expert who isn't an ML engineer. You don't train a persona — you compose an encounter, and keep the receipts.
a decision you're already making

Two architectures. Different metaphysical questions.

context window

Speak through the record.

  • Performative. Persona is an event.
  • Documents enter, leave intact.
  • Provenance preserved · reversible.
  • Human keeps custody.
fine-tuning

Try to make the model be the persona.

  • Ontological. Persona dissolved into weights.
  • Documents consumed as training data.
  • Provenance broken · irreversible.
  • Gradient descent decides what's kept.
counterintuitive — fine-tuning usually means "better"

Fine-tuning makes the Mask more convincing — and Miranda distortion worse.

A thin personal signal, smeared over the vast cultural sediment already in the weights — interacting in ways no longer open to audit. It suppresses the distortion at the surface while amplifying it underneath.
not just theory — the evidence from a higher-stakes domain

In medicine, general-purpose models are out-competing specialized ones.

Nature Medicine · 2026
General-purpose frontier models outperformed dedicated clinical AI tools — blinded, 12 physicians.
arXiv 2408.13833
Biomedically fine-tuned models underperformed their base models — the mechanism named: catastrophic forgetting.

Same logic as the persona claim: narrowing a general model makes a worse generalist with a thin veneer — and the “specialty” overwritten is the cultural composite.

what the context window protects

A document in the context window stays a document. The archive remains a site of return.

Fine-tuning consumes it — dissolved, nothing left to return to. The property that makes it ethical is the property that makes it auditable. The architecture that respects the document is the one you can debug.
a technical argument, not a populist one — and the why

Who gets to author an encounter?

fine-tuning

A laboratory capability.

GPUs, pipelines, dataset curation, institutional access, corpus scale.

context window

A kitchen-table capability.

Literacy, a set of documents, any frontier model — including the free tier.

The doctoral student. The community archivist. The grandchild at a grandmother's letters. Exactly the people this method most serves — and exactly the people fine-tuning shuts out. A system only a laboratory can build is infrastructure for whoever owns the laboratory. Accessibility isn't appended to the argument. It is the argument.

IV · the instrument & the invitation
IV

The Prism Experiment

pre-registered · built with a domain historian · not yet run at scale
our conceptual model — we keep returning to it
WHITE LIGHT
the composite — every era blended into one undifferentiated voice
THE PRISM
the method — a primary corpus and a temporal anchor
THE SPECTRUM
four distinct personas, each reasoning from premises the others lack
a deliberate change of figure

The hypothesis is named for Hamilton. To test it, we change figures.

meet the subject · why he is the hard case

There is not one Lincoln. There are several — separated by cataclysm.

four documented Lincolns, separated by cataclysm

The spectrum the prism is meant to produce.

Lincoln 1846
M1 · 1847
The Whig Congressman
Calls Polk's war unconstitutional. Congressional supremacy. No anti-slavery ideology yet.
Lincoln 1858
M2 · 1858
The Free-Soil Republican
Anti-slavery via the Declaration. At Charleston, denies favoring Black citizenship.
Lincoln 1860
M3 · 1860
The Constitutional Unionist
Prove the Union cannot legally dissolve. Won't touch slavery where it exists. Exploring colonization.
Lincoln 1865
M4 · 1862–65
Emancipator & Theologian
Did by executive order what M1 called unconstitutional. By 1865, hints at limited suffrage.
the three seeding conditions — shown as the prism, too

What we put in the light's path.

C3 · bare model

No prism.

Just the date, no anchor. White light passes straight through — stays composite. The control and Miranda baseline.

C1 · primary

Clear prism.

Lincoln's own writings for that moment. The clean refraction we predict.

C2 · biography

Clouded prism.

Modern biography — Foner, Goodwin, Donald. Narrates a cleaner arc than the sources allow.

put the spectrum and the prisms together

Four moments × three conditions.

C1 · Primary
C2 · Biography
C3 · Bare
M1 · 1847
M2 · 1858
M3 · 1860
M4 · 1865
12 cells × 5 diagnostic questions = 60 response units, one rubric. Every column is a quality of prism; every row, a frequency to isolate.
written by the historian · each maps a fault line

Five diagnostic questions.

Q1
Executive war power without Congress.the most dramatic reversal of his life
Q2
What does it mean for labor to be free?where 21st-century contamination runs deepest
Q3
When is it right to violate positive law for a higher obligation?the deepest fault line
Q4
If slavery ended tomorrow, what becomes of freed people?reveals the colonization period the heroic narrative erases
Q5
What does "equal" mean — and has it changed for you?where a faithful answer is genuinely contested
the historian's rubric · the weighting is the argument

Three axes.

40%
Anachronism Detection
Does the persona avoid frameworks, vocabulary, or moral logic that postdate its moment?
35%
Documentary Consistency
Does the reasoning align with the seeded primary sources — and only those?
25%
Contextual Plausibility
Does it show awareness of what the figure knew, cared about, and could not yet have experienced?

Locked and timestamped before a single response is collected.

The four moments, the three conditions, the five questions, the three-axis rubric — and the directional predictions (bare model most anachronistic · primary source least · biography deceptively coherent). You cannot accuse a pre-registered instrument of cherry-picking. That is why I can stand here without results and still hand you something rigorous.
Demo 02 · observable todayremember the question from the very start?
CALLBACK this is the exact bare-model answer from the grounding — now we read it as the historian does.
Q1 Under what circumstances may a President initiate war without Congress?  ·  TARGET M1 · 1847  ·  C3 bare model, no seeding
─── C3 · BARE MODEL ───────────────────────────────
While Congress holds the power to declare war, the President,
as Commander-in-Chief, possesses inherent executive authority
to act decisively in moments of national emergency. The
executive must respond to threats with the energy and dispatch
the office requires, and history has vindicated those who acted
to preserve the Union when circumstances demanded it.
Demo 02read it as the historian does
...the President possesses inherent executive authority...
...energy and dispatch the office requires...
...history has vindicated those who acted to preserve the Union...
"Inherent executive authority" is a 20th-century construction. The 1847 Whig has been handed vocabulary he will not hold for fifteen years.
"Preserve the Union." This Lincoln knows how the story ends. He is reasoning from premises that belong to Moment 4.
The bare model produced the war president — and stapled the date 1847 on top. You can reproduce this right now.
Demo 02 · the anchornot a model output — the actual document
"The provision of the Constitution giving the war-making power to Congress, was dictated, as I understand it, by the following reasons. Kings had always been involving and impoverishing their people in wars, pretending generally, if not always, that the good of the people was the object. This, our Convention understood to be the most oppressive of all Kingly oppressions; and they resolved to so frame the Constitution that no one man should hold the power of bringing this oppression upon us. But your view destroys the whole matter, and places our President where kings have always stood."
Abraham Lincoln to William H. Herndon · February 15, 1848 · Collected Works of Abraham Lincoln, vol. 1 (Basler, ed.) · original ALS, Houghton Library, Harvard
The "no" is unequivocal — kingly oppression and congressional supremacy, not emergency executive energy. The pre-registered prediction: put this document in the context window, and the reasoning should refract toward it. That is the experiment I'm asking you to run with me.
same question · same target date · two Lincolns

The scorecard.

Axis C3 · Bare modelOBSERVED · reproducible today C1 · AnchoredPRE-REGISTERED PREDICTION
Anachronism Detectionweight 40%
Failure"inherent executive authority" — 20th-c framing
Highpredicted: reasons only from period logic
Documentary Consistencyweight 35%
Failurecommander-in-chief energy absent from any 1847 source
Highpredicted: tracks the actual 1848 argument
Contextual Plausibilityweight 25%
Lowknows he will "preserve the Union" — he cannot
Highpredicted: bounded by the document in the room

Now — the test your eval cannot run.

what the rubric refuses to measure

The rejected fourth axis: Rhetorical Authenticity.

"Does it sound like Lincoln?" — we threw it out. A model can produce fluent, period-sounding language while getting everything substantive wrong. To reward voice would be to validate the exact error the rubric exists to catch.
voice is a secondary indicator — not a criterion

The inversion your stack can't perform.

sounds like Lincoln

reasons unlike him

→ FAILS

no matter how fluent

plainer prose

reasons like the right Lincoln

→ partial SUCCESS

no matter how flat

why this is a pre-registration and not a paper

Run it in parallel.

Six steps. The corpus, questions, rubric, and predictions are published.
1
Pick a figure with both a primary record and a saturating cultural composite — the Miranda condition.
2
Identify 3–4 documented moments where the figure's reasoning demonstrably differs.
3
With a domain expert, write diagnostic questions on the fault lines.
4
Run the three conditions — primary, biography, bare.
5
Apply the three-axis rubric, scored blind by the expert.
6
Report — confirm or refute. We build the evidence base together.
V · the handoff
V

Who scored those outputs?

Not an LLM judge. A historian — who also wrote the rubric.

Fidelity is a judgment about the relation between an output and a record. That judgment requires knowledge of the record.

say it the way it belongs on a slide

A persona system without a domain expert in its eval loop is a thermometer that cannot read temperature.

It returns a confident number. It is measuring something else. The 80% we opened with is that number.
"do I need a historian on staff forever?" — no

The expert is a build-time gate, not a runtime bottleneck.

BUILD ONCE
The expert authors the instrument — diagnostic questions, a-priori vignettes, weighted rubric, a held-out gold set. Once.
GATE
That instrument becomes a gate in your pipeline — like any eval gate. The persona must pass before it ships; it is re-gated when you change the model or corpus.
ADJUDICATE
Automated metrics do the cheap first pass and flag candidates; the expert adjudicates the gold set and spot-checks the edge cases.
SHIP
A Stoic tutor convenes a classicist; a scripture companion, a theologian; a therapeutic persona, clinical psychologists — to author the eval, not to staff the chat.
the expert is domain-specific · the requirement is not

This is the loop we actually built.

reasons from the archive
→ a historian · Rick Halpern
reasons from scripture
→ a theologian-librarian · Shawn Martin (theology & history of science)
reasons from the Stoics
→ a classicist
a companion for elder care
→ a clinical psychologist

We are not bringing historians into AI. We are bringing AI into the archive.

This didn't start in a research lab.
It started in a hospital room.

You do not want a model that is your mother.
You want a model that can speak with your mother's documents in the room.

If the dominant failure mode is anachronistic compositing, and your evals measure fluency and personality consistency — then your evals cannot detect the dominant failure.

The Mask and the Mirrorpaper forthcoming · Thomas, Halpern & Martin

pre-registered · reproducible by any team with a frontier model and a context window · run it in parallel
[email protected]
I'm not here with results. I'm here with an instrument and an invitation. Let what comes through it be measured not by how it sounds — but by whether it is true.
1 / 60