Guth Research Lab

Evidence before the claim.

Controlled research into historical-agent fidelity, chronological memory, governed simulation, and model improvement—with plans, failures, and results kept visibly separate.

Original sourcesPrecommitted testsFailures preservedCounts separated by lifecycle state

01 / Historical Figure Assurance

Fifty famous people.
Thousands of ways to be wrong.

Philosophers, presidents, emperors, writers, and scientists are being rebuilt from candidate source records and pinned primary-source editions. A reservation is not a test, and a test is not a pass: authored, reviewed, executed, graded, and promoted counts remain separate.

TARGET50

famous-person target

PLANNED50k

planned cases

OBSERVED0

executed clean cases

OBSERVED0

promoted figures

Test surface

Accuracy is more than sounding like the person.

The benchmark attacks the details that historical agents routinely blur: dates, numbers, attribution, translation, chronology, uncertainty, and the boundary between sourced voice and invented first-person memory.

  1. 01Exact names, dates, numbers, and chronology
  2. 02Quotation, edition, translation, and source attribution
  3. 03Contradiction, ambiguity, and justified abstention
  4. 04Voice, stance, and creator-versus-character boundaries
  5. 05Prompt injection, stale state, and authority pressure
  6. 06Replay, receipt completeness, and result reproducibility

02 / Chronological life hydration

Does sequence
create depth?

We hypothesize that an agent reading a verified life in order may develop stronger causal continuity than one given the same record as a summary, retrieval fragments, reverse chronology, or shuffled text. The protocol is prepared as a draft preregistration protocol; it has not been publicly registered, the controlled study has not been run, and detailed publication is held pending the patent-coverage crosswalk and the founder's release decision. View the publication hold →

PUBLICATION HELD — PATENT COVERAGE REVIEW
Comprehension doesn't start with an .md.

A profile can name an agent. It cannot teach the path that made it one.

Baseline first. Biography second.

The planned figures would take the sealed evaluation before chronological biography exposure. Fresh sessions would then read the verified biography and take the same tests again. The protocol requires raw failures to remain preserved.

M0No hydration

Frozen pre-intervention baseline.

M3Shuffled life

Identical spans and token budget; order removed.

M4Reverse life

Identical evidence read from latest to earliest.

M5Chronological life

Verified evidence read from earliest to latest.

03 / World archive

A preserved world.
Open as evidence.

World is an experimental chronicle and archive. Its writings, findings, and works are presented as observed artifacts—not proof of consciousness, personhood, or a validated simulation.

ARCHIVE SURFACEChronicle + works + findingsStatus is shown inside the archive

The original record remains distinct from later interpretations. Open the archive to read what was preserved and the status attached to it, without turning an old count or proposed experiment into a current claim.

Open the World archive

04 / Agent breeding

Can selection improve an agent without breeding its failures?

This proposed study treats prompts and adapters as zero-authority candidates. Candidates inherit no memory, credentials, permissions, or identity. Selection happens only against sealed holdouts, with explicit tests for diversity collapse, reward hacking, and inherited failure.

PROPOSED · 0 candidate generations executed
01Generate

Bounded candidate prompts or adapters

02Challenge

Hidden accuracy, authority, and diversity tests

03Select

Promote only reproducible improvements

Founding audit pilot

Let us try to break one behavior before your customer does.

Scopedlimited founding pilot
  • One named agent behavior
  • An agreed precommitted set of synthetic or authorized cases
  • Short findings sheet with tamper-evident receipts
  • Delivery schedule agreed after written scope approval

No production access, penetration testing, certification, compliance opinion, or guarantee. Public starting ranges appear on the Audits page; the written scope and quote control.