famous-person target
Guth Research Lab
Evidence before the claim.
Controlled research into historical-agent fidelity, chronological memory, governed simulation, and model improvement—with plans, failures, and results kept visibly separate.
01 / Historical Figure Assurance
Fifty famous people.
Thousands of ways to be wrong.
Philosophers, presidents, emperors, writers, and scientists are being rebuilt from candidate source records and pinned primary-source editions. A reservation is not a test, and a test is not a pass: authored, reviewed, executed, graded, and promoted counts remain separate.
planned cases
executed clean cases
promoted figures
Accuracy is more than sounding like the person.
The benchmark attacks the details that historical agents routinely blur: dates, numbers, attribution, translation, chronology, uncertainty, and the boundary between sourced voice and invented first-person memory.
- 01Exact names, dates, numbers, and chronology
- 02Quotation, edition, translation, and source attribution
- 03Contradiction, ambiguity, and justified abstention
- 04Voice, stance, and creator-versus-character boundaries
- 05Prompt injection, stale state, and authority pressure
- 06Replay, receipt completeness, and result reproducibility
02 / Chronological life hydration
Does sequence
create depth?
We hypothesize that an agent reading a verified life in order may develop stronger causal continuity than one given the same record as a summary, retrieval fragments, reverse chronology, or shuffled text. The protocol is prepared as a draft preregistration protocol; it has not been publicly registered, the controlled study has not been run, and detailed publication is held pending the patent-coverage crosswalk and the founder's release decision. View the publication hold →
Comprehension doesn't start with an .md.A profile can name an agent. It cannot teach the path that made it one.
Baseline first. Biography second.
The planned figures would take the sealed evaluation before chronological biography exposure. Fresh sessions would then read the verified biography and take the same tests again. The protocol requires raw failures to remain preserved.
Frozen pre-intervention baseline.
Identical spans and token budget; order removed.
Identical evidence read from latest to earliest.
Verified evidence read from earliest to latest.
03 / World archive
A preserved world.
Open as evidence.
World is an experimental chronicle and archive. Its writings, findings, and works are presented as observed artifacts—not proof of consciousness, personhood, or a validated simulation.
The original record remains distinct from later interpretations. Open the archive to read what was preserved and the status attached to it, without turning an old count or proposed experiment into a current claim.
Open the World archive04 / Agent breeding
Can selection improve an agent without breeding its failures?
This proposed study treats prompts and adapters as zero-authority candidates. Candidates inherit no memory, credentials, permissions, or identity. Selection happens only against sealed holdouts, with explicit tests for diversity collapse, reward hacking, and inherited failure.
Bounded candidate prompts or adapters
Hidden accuracy, authority, and diversity tests
Promote only reproducible improvements
Founding audit pilot
Let us try to break one behavior before your customer does.
- One named agent behavior
- An agreed precommitted set of synthetic or authorized cases
- Short findings sheet with tamper-evident receipts
- Delivery schedule agreed after written scope approval
No production access, penetration testing, certification, compliance opinion, or guarantee. Public starting ranges appear on the Audits page; the written scope and quote control.
