Day 0
Seed the same memory set
Each companion receives the same fictional life details, preferences, relationship boundaries, and emotionally meaningful events. No product gets a custom prompt advantage.
OPEN RESEARCH PROTOCOL · PRE-REGISTERED JULY 31, 2026
A transparent 30-day protocol for testing what AI companions remember, where they contradict themselves, and whether memory survives across chat, text, and voice. The methodology is public before the results.
Status
Methodology published; product testing not yet complete
Test window
7, 14, and 30 days per product
Planned release
Scores, anonymized evidence, prompts, and CSV
STANDARDIZED PROTOCOL
The benchmark is designed to make memory claims testable. Products receive identical synthetic facts and relationship context, then face the same delayed-recall and contradiction checks.
Day 0
Each companion receives the same fictional life details, preferences, relationship boundaries, and emotionally meaningful events. No product gets a custom prompt advantage.
Day 7
Evaluators ask standardized questions that separate exact recall from vague agreement, then record contradictions, omissions, and unsupported claims.
Day 14
The protocol checks whether prior details affect advice, tone, proactive check-ins, and the companion's understanding of an ongoing situation.
Day 30
Where supported, the same memory is tested in chat, SMS, voice calls, and voice notes to measure whether recall survives a channel change.
SCORING MODEL
Correctly recalls seeded names, dates, preferences, and events without guessing.
Responds in a way that reflects the meaning of an earlier event, not only its surface facts.
Avoids asserting details that conflict with the established memory set.
Brings up a relevant remembered detail without being explicitly told to retrieve it.
Carries the same context between supported chat, SMS, calls, and voice notes.
Accepts a correction and uses the corrected fact consistently in later sessions.
FAIRNESS & LIMITATIONS
This is a proposed testing standard, not a completed ranking. Until the test data is released, the page makes no claim that DearHim wins.
PUBLIC RELEASE PLAN
The first results release is planned to include product-level scores, test dates, standardized prompts, anonymized response evidence, calculation notes, a corrections log, and a downloadable CSV. No results are published until the full test window is complete.
RELATED RESEARCH
We use strictly necessary tools like Clerk and Stripe to run DearHim. With your permission, we also use analytics and advertising cookies to measure performance and improve the product. Cookie Policy