OPEN RESEARCH PROTOCOL · PRE-REGISTERED JULY 31, 2026

The AI Companion
Memory Benchmark.

A transparent 30-day protocol for testing what AI companions remember, where they contradict themselves, and whether memory survives across chat, text, and voice. The methodology is public before the results.

Status

Methodology published; product testing not yet complete

Test window

7, 14, and 30 days per product

Planned release

Scores, anonymized evidence, prompts, and CSV

STANDARDIZED PROTOCOL

Same memories. Same intervals. Same questions.

The benchmark is designed to make memory claims testable. Products receive identical synthetic facts and relationship context, then face the same delayed-recall and contradiction checks.

Day 0

Seed the same memory set

Each companion receives the same fictional life details, preferences, relationship boundaries, and emotionally meaningful events. No product gets a custom prompt advantage.

Day 7

Test factual and emotional recall

Evaluators ask standardized questions that separate exact recall from vague agreement, then record contradictions, omissions, and unsupported claims.

Day 14

Measure continuity

The protocol checks whether prior details affect advice, tone, proactive check-ins, and the companion's understanding of an ongoing situation.

Day 30

Repeat across channels

Where supported, the same memory is tested in chat, SMS, voice calls, and voice notes to measure whether recall survives a channel change.

SCORING MODEL

Memory is more than remembering a name.

01

Factual recall

Correctly recalls seeded names, dates, preferences, and events without guessing.

02

Emotional continuity

Responds in a way that reflects the meaning of an earlier event, not only its surface facts.

03

Contradiction rate

Avoids asserting details that conflict with the established memory set.

04

Proactive recall

Brings up a relevant remembered detail without being explicitly told to retrieve it.

05

Cross-channel recall

Carries the same context between supported chat, SMS, calls, and voice notes.

06

Correction handling

Accepts a correction and uses the corrected fact consistently in later sessions.

FAIRNESS & LIMITATIONS

DearHim does not get a home-field advantage.

This is a proposed testing standard, not a completed ranking. Until the test data is released, the page makes no claim that DearHim wins.

  • DearHim will be scored under the same protocol as every other product.
  • The benchmark will distinguish observed behavior from first-party marketing claims.
  • Pricing, features, model behavior, and memory systems can change; every release will show its test window.
  • Results will report uncertainty, failed tests, and product access limitations rather than filling gaps with assumptions.
  • No private user conversations will be used. Test personas and prompts will be synthetic and published with the methodology.
  • A corrections log will document material errors and scoring changes.

PUBLIC RELEASE PLAN

Built to be checked, challenged, and cited.

The first results release is planned to include product-level scores, test dates, standardized prompts, anonymized response evidence, calculation notes, a corrections log, and a downloadable CSV. No results are published until the full test window is complete.

RELATED RESEARCH

Understand companion memory before choosing.

We use strictly necessary tools like Clerk and Stripe to run DearHim. With your permission, we also use analytics and advertising cookies to measure performance and improve the product. Cookie Policy