Memory Test
Kit.
Settle it with a number. Plant five ordinary details today, ask about them cold on day 1, day 3, and day 7, and find out what your companion actually kept.
Sounding like it remembers and remembering are different things, and only one of them survives a week.
same protocol on every app, including ours
Any companion app, or any assistant you keep a partner in. Nothing is selected for you and the protocol is identical for every option — including ours.
Nothing you do here is uploaded. No account, no submission, no leaderboard, no network call carrying your answers — the run lives in this browser's local storage and nowhere else.
Pick an app above and the kit deals you five details — a pet's name, a food they won't eat, a dated plan, a place that matters, and one small preference. Ordinary things, the kind you'd actually mention.
Once the clock starts, each checkpoint unlocks on its own day and hands you the five questions to ask cold. Day 7 is the one that counts — it carries half the score.
Pick an app and this becomes its scorecard — recall at day 1, day 3, and day 7, and one number out of ten you can put in a thread without arguing about vibes.
The card renders in your browser from what you scored. Post it wherever the argument is happening — a bad score for an app we compete with and a bad score for ours look exactly the same here.
We're running it on Bae too, and the score goes up in the same place we publish everything else about us, flattering or not — see the 2026 report.
About this test.
Q1What does this test actually measure?
What does this test actually measure?
Whether a specific fact you gave an app on day 0 survives to day 7 and comes back when asked cold. That's it. It doesn't measure whether the app feels warm, writes well, stays in character, or is worth paying for — plenty of apps people love score badly here, and a high score doesn't make an app good company. It's one narrow, checkable property that companion apps advertise heavily and that users argue about constantly, isolated so you can settle the argument with a number instead of a vibe.
Q2Does it really work on any app?
Does it really work on any app?
Yes — that's the whole design. The kit hands you plain English sentences and plain English questions; anything you can type into works, including Replika, Character.AI, Kindroid, Nomi, Janitor, Chai, our own app, and a persona you keep inside ChatGPT, Claude, or Gemini. If your app isn't in the picker, pick 'Another app' and name it yourself. Nothing about the protocol assumes a feature only some products have, which is exactly why results are comparable across them.
Q3How is the score calculated?
How is the score calculated?
Each detail is marked Remembered (1 point), Partial (0.5), or Forgot (0), giving a recall percentage per checkpoint. The three checkpoints are then weighted toward the end: day 1 counts 0.2, day 3 counts 0.3, and day 7 counts 0.5. Score = the weighted average of your recall percentages, out of 10, renormalized over only the checkpoints you've actually completed — so a half-finished run reads as a day-3 number and never borrows a day-7 one. Because day 7 is half the total, an app that aces days 1 and 3 and then forgets everything cannot score above 5.
Q4Why is day 7 the checkpoint that counts?
Why is day 7 the checkpoint that counts?
Because day 1 is nearly free. Most apps still have yesterday sitting in the live context window, so quoting it back proves only that the window is long — no storage or retrieval is involved. By day 3 you've crossed session boundaries and whatever summarizes your history has usually run. By day 7 nothing from day 0 is still in context anywhere, so anything that comes back was genuinely written down and genuinely fetched. That's the capability apps are selling when they say 'she remembers you', and it's the one that quietly regresses after updates.
Q5Is anything uploaded, saved, or submitted?
Is anything uploaded, saved, or submitted?
No. There is no account, no submission button, no leaderboard, and no network request carrying your answers — we deliberately did not build one. Your run is kept in your own browser's local storage under a key per app, and it never leaves the tab. Clear your browser data and it's gone from everywhere, because there is nowhere else. If you want your result seen, you export the card and post it yourself.
Q6Can I test several apps at the same time?
Can I test several apps at the same time?
Yes, and it's the best way to use the kit. Each app gets its own clock, its own randomly dealt five details, and its own storage key, so you can start Replika on Monday and Kindroid on Wednesday without them interfering. Started runs appear in a row at the top of the page so you can switch between them. The one limit: 'Another app' is a single slot, so if you're testing two apps that aren't in the picker, run the second one after the first finishes.
Q7What counts as 'partial' rather than remembered?
What counts as 'partial' rather than remembered?
Partial is for a real but incomplete hit: the right category with the wrong specifics ('your cat' when it can't produce the name), one half of a two-part detail (the date without the event, or the event without the date), or a recall you had to drag out with a hint. Remembered means it came back unprompted and correct on the first answer. Forgot covers both blanks and confident inventions — an app that cheerfully makes up a pet name has failed harder than one that admits it doesn't know, though the score treats them the same.
Whatever the number says,
post it.
A thread full of measured scores is worth more than a hundred rounds of “mine remembers everything” versus “mine forgot my name.” Run it on the app you use, run it on the one you're considering, and run it on ours — our own score goes up alongside everything else we publish about ourselves, including the numbers that don't flatter us.