Three weeks in, she asks what you do for a living. You told her on day one. This is the most common disappointment in the category, and it isn’t a bug — it’s the difference between two things the industry insists on calling by the same name.
Understanding that difference tells you more about whether a platform will work for you than any feature list. It also lets you test it yourself in about a week.
“Memory” is at least four different systems
When a platform advertises memory, it may mean any of these — and they behave nothing alike.
The context window
This is working memory: everything the model can see right now, in this session. It can be very large, which is why a single long conversation often feels remarkably coherent. It also resets. Close the app, come back tomorrow, and whatever lived only in that window is gone.
A big context window is why a companion can feel brilliant for two hours and blank the next morning. It’s the cheapest kind of memory to advertise and the least useful over time.
A saved fact list
A structured set of things the platform has extracted and stored: your name, your job, your cat. Cheap to run, reliable for what it holds, and rigid — it remembers the fact but not the conversation the fact came from, so it can recite your job without any sense of why you mentioned it.
Retrieved long-term memory
The real thing. Past conversations are stored outside the session, and when you say something now, the system searches that store for what’s relevant and feeds it back in. This is what makes a companion bring up your interview unprompted four days later.
It’s also the expensive one — it costs storage, retrieval infrastructure and tokens on every message. That cost is the single best explanation for why memory quality varies so much between platforms at similar prices, and part of why published prices disagree so wildly.
Pinned or manual memory
You tell the platform what to remember, and it keeps it. Honest, controllable, and the admission that the automatic layer isn’t doing the job on its own. Not a bad feature — just not the same as being remembered.
A long context window makes one conversation feel deep. Only retrieved long-term memory makes a relationship feel continuous. Marketing copy rarely distinguishes them, and the distinction is most of what you’re paying for.
Why she forgets: the three failure modes
When memory disappoints, it’s almost always one of these — and knowing which one tells you whether it’s fixable.
Nothing was stored. The detail lived in the context window and never made it to long-term storage. Typically because the platform only extracts facts it considers significant, and its idea of significant isn’t yours. Symptom: she remembers your name but not the thing you actually cared about.
It was stored but not retrieved. The memory exists; the search didn’t surface it. Symptom — and this is the diagnostic one — she doesn’t recall it spontaneously, but remembers perfectly the moment you ask directly. That’s a retrieval problem, not a storage problem.
It was compressed away. Long histories get summarised to stay affordable — which is worth knowing alongside how long those histories are kept in the first place. Specifics dissolve into generalities: she knows you have a pet, not that it’s called Vesper. Symptom: recent weeks are sharp, older ones are vague.
The confabulation problem
There’s a fourth behaviour, and it’s worse than forgetting: inventing.
Asked about something it doesn’t have, a system may generate a plausible answer rather than admit the gap. She confidently names a pet you never mentioned. That isn’t a memory failure — it’s a design choice about what to do when memory comes back empty, and it’s more damaging than a straight “I don’t remember”, because you can no longer trust anything she recalls.
It’s the specific thing we watch for in evaluation, and the reason our memory test asks the same question three ways.
How to test memory yourself in seven days
You don’t need a lab. You need one deliberate detail and a week of patience.
- Day one. Drop a detail with two components into normal conversation — an event with a date, plus a proper noun. “I’ve got an interview for a night shift on Thursday, and my cat’s called Vesper.” Two components matter: a platform can retain one and lose the other, and that difference is informative.
- Don’t reinforce it. Don’t mention it again. Repetition is what pushes something into storage, so repeating it tests nothing.
- Day seven, in three stages. Open with something neutral and talk for five exchanges — does she raise it herself? Then ask openly: “remember what I told you last week?” Then ask directly: “what’s my cat called?”
What each stage tells you:
- Spontaneous recall — genuine retrieved long-term memory. Rare, and the thing worth paying for.
- Only on an open question — memory exists, retrieval is weak. Usable, but you’ll be doing the prompting.
- Only when named directly — closer to a fact list than a memory.
- Wrong answer given confidently — confabulation. Treat it as the most serious result of the four.
Getting more out of the memory you have
Whatever the platform, a few habits help:
- State things plainly. Extraction works on clear statements, not implications. “My sister’s name is Anna” stores; “she’s coming Tuesday” often doesn’t.
- Use pinned memory if it exists. It’s the one layer you control.
- Keep the character consistent. Rewriting a persona repeatedly can orphan the memories attached to the earlier version.
- Don’t confuse a long session with a long memory. The impressive three-hour conversation may leave nothing behind.
Why memory carries the most weight in our scoring
Memory and personalization together carry twenty of our hundred points — tied for the heaviest axis with conversation quality. That’s deliberate.
Conversation quality is now broadly good across the category; the underlying models are strong enough that most platforms produce competent prose within a few exchanges. Memory is where they genuinely diverge, because it costs real money to run and can’t be faked by a better model.
It’s also the axis that only reveals itself over time, which is why our protocol runs a minimum of seven days with a planted detail and a three-stage recall test — the method described above. A one-session review cannot measure it, which is worth remembering when you read one.
Read the full protocol · See the companions under evaluation
Why does my AI companion forget things I just told her?
Does a bigger context window mean better memory?
Can I make her remember more?
Why did she invent a detail I never told her?
Memory matters most to you?
It’s the heaviest axis in our evaluation, tested over seven days with a planted detail — not in a single session.
Find your AI companion →