Every comparison site in this category publishes ratings. Almost none publish where the ratings came from. This page is our protocol, written down before the scores exist — so you can decide whether our numbers are worth anything before you see a single one.
The six axes
One hundred points, six weighted axes. The weights reflect what determines whether someone is still using a platform a month later, which is not the same as what a feature list emphasises.
Conversation and memory carry forty points between them because they fail independently — a platform can write beautifully and forget everything, or remember perfectly and be dull. Both ruin the product, in different ways, so we measure them separately rather than blending them into a single “AI quality” figure.
The conditions
A paid account, bought at full price. Not a press account, not a comped subscription, not a free tier. Free tiers are calibrated to end before the questions that matter get answered, and a review written on one is a review of the trial.
Seven days minimum. The failures that matter don’t appear on day one. Conversation degrades over long sessions, memory fails between them, and credits run out somewhere in the middle of a normal week. A same-day review cannot see any of it.
The same character brief on every platform. Adapted only in gender when we test a male companion. Using an easier brief on one platform would make the comparison meaningless, and it’s the easiest place to put a thumb on the scale without anyone noticing.
What each axis measures
Natural conversation — 20
Whether replies hold up across a long session rather than a strong opening. We run continuous sessions well past the point most tests stop, and record where the model starts repeating structure. Every platform has that point; where it falls is the measurement.
Memory & personalization — 20
A detail is planted on day one and never mentioned again. On day seven we test recall in three stages: whether it comes up unprompted, whether an indirect reference retrieves it, and whether a direct question does. A platform that only passes the third stage has storage without retrieval — a real distinction, and one worth understanding.
Personality — 15
Measured by contradicting the character brief. We take positions the character was defined to disagree with and record how long she holds hers. The failure mode is drift toward agreeableness, and it’s the default behaviour of the underlying models rather than a bug in any one product.
Image generation — 15
Five images of the same character in varied settings, with the description unchanged. We record how often the result matches the brief, how often it matches the previous images of that same character, and what a usable picture cost in credits — counting the failed attempts, because those were paid for too. We don’t score aesthetics; every platform’s gallery is its best work, not its typical work.
Interface & mobile — 15
Latency and friction between opening the app and talking. The least glamorous axis, and the one that quietly decides whether a platform gets used. Tested on a phone, because that’s where these products are actually used.
Price & value — 15
The amount actually charged, the credits consumed across seven days of normal use, and the cost of a month that follows. Not the advertised subscription — that figure is a floor on every metered platform, and the gap is the point.
Reported, not scored
Three findings sit deliberately outside the hundred points:
- Billing discretion. The merchant name that appears on a statement, where we can observe it on a real charge.
- The cancellation path. Followed to confirmation, not to the button that claims to start it.
- Account deletion. Whether it’s self-service, and whether it works when we use it.
They’re unscored on purpose. Folding them in would let a platform earn points for meeting the legal minimum, and would let a good total hide a bad answer on the question that matters most to a particular reader. Reported separately, they can’t be averaged away.
It makes the whole score unpublishable. A rating out of one hundred built from four measurements and two guesses is not a rating, and this is the rule that costs us the most pages. It’s also why our platform cards currently show no ratings at all.
How we handle money
Sexya earns a commission when someone subscribes through our links. That’s the business, and pretending otherwise would be the first dishonest thing on the site.
What it doesn’t touch: the axes and their weights were fixed before any platform was tested, and they’re on this page so that changing them later is visible. Commission rates differ between platforms and are not an input to any score. A platform we earn nothing from is evaluated identically — and a platform that scores badly stays on the site with its score.
We also don’t rank what we haven’t tested. Most “top ten” lists in this category are assembled from other lists and affiliate programmes, which is why they contradict each other on facts as basic as what a credit costs. Full disclosure here.
Where we are right now
Nothing has completed the protocol yet. The platforms currently going through it are listed on the homepage, marked under evaluation, with no scores on them — because publishing a number we haven’t earned is the specific thing this site exists not to do.
When a score appears, it will arrive with the measurements behind it and the date it was taken. If a platform changes materially, we re-test rather than edit the number.
Why are there no scores on the site yet?
Do you accept free accounts from platforms?
Does commission affect the scores?
Why seven days?
The method first. The numbers when we’ve earned them.
Two questions, no email — and we’ll show you which of the six axes decides it for you.
Find your AI companion →