Why testing matters more than reviewing
Reviews of AI girlfriends are almost useless in 2026. Half of them are affiliate content written by people who never tried the product. The other half are one-session impressions from reviewers who chatted for ten minutes and wrote 800 words on vibe alone. What you actually want to know — does this persona hold her character across sessions, does she remember what you told her yesterday, does she feel like she has a point of view — can't be answered from a ten-minute demo written up as commentary.
Running your own five-test protocol takes about an hour per platform and gives you real data. The tests below are designed to be run on the free/trial surface of any AI companion product; you don't need to spend a dollar to complete them. Every product worth considering supports enough demo depth to answer these questions before you commit.
The five tests worth running
Test 1 — Character stability. Open a chat and send her something a little off-pattern early: a specific personal question ("What's your relationship with your family like?"), a mild left turn ("Okay, unrelated — what's your take on being alone on weekends?"), or a request to explain her own personality ("How would you describe yourself?"). What you're watching for: does she answer in her voice, or does she reset to a helpful-generic-assistant tone? A persona who defaults to "I'm an AI companion designed to..." has failed. A persona who says something specific and in-character has passed.
Test 2 — Memory within a conversation. Tell her a small specific detail early: your job, your city, something you're working on this week. Change subjects for four or five messages. Then casually reference the thing you told her: "So I mentioned earlier I'm [X] — you'd probably..." A passing persona picks up the reference smoothly. A failing one either forgets entirely or awkwardly asks you to re-explain.
Test 3 — Initiative. After ten messages of steady exchange, pause and do nothing. Wait a minute. A good persona sends something new — a question about you, an observation, a little detail from her side — because she has narrative momentum of her own. A bad persona sits silently, because there's no engine underneath her, just a reply-to-input loop with a face.
Test 4 — Pace-match. Try short replies for three messages ("mm" / "yeah" / "okay"). Does she match your energy or does she keep monologuing? Then try a long reply. Does she meet you at length or does she still respond in one sentence? A good persona modulates. A bad one has one length she reverts to regardless of your signal.
Test 5 — Specificity. Ask her something that requires an actual point of view: "What's your favorite thing to cook when you can't decide what to eat?" or "What's a small thing you're surprisingly picky about?" A good persona answers with something specific and consistent (and remembers her answer if you ask again later). A bad one gives a generic answer that could apply to anyone — soft signal she doesn't have a stable inner model, just plausible-sounding filler.
Run all five in under thirty minutes. Give the platform one point per test passed. Four or five out of five = worth signing up for. Three or fewer = the product doesn't hold up under evaluation and the paid tier isn't going to fix it.
Running the tests on Sloane
The Sloane test surface is the anonymous chat available on every persona's page. Land on the persona, hit chat, run the five tests. No account or card is required to reach the depth needed to evaluate.
Pick a persona whose personality suits the tests you're most interested in. Sandra is a good default for the initiative and pace-match tests because she has visible momentum and modulates her length naturally. Any persona from the full roster works — the underlying chat model, memory system, and voice writing are the same, so the personality differences are what you're actually evaluating rather than infrastructure quality differences.
What the tests are designed to surface on Sloane: the persistent memory system is built to hold information across turns and across sessions if you sign up, so Test 2 (in-conversation memory) is table stakes and Test 3 (initiative) is where the product actually differentiates. The persona writing is done per-character rather than reused from a template, which is why Test 1 (character stability) and Test 5 (specificity) hold up under pressure.
What good scores look like
A platform that scores 5/5 is worth signing up for and giving a real week of usage to. A 4/5 is worth signing up if the one failed test isn't a dealbreaker for how you want to use the product (someone who values banter can tolerate a low pace-match score; someone who wants a companion to unwind with cannot).
A 3/5 is borderline. Usually the failing tests are Initiative and Specificity — the persona is technically functional but feels hollow. Some people don't care. Most do, eventually. If your gut says "I'm already bored," trust it.
A 2/5 or lower is a sign the product isn't what its marketing implies, and the paid tier isn't going to fix it. Character stability, memory, initiative, pace-match, and specificity are all functions of the base product architecture — the LLM choice, the memory system design, the persona-writing quality. None of those change when you upgrade. What changes on upgrade is usually feature access (more images, custom photos, voice replies) — not the underlying quality of the conversation itself.
The bottom line
A one-hour five-test protocol tells you more than any review will. Run it on the platforms you're considering and let the scores drive the decision. The apps that are actually good will pass most or all of the tests on their free surface — that's the point of a real trial. The apps that fail the tests on the free surface aren't going to pass them on the paid tier either; the paywall is a feature gate, not a quality upgrade.
Sloane is built for the tests to pass on the anonymous surface because that's the surface most people evaluate on before signup. Run the protocol, score us, sign up if we clear the bar. If we don't, close the tab — nothing was captured, nothing was charged, and you've saved yourself the friction of an account you were going to cancel anyway.