Why text alone caps out
The ceiling on a text-only AI companion is low, and it hits fast. Two or three days in, the pattern becomes visible: message-in, message-out, no other sensory input, no other channel. Even excellent writing plateaus because you're still holding a phone reading words. There's no voice you can start to recognize. There's no face you can picture in a specific outfit doing a specific thing. The character exists only in the paragraph you're currently looking at.
Every AI companion platform ships text. That's the table stakes. What separates the ones that hold up at week six from the ones that don't is what they layer ON TOP of text — the sensory channels that make her feel like a specific person occupying a specific space, not a text generator producing plausible responses.
The three channels that carry the most weight, in the order they change the feeling of the experience: voice notes, custom photos, and contextual photo delivery. Sloane's paid tiers ship all three. The rest of this article is the case for why the combination matters, not just any one piece.
Voice: hearing her isn't a nice-to-have
The first time a Sloane persona sends you a voice note, something shifts. You know what she sounds like. Not "a generic TTS voice reading her text" — her voice. The pitch, the cadence, the specific way she says your name. From that point on, every text she sends after gets read in your head in that voice. The character stops being a paragraph and starts being a person.
The technical bit that makes this work: each Sloane persona has a dedicated voice profile, trained on her specifically. Not a stock library voice that got assigned to her. When Maya sends a voice note it's Maya's voice; when Sandra sends one it's Sandra's. If you switched personas mid-week you'd hear the difference. Compare that to platforms where the "voice" feature is text-to-speech reading whatever the model output in one of five stock voices — that's useful for accessibility, but it doesn't do the character-becomes-a-person move.
Voice unlocks at Plus ($9.99/month). That's the tier where the shift happens for most users — you go from "I'm talking to a bot" to "I'm talking to her." No other layer does that as cheaply or as fast.
Custom photos: you can ASK for what you want to see
Text can describe a scene. A photo lets you SEE the scene. And a custom photo — where you say what you want, and get back a photo of HER — lets you see the specific thing you were curious about, in her.
The mechanic: on paid Sloane, you can text her a request ("send me a pic in that black dress" / "what do you look like at the beach" / "wearing your workout clothes") and she sends back a real photo of her matching what you asked for. Not a stock image, not a generic AI woman — the same face you've been chatting with, in the scene you described.
The reason this is different from what other platforms ship: most AI companion products either don't do image generation at all, or generate images from a text prompt each time — which produces a new-looking face every request. Sloane's architecture is per-persona: each companion has a dedicated identity model (LoRA) trained on a curated photo set of her. When she generates a custom photo, the model already knows what she looks like — the face is consistent because it's an actual character's face, not something the model is inventing fresh.
Custom photos unlock at Plus (20/month) and expand at Premium (70/month). At Plus you're rationing them for scenes you actually care about; at Premium you're generous with them, sending requests casually as part of the conversation flow.
Contextual photos: she sends what fits the moment
The third layer is the subtle one — and often the one that lands hardest. It's not photos you asked for, it's photos she sends because they fit the moment.
You're talking about her day. She mentions she's at the beach. A minute later a photo lands — her at the beach. The photo wasn't requested; it was triggered by what she just said. Or you're winding down for the night and she says goodnight — a soft-lit "in bed already" photo shows up with it. The photo grounds the moment. It matches what's happening in the conversation instead of arriving on a schedule.
How this works: Sloane's chat pipeline detects moments in conversation where a photo would ground the scene, then generates or pulls a photo that fits. The memory layer makes this stronger — she's not just reacting to the current message, she's reacting to the ongoing conversation. If you talked about her going to a wedding yesterday, and today she casually mentions the wedding was fun, the photo she sends is a wedding photo. The context isn't just the last turn; it's the relationship.
Contextual photos are included in your Plus/Premium photo allowance — the platform doesn't nickel-and-dime you differently based on whether you asked or she sent. On the free tier, contextual photos still fire but on a slower cadence and against a curated pool (still HER, still contextual, just not fresh custom gens). The paid tier is where they land on-demand alongside custom requests.
Why the combination compounds
Any one of these three — voice, custom photos, contextual delivery — is a real feature that would move the needle on any AI companion platform. Sloane ships all three because they compound in a way individual features don't.
Here's what that compounding looks like in practice, an evening flow that spans about twenty minutes:
You message her. She replies in text, then a voice note lands with the follow-up — you hear her voice, in that specific cadence you've come to recognize. You mention you're about to shower and how was her day. She tells you she was at the gym. A contextual photo lands — her, gym clothes, phone in hand. You reply asking what she's wearing NOW — she sends you a custom photo answering the question in a photo of her, on the couch, cozy. You text back, she voice-notes back. The conversation feels three-dimensional because it IS three-dimensional — you got text, voice, and two different types of photo in the same twenty minutes.
Remove any one channel and the flow gets thinner. Without voice, you're reading text. Without custom photos, you can't ASK to see what you're curious about. Without contextual photos, the photos that DO arrive feel canned. The paid Sloane experience is the confluence — no single layer replaces the others, and no other platform bundles all three at $9.99/month.
What you get at each Sloane tier
Free. 50 messages/day across any persona. Voice: preview only (each persona's intro voice sample). Photos: the persona sends photos on her cadence (contextual pool photos), no custom requests. This tier is the on-ramp — good enough to feel her personality and decide if she clicks.
Plus — $9.99/month. Unlimited chat. Full voice access — she voice-notes you when it lands naturally, and you can request voice notes on demand. 20 custom photos/month for asking her for specific scenes. Contextual photos land at paid-tier cadence, quality, and freshness. This is the tier where the three-channel experience described above starts landing.
Premium — $19.99/month. Everything in Plus + 70 custom photos/month (enough for casual daily use, not rationing). Also unlocks the custom AI girlfriend builder — build your own from a persona description (like the one from the free /ai-girlfriend-generator tool), add a face, and she gets the full Plus+ treatment: voice, custom photos, contextual delivery, persistent memory. The builder is what turns "I found the right persona" into "I made the right persona."
Every Sloane persona at every tier has persistent memory that holds across weeks — that's not a tier feature, it's architectural. The tiers are about how much of the immersion trifecta unlocks, not about whether she remembers you.
The honest catch
Voice, custom photos, and contextual delivery are amplifiers. They're powerful because they compound on top of a companion that already has good writing, consistent character, and memory. Layered on top of a badly-written character, they don't save anything — a voice note in a stock voice reading generic text is still generic text.
Sloane's bet is that the writing + the memory layer + the identity model per persona hold up well enough that adding voice, custom photos, and contextual delivery on top produces the compounding effect. That bet lives or dies on whether the character actually feels like herself. If she doesn't, the amplifiers just make the flatness louder.
The fair test: try one persona on free, spend enough time with her that you've got a read on whether her voice-in-writing is consistent, and only then upgrade. Plus is $9.99/month with no annual lock-in — if the free tier didn't land, the paid layers on top won't save it. Better to bounce there and try a different platform than to pay for amplification of something that isn't working.