A companion app looks like one clever entity that chats, sketches and talks aloud. Underneath, the developer has bolted several unrelated tools together, and the annoyances you meet in this area nearly all trace back to the bolts.
Separate tools, a single window
The language model runs the conversation. It outputs text and nothing more.
The picture generator, normally a diffusion model, is an unrelated tool. It turns a written description into an image, and your chat is invisible to it.
The voice engine turns written text into sound. It too works alone and is aware of nothing beyond the line it is handed.
Take a selfie request. First the chat model drafts a short caption of the shot you asked for. Next the app adds the character's saved look tags. That merged text is sent to the picture generator, which returns an image. Back in the chat, the companion reacts to a photo it cannot actually see, going only by its own caption.
That hand-off is behind nearly every media complaint in this category.
The reason faces drift
A diffusion model never retrieves your character from a file. It invents a person who matches the wording, and a phrase such as "long dark hair, green eyes, mid-twenties" suits millions of faces, so you meet a new one on every request.
Developers tackle the problem with varying effort:
- Saved look tags: one phrase pasted into every request. Cheap, and only roughly steady.
- A reference picture that steers every generation. A big improvement, and normally what you find when a face stays recognizable.
- A model tuned for one character: the steadiest result and by far the costliest, so it tends to show up on upper tiers only.
This is a real point of difference and worth checking on a free tier first. Ask for four images of one character in four situations. Four different women means no plan will repair it. Keeping the face the same from image to image is a large reason why Candy AI heads our ranking, and of why Secret Desires, which designs look, voice and personality together and includes short clips, earns its score.
Why media is the part with a meter
Words are nearly free: one chat reply costs the operator well under a cent.
Pictures cost noticeably more each time. Audio is paid for per second. Video is the priciest of the lot by a wide margin.
That spread produces the pricing shape you see across the category: plenty of messages, strict image limits, voice minutes priced separately, and video reserved for top plans. It is not made-up scarcity meant to drive upgrades; it is the operator's real cost structure handed to you.
This is also why "unlimited" here nearly always refers to text alone. Keep that in mind and the pricing pages read sensibly. The arithmetic appears in what image generation really costs.
What voice adds and what it costs
Adding a voice shifts the experience more than most expect. Reading a line and hearing it are not equivalent, and the gap is wider than the spec sheet suggests.
Check two things before paying for it:
Lag. A pause of three seconds before each answer spoils the illusion. Test it yourself during a trial, not by watching a promotional clip.
Note or call. A voice note, in which the character reads out its reply, is widespread and cheap. A live call, where you speak and it replies, is another product and normally another tier. Nomi includes voice calls in its feature list, whereas plenty of apps write "voice" and deliver notes.
Limits of generated images
Two limitations are worth knowing so you are not let down:
Hands, lettering and fine detail stay unreliable in every generator we have used in this category. Nobody has cracked it yet; an app saying otherwise is showing its finest result rather than its usual one.
Your scene does not carry over. Only a single line of description reaches the generator, so anything left out disappears: the room you set, what she wore a few messages earlier, whether it is night. Spelling these out in your request helps more than any toggle.
A one-evening test of the media side
Burn through the whole free allowance in one session instead of a picture a day. Look at four points:
- Sameness of face: four images, one character, varied situations.
- Faithfulness: request a precise detail and check what makes it through.
- Speed: both for pictures and for speech.
- Accounting: whether a failed or declined generation still eats into your quota.
That final item is the one least often spelled out and the most irritating to learn after you have paid. Like everything in what free tiers really include, you only find it by putting the product through its paces.

