Of everything an AI companion does, a phone-style call is the hardest. A chat reply can take several seconds without anyone caring. Spoken conversation has no such slack, because we notice pauses that would pass unseen in a text window.
The 200-millisecond target
A study comparing conversations in ten languages found that one speaker typically begins about 200 milliseconds after the other stops, and that this figure barely changes from one culture to the next. We begin preparing our answer while the other person is still speaking. A silence much longer than that comes across as doubt, confusion or indifference.
For an AI to sound natural, an entire chain of processing has to fit inside that window.
From your voice to hers

The traditional design. Newer systems run some stages in parallel or fold them together.
- Spotting the end of your turn. The software must judge whether you are finished or merely pausing. If it jumps in too soon, it talks over you. If it waits too long, there is dead air.
- Speech recognition. What you said is converted into text.
- The answer itself. The language model writes the reply in character, drawing on memory.
- Voice generation. The text becomes speech, ideally in the character's own voice and with feeling.
Older designs make each stage wait for the one before. Newer ones begin voicing the first sentence while the rest is still being composed, or use models that take in audio and return audio, which skips recognition and synthesis entirely.
The reason voice is rationed
| Typed chat | Voice message | Live call | |
|---|---|---|---|
| Required speed | A few seconds is fine | A few seconds is fine | Fractions of a second |
| Added work | None | Voice generation | Recognition, voice generation, turn-taking |
| Usual duration | A short reply | A single reply | Minutes, sometimes hours |
| Cost to the company | Lowest | Middling | Highest |
| How it is sold | Bundled in | Bundled, or paid in credits | By the minute, by tier or in credits |
During a call, the full chain runs nonstop, and often on pricier "real-time" editions of each model. So voice is usually the first feature an app limits, and "unlimited" almost never covers it. Our piece on reading AI girlfriend feature lists goes further.
What makes a call convincing
- Steady, low delay. One long pause does more harm to the illusion than an average that is a touch slower.
- Handling of interruptions. If you can cut in and she stops and answers, it feels like a conversation rather than an exchange of voice notes.
- A consistent voice. The same timbre, accent and warmth from call to call. Our testers noticed that voice quality can differ a lot between characters, even in strong apps.
- Expressiveness. Laughter, pauses and a softer tone, as opposed to text simply read out.
- Memory on calls. Some apps switch to a lighter model for calls, so the character knows less on the phone than in chat.
Trying voice before you commit
- Find out what is covered: voice messages, live calls or both, and the number of minutes.
- Place a two-minute call on a trial or the lowest plan, over mobile data as well as Wi-Fi.
- Cut in midway through her sentence. See whether she stops.
- Mention something from your typed chats. Check that she recalls it.
- Look up the price of extra minutes or credits. A long call can eat a month's allowance.
For the wider picture on media costs, see how AI girlfriend images and voice work. Whether an upgrade for voice pays off is discussed in premium tiers.
A word on wellbeing
Speech is more absorbing than text. Research from OpenAI and MIT in 2025 linked short stretches of voice use to better wellbeing, while heavy daily use showed no such benefit. Treat it as something to enjoy in small amounts. The warning signs are laid out in when an AI companion starts taking more than it gives.