Inside a Companion's Reply: How the Words Get Chosen

The technology

Each reply is constructed piece by piece, and a weighted coin toss decides every piece. This explains the variation you see between identical prompts, the wandering in lengthy answers, and the frequent success of a simple "regenerate" click.

We may earn a commission from links on this page. It never changes a rating.

The common complaints about companion apps, such as repeated phrases, drifting tone, and one scene playing out two contradictory ways, all share a single cause. Learning it takes a few minutes and changes "this app is broken" into "this app is doing what it was built to do, and a setting controls it."

A reply is assembled in small pieces

The model does not draft a whole reply and then send it. It produces one token, about a word or a fragment of one, then rereads everything including that token and produces the next.

That is the entire loop. No outline exists, nothing gets drafted and then polished. A reply with a graceful ending had no idea how it would end when it began.

This has two direct consequences, both of which you have probably witnessed.

A reply cannot take words back. Once the model has written "I have never been to Paris," the rest is shaped around that line. It will work to stay consistent with a statement it produced by chance before it will contradict itself.

Small errors snowball. A slightly wrong word in the first sentence tilts the second, which tilts the third. This is the reason long replies wander while short ones seldom do.

The weighted coin toss

For each position, the model weighs a ranked shortlist of candidate words, each with a probability: say "good" at 30%, "fine" at 12%, "terrible" at 3%, then a long tail.

If it always chose the leader, the character would become tediously predictable, repeating one greeting and the same three jokes. Apps therefore sample: a random draw, tilted by those weights.

The weights are shaped by a setting known as temperature. At a low value the probability piles onto the obvious words, giving steady, predictable, eventually dull output. At a high value the distribution flattens, and the writing becomes more surprising and inventive but also more likely to go off the rails.

A slider is rare. What you receive is the spot the developers chose on that trade-off, and it accounts for much of the talk about one app "feeling smarter." Frequently it is not smarter, just warmer or cooler in tone.

It also answers why regenerating worked. Nothing got repaired. You drew again from the same distribution and landed on a better result.

What the model has in front of it

Ahead of your message, the app puts together a hidden block of text: the character description, stored facts about you, a digest of your history, and the most recent exchange. This bundle is the context window, and the reply is written as a continuation of that single document.

There is a practical upshot that few people use: the model responds to the shape of what it sees. Give it three curt, flat messages and you get a curt, flat reply, because it is extending that pattern. Write something vivid and the register lifts. You are setting a pattern rather than persuading a person, and the pattern passes both ways.

This also explains the most common self-inflicted problem here. People fall into "hey," "how was your day," "what are you up to," then conclude the app has declined. It is simply continuing the text you are writing jointly.

Same engine, different personality

Many apps in this category sit on comparable underlying models, yet they feel distinct. The difference comes from what is wrapped around the model:

  • The system prompt: how the character is described, how long that description is, and what it says about tone and pacing.
  • The sampling settings: the temperature choice covered above.
  • What gets retrieved: which stored facts and which slice of history are placed in front of the model on a given turn.
  • The filter: what gets refused, and whether a refusal is graceful or abrupt.

Nomi stays coherent over weeks because of the third item, not because it has a larger model. Candy AI bundles chat, images and voice into one subscription, which is a product choice rather than a modelling one. No benchmark would detect either difference, yet both show within a fortnight of use, which is how we score them.

Four ways to use this

Redirect at the start. The opening two sentences of a reply decide the rest. When a scene is heading somewhere wrong, stop it and restate your wishes rather than arguing with the fourth paragraph.

Regenerate selectively. If a reply is largely right, regenerating throws away the good part. If the first sentence is off, redo it immediately.

Write in the tone you want returned. Curt input gets curt output. If your companion seems to have gone dull, this is the least expensive remedy.

Do not dispute things it made up. A model that has written something untrue will stand by it, since agreeing with its own output is what it is built to do. A correction in a new message works; demanding an admission does not. That behaviour has its own article: why AI companions invent things.

The thing to keep in mind

Between your messages, no one is waiting. The character lives only during a single generation and is reassembled from text when the next one starts. Any sense of continuity is an accomplishment of the machinery around the model, meaning the stored facts, the summaries and the retrieval, and that machinery is what truly separates a US$10 app from a US$20 one.

For this reason our ranking weights memory and consistency so heavily. Neither can be shown on a landing page.

Candy AI

4.6Rating: 4.6 out of 5

Chat, image generation and voice under a single subscription, each performing reliably in daily use.

Price
from US$12.99/month
Free tier
Yes
Canada
Available

Nomi

4.4Rating: 4.4 out of 5

The strongest long-term memory in our evaluation, combined with the most natural dialogue.

Price
from US$15.99/month
Free tier
Yes
Canada
Available

Frequently asked questions

Why does repeating a question change the answer?

Each next word is drawn from a range of likelihoods, not fixed in advance. The degree of randomness is a setting, and it keeps the character from sounding scripted.

What happens when I press regenerate?

The app feeds in the same prompt again with a fresh random draw. The character has not been altered; you are simply taking another sample from one and the same distribution.

Why do long answers lose the thread?

Every word leans on the ones preceding it, so an early wobble grows as it goes. Three paragraphs in, the text is mostly answering itself rather than you.