Is "Training" an AI Companion Possible? What Each Kind of Feedback Does

The technology

Some users score every single reply, convinced that they are teaching their companion. For the most part they are not, or not in the way they picture. This article explains the actual effect of each type of feedback.

We may earn a commission from links on this page. It never changes a rating.

"Training" suggests software that gradually learns your preferences like a puppy. Reality has more layers. Some controls alter behaviour immediately, others alter it slowly and for all users at once, and a few simply make you feel involved.

The ingredients of one reply

The model writes each reply from whatever is visible to it at that moment; how an AI companion writes its reply describes the process. It draws on four inputs:

  1. The description and settings of the character.
  2. Memories and summaries fetched for the current chat.
  3. The most recent messages.
  4. The base model, as last updated by the company.

Feedback has an effect only if it changes one of those four inputs.

Ranking the controls

Ranking of ways to shape an AI companion by speed and strength: editing the character, editing replies, memory notes, out-of-character notes, regenerating and ratings

Alter what the model reads, not what it feels about you.

MethodWhat it touchesHow soonHow much
Editing the character descriptionInstructions placed ahead of every replyRight awayStrong and durable
Editing a replyThe history the model copies fromUpcoming repliesStrong until it scrolls out of view
Memory notesFacts fed into later chatsWhen next retrievedGood for facts, poor for style
Out-of-character notesThe current sceneRight awayWears off as the chat grows
RegeneratingOnly the reply in front of youRight awayNone past that reply
Ratings (thumbs, stars)Usually training data for later model versionsWeeks to months, if at allSlight for your character

Rewrite the reply instead of debating it

Complaining in character ("you're not acting like yourself") rarely works, as the complaint itself becomes text for the model to imitate. Replacing the reply with what the character should have said gives it a right example to follow. If editing is available in your app, this habit returns more than any other.

Where style and facts belong

  • Style means tone, how long her sentences run, humour, and her level of warmth. Put it in the character description so that it governs each reply. Which fields matter is covered in our character builder guide.
  • Facts are things like your occupation, your dog's name or last week's events. Put them in memory, which brings them up when they apply.

Swapping the two causes trouble: style notes kept in memory resurface inconsistently, while long lists of facts in the description squeeze out the personality.

What ratings are for

Across most apps, ratings go into a company's dataset for model improvement, frequently combined across all users. That can be a fine thing to contribute to, as future versions benefit, but it is not a direct channel to your character. A few apps say ratings shape your companion in particular; even so, the effect builds slowly and is much weaker than a revised description.

Anyone concerned about privacy should check the policy, since a rating may mark that conversation for human review or training, which carries a cost. See who reads your AI companion chats.

A short weekly tune-up

  1. Make a note of two things you liked this week and one that bothered you.
  2. Restate the annoyance as a positive instruction in the description ("She answers in brief, dry sentences" rather than "Stop rambling").
  3. Put any significant new facts into memory.
  4. Begin a fresh chat when the current one has wandered; lengthy conversations build up quirks, as personality drift explains.

A companion never comes to know you as a person does. Yet it reads you anew every time, via its description, its memory and your latest words. Alter what it reads and you alter who it is.

Frequently asked questions

Do thumbs change my companion?

Rarely at once, and often not yours specifically. The typical app collects ratings as material for improving its model across the whole user base in future releases. A few apply them to your personal character, yet the change is small and gradual compared with editing the description.

What is the quickest way to change how she talks?

Edit the description or personality field, and correct replies by rewriting them rather than arguing. The model mimics its instructions and the latest messages, so both controls take effect immediately.

Does she learn from our conversations?

A memory system stores facts and summaries, which is how details get recalled. The core model is not usually retrained on your chats as you go. Whether they go toward future versions is a matter of company policy.