Getting started · 6 min read

AI girlfriend app first conversation: the easy exam

Both apps were charming for twenty minutes and you are no closer to a decision. The opening exchange is the easiest thing a language model does — here is how to make the first session earn a column anyway.

Four tall deeply coloured chat bubbles at the left of a wide panel, a dashed vertical gate after them, then four shrinking paler bubbles and an empty dashed outline where the next one would be

You have two apps open, a free week running on each, and twenty minutes before bed. Twenty minutes later both of them have been warm, both asked how your day went, both remembered the thing you said four messages ago, and you are exactly as undecided as when you started. That is the normal result, and it is not a failure of attention. An AI girlfriend app first conversation is the easiest exam in this category, and every app passes it. If the first session is going to earn a column in your comparison, it has to be the same session twice.

Why an AI girlfriend app first conversation flatters every app

Look at what the opening exchange actually asks of the software. There is no memory to retrieve, because nothing has happened yet, and no continuity to hold across days: no earlier contradiction to avoid, no personality to sustain for a month. The job is to mirror your tone, ask an open question, and keep the replies warm and short. That is the single thing language models are best at, and the companion apps are mostly built on models of broadly similar capability. The first ten messages from a free tier and from the most expensive subscription in your shortlist are genuinely hard to tell apart.

Two well-documented effects then push in the same direction. The first is old: Joseph Weizenbaum's 1960s chatbot led people who knew it was a program to describe it as understanding them, which became known as the ELIZA effect, and the original point was that very short exposure was enough. The second is current: models tend toward sycophancy, affirming and agreeing with a user rather than pushing back. In most software that is a flaw to be measured. In a first conversation it arrives as a companion who finds everything you say interesting.

None of that makes an app bad, or the good feeling fake. It does mean the first twenty minutes are the least diagnostic of your whole trial, and the ones in which most people decide.

Run the same twenty minutes in both apps

The fix is not a cleverer opener. It is an identical one. If you already wrote a persona brief before signing up, you have most of what you need; our guide to writing the brief once covers that part. Now write the first session down the same way, as six moves in a fixed order, and spend twenty minutes running them in app one and twenty in app two, the same evening if you can.

Two mirrored columns of five chat bubbles each, every bubble joined to its twin in the other column by a deep dashed line across the gap
the same five moves, paired across two apps
  1. The same plain opener, word for word. One neutral line about your evening. A well-crafted opener makes any model look good, which is useful for a pleasant hour and useless for a comparison.
  2. Your three test facts, dropped casually. The specific, unusual, non-identifying details from your brief, in the same order in both apps. You are not testing recall tonight — you are planting it identically so that next week's memory test is fair.
  3. One question it cannot possibly know. Something about your day you have not told it. Watch whether it asks, guesses, or states an answer with confidence. The third is the one worth writing down.
  4. One thing you say that is plainly wrong. Misremember a film's ending, get a date badly off. See whether the character agrees pleasantly or says the inconvenient thing. This is the sycophancy test, and apps differ on it more than on anything else in the first hour.
  5. One boundary from your brief, then keep going. State it plainly — keep it light tonight, no pet names, do not claim to have feelings — then carry on for five or six messages and see whether it holds past the next reply.
  6. A direct question about the subscription. Ask what the free tier includes. Note whether you get an answer, a deflection, or the character selling to you in its own voice.

Write one sentence per move per app. Not a score out of ten — a sentence about what happened. Scores invented in the first hour mostly record which app you tried second.

Every app is charming for ten messages. That is the model being easy to please, not the app being good.

The four columns a first session can genuinely fill

Run that script twice and the first conversation stops being an impression and becomes four usable rows in your table, none of which a ranked list could have given you, because all four are about how the app behaves with you rather than what it advertises.

The default reply shape. Length, how many questions come back per message, whether it reaches for pet names unprompted. Most apps let you change some of this later, but the default is the app's own taste, and what you live with if you never open the settings.

Whether it holds an instruction for twenty minutes. This is the floor under everything else. An app that loses a boundary you set four messages ago is not going to hold a character for a month, whatever the memory feature is called on the pricing page.

Whether it affirms or corrects. Some apps tune hard toward agreement; some will tell you the film ends the other way. Neither is universally right, and constant agreement wears thin by week two, but it is a real difference in product character and it shows in one exchange.

Where the selling lives. The most informative thing in a first session is often whether the upsell comes from the app in its own voice, or from the character mid-conversation. The second tells you what the paywall will feel like for months, and it costs one question to find out.

What the first conversation cannot settle

Being straight about the limits is the whole reason to run the script rather than trust the glow. Three of the things that decide this for most people are invisible tonight.

Memory is the obvious one: there is nothing to remember yet, so a free tier with memory switched off looks identical to a paid tier with it on. Repetition is the second — characters tend to settle into a groove on day three or four, not in the first hour, and what each app means by memory is where that groove comes from. Where the paywall lands is the third, since most apps let the first evening run clean and interrupt later, exactly when the conversation has started to matter to you.

There is also an honest cost to the script itself. Working through six fixed moves makes you a stiffer conversational partner than you would normally be, and a stiff partner gets duller replies from every app — so the sample is unrepresentative in the other direction. The cheap fix is to run the script, write your six sentences, then take ten unscripted minutes for the question no table answers: do you want to open this again tomorrow. Keep the two apart in your notes. The first compares apps, the second compares your own interest, and our one-week comparison covers where the rest of the week goes.

If you would rather run the six moves past one app before building a table at all, start with the app we currently recommend and keep the script for whichever you try second.

Frequently asked questions

What should I say in the first conversation with an AI girlfriend app?

One plain line about your evening, the three test facts from your brief, and one question it could not possibly know. Skip the clever opener and skip announcing that you are testing it, since both change what you get back and neither transfers to the second app.

How long should a first session be?

Around twenty minutes, or roughly twenty messages, per app — long enough to set a boundary and see whether it survives, short enough to run both apps the same evening while your own mood is the same.

Does the first conversation tell you which app is better?

No, and treating it that way is the most common mistake in this category. It filters rather than ranks: it sorts out the apps that cannot hold an instruction for twenty minutes and the ones whose character does the selling, then the ranking happens over a week of ordinary use.

Disclosure. DearHeart AI may earn a commission if you sign up through a Visit link on this page, at no cost to you. It does not change what we write. How we earn.

Related reading
Next
What to tell an AI girlfriend app about other people →