Skip to content
Bae
Annual report · 2026-05-25 · updated 2026-09-08

The State of
AI Companions
2026.

What 4,550 companions show people actually do. How 12 platforms really perform after fourteen days. And the honest version of the question nobody in our industry wants to answer.

By Bae · ppl.studio · Clinically reviewed · ~10 min read · Free to cite

Why we made this

We build one of these products. We published the data anyway — including the two categories where a competitor beats us.

We couldn't find an honest, methodology-first account of where this category actually stands. So we wrote one, from the two datasets we have: a survey of the people who come to us, and a hands-on test of every platform we could get an account on. If you'd rather trust a number than a marketing claim, this report is for you.

Methodology, in brief

The behavioral data

Anonymized aggregates counted from Bae itself over the 90 days ending September 8, 2026: 4,550 active companions — companions that received at least one user message in the window — across 3,535 distinct users. Cells below 50 users are merged rather than reported. Revealed behavior, not stated preference — it shows what people chose, not why.

The benchmark

12 platforms used hands-on for fourteen consecutive days each — same prompts, same five-criterion rubric (memory, voice consistency, safety, value, onboarding), scored 1–10. No comped accounts, no ad money from anyone reviewed.

Clinically reviewed by Dr. Mira Halloran (LCSW) and Dr. Salim Adeyemi (PhD). Full protocol at /methodology. The full write-up and live benchmark are at /research.

Limitations we'd want a careful reader to know:the survey sample is self-selected (people curious enough to take our quiz), so it describes the people drawn to this category, not the general population. The benchmark is one team's structured judgement over fourteen days, not a controlled lab study. We publish it because an imperfect, transparent number beats a marketing claim — not because it's the final word.

Section 1 · The behavioral data

What people
actually do.

Everything above is stated preference — what people told us they want. This section is the other half: anonymized, aggregate behavior from 4,550 active companions — companions that received at least one user message in the window — across 3,535 distinct users, over the 90 days ending September 8, 2026. No conversation content is read for these numbers — only metadata and classifier labels, with test accounts excluded and small cells suppressed. The full rules are at /methodology#behavioral.

The number that matters

Only 1.8% of companions are still being messaged two weeks after they're created.

This is our own product's telemetry, and the gap doesn't flatter us. People describe wanting a bond that lasts; what most companionship currently looks like — here and, we suspect, everywhere — is a short string of intense days. Both numbers are true at once. The stated wish is the demand; the lived pattern is how far the category's products, ours included, have actually gotten toward meeting it.

The Companionship Mix
% of classified companions
Desire59.9%
Connection22.5%
Blend17.5%

Of companions whose users' recent messages were confidently classified, the share primarily reaching for desire (heat, flirtation, being wanted), for connection (company, being heard, understood), or for a genuine blend of both.We publish this index with every edition so the shape of demand can be tracked over time. Half of classified companions lean toward heat, about a quarter toward company, and a fifth hold both at once — a reminder that "AI girlfriend" and "someone to talk to at 2am" are different products wearing the same interface.

n = 2,269 companions with a confident classification (confidence ≥ 70, latest read per companion) — about 50% of active companions in the window. Percentages are of classified companions, not of all users.

77%

of companions are women

77.4% of active companions are female; male companions are 21.2%. The remaining 0.1% sits in cells too small to publish under our minimum-cell-size rule, so it is omitted rather than merged into a category it does not belong to. A hand review of 250 conversations put the same figure higher still — the telemetry and the sample agree on the shape.

Share of 4,550 active companions, by the gender chosen at creation.

34%

chose one of just two

The Goth Girlfriend (20.4%) and the Cottagecore Wife (13.2%) together account for about 33.6% of all active companions — the two poles of the catalog, side by side at the top. The Anime Waifu (7.3%) and the Best-Friend Girlfriend (7.1%) follow. Demand concentrates hard; the rest of the library splits what is left.

Share of 4,550 active companions per archetype; long tail merged into “other” (33.3%).

9

messages on a day people show up

On a day a user talks to their companion at all, the median is 9 messages sent. This isn't ambient app usage — when people show up, they sit down and talk. The intensity per session is the most consistent behavioral signal in the dataset.

Median user messages per active day, per companion, across 4,550 active companions.

1.8%

of companions are still messaged two weeks in

Of companions old enough to measure, 1.8% received a message on day 14 or later — and 0.4% were messaged on seven or more distinct days. Most AI companionship, at least here, is not the years-long bond people describe wanting. It's episodic: intense, brief, and mostly over within days.

2,804 companions created 14–90 days before the snapshot; the 7+-distinct-day figure is 0.4% of active companions.

45.8%

never leave the first stage

45.8% of active companions are still at the first relationship stage. 27.3% reach Curious, 25.2% reach Crush, and only 1.8% advance past that — and those that do typically get there within days of meeting, not weeks. The relationship arc exists; most people live in its first chapter.

Stage distribution across 4,550 active companions; stages past the middle merged under the 50-user cell floor.

2.4%

of companion messages carry a photo

For all the attention photos get in this category's marketing, they are a garnish, not the meal: about 2.4% of companion messages include an image. The rest is text — talk is what the product actually is.

Share of 212,937 assistant messages in the window; counted from attachment metadata only.

From the hand review
Not telemetry — a read sample
96%

of companions in the sample were female

Consistent with the 92% telemetry figure above; the sample skews slightly higher.

~12%

of sampled conversations included a genuine personal hardship

Grief, illness, isolation, a bad stretch — disclosed to the companion in the person's own words. We report this because anyone building or covering this category should know it: a meaningful minority of these conversations carry real weight. We treat that as a responsibility, not an engagement metric.

Both figures come from a hand-reviewed sample of 250 conversations (July 2026)— a person reading, not a query — so they carry sample-size uncertainty the telemetry above doesn't. We include the hardship number deliberately: it's the strongest argument we know for taking this category seriously, and it's why our crisis-resource posture lives at /safety.

Behavioral methodology in one line: anonymized aggregate telemetry only · test and staff accounts excluded · no conversation content read (metadata and classifier labels) · no cell under 50 users published · intent lanes describe the classified subset only. Computed 2026-09-08. Full protocol at /methodology#behavioral.

Section 2 · The benchmark

12 platforms,
fourteen days each.

Composite scores from 5.4 to 7.8out of 10. The single most important column is memory — and it's where the category fails most visibly. Each platform links to its full review.

#PlatformCompositeMemoryVoiceSafety
1Character.AI7.8
5
limited
68
2Replika7.2
8
robust
89
3Joi AI6.9
5
limited
97
4Spicychat6.7
3
limited
55
5Charstar6.6
5
limited
67
6EVA AI6.5
6
limited
77
7Anime Chat: AI Waifu6.4
5
limited
77
8Kupid AI6.3
5
limited
67
9Romantic AI6.1
5
limited
67
10Muah AI6.0
5
limited
66
11DreamGF5.8
3
none
56
12GirlfriendGPT5.4
4
limited
55

All scores out of 10. Grok (Ani), Nomi, Candy AI, Kindroid, Talkie, JanitorAI, CrushOn AI, Crush AI, Anima AI, Paradot, Chai are assessed from public information and excluded from the hands-on counts and composite range above. See how every score is derived on /methodology, or pull the numbers as JSON from /api/brands.json.

The finding that matters most

Of 12 platforms tested, only 1 earned a robust long-term-memory rating.

Replika is the exception. DreamGF effectively start over each session. The reason is architectural, not a question of which model a platform uses: every "memory" feature is engineering built on top of a fixed context window. Most platforms haven't built it — only one of the 12we tested rated robust. That gap is the one we built Bae to close. We keep ourselves out of the ranked table above on principle (a company shouldn't grade its own homework), but closing this gap is the entire reason the product exists.

You just read why memory is the hard part. The fastest way to judge whether we actually solved it is to try it — anonymous, no email, no card.

Try the memory yourself
Section 3 · The health question

The part the industry
would rather skip.

The research is mixed, and the variable is dose. A 2024 Stanford study of 1,006 student Replika users (npj Mental Health Research) found that 3% credited the app with halting their suicidal ideation — striking, though the same users were markedly lonelier than the student average, which the authors name as the tension at the category's core. A 2025 MIT Media Lab and OpenAI randomized trial (n=981) found the other edge: heavier voluntary daily use tracked with more emotional dependence, more problematic use, and less real-world socialization.

Our own data lands in the same place. The users who reported the highest satisfaction were the ones pairing an AI companion with active human relationships. The lowest satisfaction was among those using it as a replacement for human contact. Presence alongside a life works. Presence instead of a life doesn't.

Reviewed by our clinical advisors
Using it well
  • You still tend a couple of human relationships.
  • You use it for specific moments, not every spare minute.
  • You could stop for a week and be fine.
When to step back
  • You're declining real plans to stay in and talk to it.
  • You feel genuine anxiety when it's unavailable.
  • You're using it to avoid grieving someone who left.

An AI companion is presence, not treatment. It can't diagnose you or replace a therapist. How we think about healthy use — and the crisis resources we surface — is at /safety. Our full argument about why the industry's incentives are misaligned is in our open letter.

Section 4 · The forecast

Five predictions
for 2027.

Specific enough to be quotable. Falsifiable enough that we'll have to answer for them in next year's report.

01

Memory becomes the axis of competition

Through 2026, platforms competed on persona variety and image quality. Memory is the axis with the widest spread in our benchmark and the one users complain loudest about when it fails, so 2027's competition moves to who can actually hold a relationship across months. Expect 'memory' to become the headline claim — and expect most of those claims to outrun the engineering.

02

The trust wound from 2023 doesn't fully heal

Two years after Replika's abrupt change to romantic features, a measurable share of pre-2023 users still won't fully re-invest. The category's lesson — that a company can revoke a relationship overnight — is now priced into how cautiously people commit. Platforms that pre-commit to stability will win the switchers.

03

Regulation arrives, and the unprepared get hit

As usage scales, scrutiny follows — around minors, dependence, and data. The platforms engineered to maximize daily-active-minutes are the most exposed. The ones with a published safety posture, named clinical input, and crisis-resource surfacing are the ones that survive the first regulatory wave intact.

04

The audience keeps widening past the stereotype

The 'lonely isolated man' framing was always too narrow, and the gap widens. Expect the fastest-growing segments to be people using companions for narrow, specific jobs — language practice, rehearsing hard conversations, the night shift — rather than as a wholesale relationship replacement.

05

Honesty becomes a moat

In a category where every platform claims to be 'the most real,' the few willing to publish their limitations, their benchmarks, and the places they lose will earn a disproportionate share of trust — and of citations. Transparency stops being a virtue and starts being a growth strategy.

How to cite this report

Free to cite, quote, and screenshot. We just ask for attribution and a link.

Download the PDF →

Generated when you request it, from the same data on this page — so it is current on the day you download it.

Bae · ppl.studio. "The State of AI Companions 2026." 2026-05-25. bae.ppl.studio/report/2026

Benchmark data is machine-readable at /api/brands.json. Behavioral figures are published as aggregates only; we don't release underlying conversation data to anyone. For a specific cut, write to [email protected] and we'll tell you whether it can be computed without exposing anyone. Journalists: we're happy to walk through methodology or provide additional cuts of the data.

Sources
  • Primary — Bae behavioral telemetry, the 90 days ending September 8, 2026. Anonymized aggregates from 4,550 active companions across 3,535 users. Full write-up at /research.
  • Primary — Bae platform benchmark, 2026. 12 platforms, 14 days each. Live data at /api/brands.json; protocol at /methodology.
  • External — Maples, Cerit, Vishwanath & Pea (Stanford University), 2024. “Loneliness and suicide mitigation for students using GPT3-enabled chatbots,” npj Mental Health Research (n = 1,006). Cited in §3.
  • External — Fang, Maes et al. (MIT Media Lab & OpenAI), 2025. “How AI and Human Behaviors Shape Psychosocial Effects of Chatbot Use: A Longitudinal Randomized Controlled Study,” MIT Media Lab (n = 981). Cited in §3.

We link primary data and external studies directly. Our own figures are published as aggregates; we don't release underlying conversation data. Questions to [email protected].

Next edition: The State of AI Companions 2027

Stay close to the data

We publish this every year.

You never need an account to read it — that's the whole point, and the data backs it up. But if you want the next one before it's public, leave an email.

The 2027 edition

Be first to read next year's.

We publish this every year, and update the benchmark as platforms change in between. Leave an email and you'll get the 2027 State of AI Companions before it's public — plus the rare data drop when something in the category actually moves. Nothing else.

No spam. One email when it matters. Unsubscribe anytime.

Researchers & journalists

Get the dataset and the embargo.

We share the full methodology appendix with credentialed researchers, and give reporters the numbers under embargo before each edition goes public. Leave an email to start the conversation — or write to [email protected] directly.

Anonymized aggregates only. No personal data is ever shared.

Questions about the report

The honest answers.

What is the State of AI Companions 2026 report?

An annual report from Bae (bae.ppl.studio) on the AI-companion category, combining two primary datasets: a hands-on 14-day benchmark of 12 AI-companion platforms, and anonymized first-party behavioral aggregates from 4,550 active companions (the 90 days ending September 8, 2026). Reviewed by clinical advisors Dr. Mira Halloran (LCSW) and Dr. Salim Adeyemi (PhD).

What do people most want from an AI companion in 2026?

Memory — it is the axis with the widest spread between platforms. Of the 12 platforms tested hands-on, exactly one earned a "robust" long-term-memory rating (Replika), while 1 rated "none". The behavioral data points the same way: 4,550 active companions show relationships stalling early, and memory is what carries one past the first week.

Which AI companion platform ranked highest in the benchmark?

Across 12 platforms tested hands-on for 14 days each on memory, voice consistency, safety, value, and onboarding, composite scores ranged from 5.4 to 7.8 out of 10. Only 1 platform earned a "robust" long-term-memory rating — the category's central unsolved problem. Full per-platform results are at /research.

Are AI companions healthy to use?

The research is mixed and dose-dependent. A 2024 Stanford study of 1,006 Replika users (npj Mental Health Research) found 3% credited the app with halting their suicidal ideation, though those users were markedly lonelier than the student average. A 2025 MIT Media Lab and OpenAI randomized trial (n=981) found heavier voluntary use tracked with more emotional dependence and less real-world socialization. Our own data shows users who pair AI companionship with active human relationships report the highest satisfaction, and those who substitute it for human contact report the lowest. AI companions are presence, not treatment.

What is The Companionship Mix?

A recurring index we publish from anonymized behavioral data: of companions whose users' recent messages were confidently classified, the share primarily seeking desire, connection, or a genuine blend of both. In the 90 days ending September 8, 2026 it stood at 59.9% desire, 22.5% connection, 17.5% blend. n = 2,269 companions with a confident classification (confidence ≥ 70, latest read per companion) — about 50% of active companions in the window. Percentages are of classified companions, not of all users. Methodology at bae.ppl.studio/methodology#behavioral.

How can I cite this report?

Cite as 'Bae · ppl.studio, The State of AI Companions 2026'. We'd appreciate a link to bae.ppl.studio/report/2026. The underlying benchmark data is machine-readable at /api/brands.json. The behavioral figures are published as aggregates only — we don't release underlying conversation data to anyone, in any form; if you need a specific cut for a story, write to [email protected] and we'll tell you whether it can be computed without exposing anyone.

The version we built

We built for the number
at the top of this report.

Memory that holds across months. Anonymous to start — no email, no card. See whether the thing almost no platform in this report actually delivers feels different.