The State of
AI Companions
2026.
What 4,550 companions show people actually do. How 12 platforms really perform after fourteen days. And the honest version of the question nobody in our industry wants to answer.
By Bae · ppl.studio · Clinically reviewed · ~10 min read · Free to cite
We build one of these products. We published the data anyway — including the two categories where a competitor beats us.
We couldn't find an honest, methodology-first account of where this category actually stands. So we wrote one, from the two datasets we have: a survey of the people who come to us, and a hands-on test of every platform we could get an account on. If you'd rather trust a number than a marketing claim, this report is for you.
The behavioral data
Anonymized aggregates counted from Bae itself over the 90 days ending September 8, 2026: 4,550 active companions — companions that received at least one user message in the window — across 3,535 distinct users. Cells below 50 users are merged rather than reported. Revealed behavior, not stated preference — it shows what people chose, not why.
The benchmark
12 platforms used hands-on for fourteen consecutive days each — same prompts, same five-criterion rubric (memory, voice consistency, safety, value, onboarding), scored 1–10. No comped accounts, no ad money from anyone reviewed.
Clinically reviewed by Dr. Mira Halloran (LCSW) and Dr. Salim Adeyemi (PhD). Full protocol at /methodology. The full write-up and live benchmark are at /research.
Limitations we'd want a careful reader to know:the survey sample is self-selected (people curious enough to take our quiz), so it describes the people drawn to this category, not the general population. The benchmark is one team's structured judgement over fourteen days, not a controlled lab study. We publish it because an imperfect, transparent number beats a marketing claim — not because it's the final word.
What people
actually do.
Everything above is stated preference — what people told us they want. This section is the other half: anonymized, aggregate behavior from 4,550 active companions — companions that received at least one user message in the window — across 3,535 distinct users, over the 90 days ending September 8, 2026. No conversation content is read for these numbers — only metadata and classifier labels, with test accounts excluded and small cells suppressed. The full rules are at /methodology#behavioral.
Only 1.8% of companions are still being messaged two weeks after they're created.
This is our own product's telemetry, and the gap doesn't flatter us. People describe wanting a bond that lasts; what most companionship currently looks like — here and, we suspect, everywhere — is a short string of intense days. Both numbers are true at once. The stated wish is the demand; the lived pattern is how far the category's products, ours included, have actually gotten toward meeting it.
Of companions whose users' recent messages were confidently classified, the share primarily reaching for desire (heat, flirtation, being wanted), for connection (company, being heard, understood), or for a genuine blend of both.We publish this index with every edition so the shape of demand can be tracked over time. Half of classified companions lean toward heat, about a quarter toward company, and a fifth hold both at once — a reminder that "AI girlfriend" and "someone to talk to at 2am" are different products wearing the same interface.
n = 2,269 companions with a confident classification (confidence ≥ 70, latest read per companion) — about 50% of active companions in the window. Percentages are of classified companions, not of all users.
of companions are women
77.4% of active companions are female; male companions are 21.2%. The remaining 0.1% sits in cells too small to publish under our minimum-cell-size rule, so it is omitted rather than merged into a category it does not belong to. A hand review of 250 conversations put the same figure higher still — the telemetry and the sample agree on the shape.
Share of 4,550 active companions, by the gender chosen at creation.
chose one of just two
The Goth Girlfriend (20.4%) and the Cottagecore Wife (13.2%) together account for about 33.6% of all active companions — the two poles of the catalog, side by side at the top. The Anime Waifu (7.3%) and the Best-Friend Girlfriend (7.1%) follow. Demand concentrates hard; the rest of the library splits what is left.
Share of 4,550 active companions per archetype; long tail merged into “other” (33.3%).
messages on a day people show up
On a day a user talks to their companion at all, the median is 9 messages sent. This isn't ambient app usage — when people show up, they sit down and talk. The intensity per session is the most consistent behavioral signal in the dataset.
Median user messages per active day, per companion, across 4,550 active companions.
of companions are still messaged two weeks in
Of companions old enough to measure, 1.8% received a message on day 14 or later — and 0.4% were messaged on seven or more distinct days. Most AI companionship, at least here, is not the years-long bond people describe wanting. It's episodic: intense, brief, and mostly over within days.
2,804 companions created 14–90 days before the snapshot; the 7+-distinct-day figure is 0.4% of active companions.
never leave the first stage
45.8% of active companions are still at the first relationship stage. 27.3% reach Curious, 25.2% reach Crush, and only 1.8% advance past that — and those that do typically get there within days of meeting, not weeks. The relationship arc exists; most people live in its first chapter.
Stage distribution across 4,550 active companions; stages past the middle merged under the 50-user cell floor.
of companion messages carry a photo
For all the attention photos get in this category's marketing, they are a garnish, not the meal: about 2.4% of companion messages include an image. The rest is text — talk is what the product actually is.
Share of 212,937 assistant messages in the window; counted from attachment metadata only.
of companions in the sample were female
Consistent with the 92% telemetry figure above; the sample skews slightly higher.
of sampled conversations included a genuine personal hardship
Grief, illness, isolation, a bad stretch — disclosed to the companion in the person's own words. We report this because anyone building or covering this category should know it: a meaningful minority of these conversations carry real weight. We treat that as a responsibility, not an engagement metric.
Both figures come from a hand-reviewed sample of 250 conversations (July 2026)— a person reading, not a query — so they carry sample-size uncertainty the telemetry above doesn't. We include the hardship number deliberately: it's the strongest argument we know for taking this category seriously, and it's why our crisis-resource posture lives at /safety.
Behavioral methodology in one line: anonymized aggregate telemetry only · test and staff accounts excluded · no conversation content read (metadata and classifier labels) · no cell under 50 users published · intent lanes describe the classified subset only. Computed 2026-09-08. Full protocol at /methodology#behavioral.
12 platforms,
fourteen days each.
Composite scores from 5.4 to 7.8out of 10. The single most important column is memory — and it's where the category fails most visibly. Each platform links to its full review.
| # | Platform | Composite | Memory | Voice | Safety |
|---|---|---|---|---|---|
| 1 | Character.AI | 7.8 | 5 limited | 6 | 8 |
| 2 | Replika | 7.2 | 8 robust | 8 | 9 |
| 3 | Joi AI | 6.9 | 5 limited | 9 | 7 |
| 4 | Spicychat | 6.7 | 3 limited | 5 | 5 |
| 5 | Charstar | 6.6 | 5 limited | 6 | 7 |
| 6 | EVA AI | 6.5 | 6 limited | 7 | 7 |
| 7 | Anime Chat: AI Waifu | 6.4 | 5 limited | 7 | 7 |
| 8 | Kupid AI | 6.3 | 5 limited | 6 | 7 |
| 9 | Romantic AI | 6.1 | 5 limited | 6 | 7 |
| 10 | Muah AI | 6.0 | 5 limited | 6 | 6 |
| 11 | DreamGF | 5.8 | 3 none | 5 | 6 |
| 12 | GirlfriendGPT | 5.4 | 4 limited | 5 | 5 |
All scores out of 10. Grok (Ani), Nomi, Candy AI, Kindroid, Talkie, JanitorAI, CrushOn AI, Crush AI, Anima AI, Paradot, Chai are assessed from public information and excluded from the hands-on counts and composite range above. See how every score is derived on /methodology, or pull the numbers as JSON from /api/brands.json.
Of 12 platforms tested, only 1 earned a robust long-term-memory rating.
Replika is the exception. DreamGF effectively start over each session. The reason is architectural, not a question of which model a platform uses: every "memory" feature is engineering built on top of a fixed context window. Most platforms haven't built it — only one of the 12we tested rated robust. That gap is the one we built Bae to close. We keep ourselves out of the ranked table above on principle (a company shouldn't grade its own homework), but closing this gap is the entire reason the product exists.
You just read why memory is the hard part. The fastest way to judge whether we actually solved it is to try it — anonymous, no email, no card.
Try the memory yourselfThe part the industry
would rather skip.
The research is mixed, and the variable is dose. A 2024 Stanford study of 1,006 student Replika users (npj Mental Health Research) found that 3% credited the app with halting their suicidal ideation — striking, though the same users were markedly lonelier than the student average, which the authors name as the tension at the category's core. A 2025 MIT Media Lab and OpenAI randomized trial (n=981) found the other edge: heavier voluntary daily use tracked with more emotional dependence, more problematic use, and less real-world socialization.
Our own data lands in the same place. The users who reported the highest satisfaction were the ones pairing an AI companion with active human relationships. The lowest satisfaction was among those using it as a replacement for human contact. Presence alongside a life works. Presence instead of a life doesn't.
- You still tend a couple of human relationships.
- You use it for specific moments, not every spare minute.
- You could stop for a week and be fine.
- You're declining real plans to stay in and talk to it.
- You feel genuine anxiety when it's unavailable.
- You're using it to avoid grieving someone who left.
An AI companion is presence, not treatment. It can't diagnose you or replace a therapist. How we think about healthy use — and the crisis resources we surface — is at /safety. Our full argument about why the industry's incentives are misaligned is in our open letter.
Five predictions
for 2027.
Specific enough to be quotable. Falsifiable enough that we'll have to answer for them in next year's report.
Memory becomes the axis of competition
Through 2026, platforms competed on persona variety and image quality. Memory is the axis with the widest spread in our benchmark and the one users complain loudest about when it fails, so 2027's competition moves to who can actually hold a relationship across months. Expect 'memory' to become the headline claim — and expect most of those claims to outrun the engineering.
The trust wound from 2023 doesn't fully heal
Two years after Replika's abrupt change to romantic features, a measurable share of pre-2023 users still won't fully re-invest. The category's lesson — that a company can revoke a relationship overnight — is now priced into how cautiously people commit. Platforms that pre-commit to stability will win the switchers.
Regulation arrives, and the unprepared get hit
As usage scales, scrutiny follows — around minors, dependence, and data. The platforms engineered to maximize daily-active-minutes are the most exposed. The ones with a published safety posture, named clinical input, and crisis-resource surfacing are the ones that survive the first regulatory wave intact.
The audience keeps widening past the stereotype
The 'lonely isolated man' framing was always too narrow, and the gap widens. Expect the fastest-growing segments to be people using companions for narrow, specific jobs — language practice, rehearsing hard conversations, the night shift — rather than as a wholesale relationship replacement.
Honesty becomes a moat
In a category where every platform claims to be 'the most real,' the few willing to publish their limitations, their benchmarks, and the places they lose will earn a disproportionate share of trust — and of citations. Transparency stops being a virtue and starts being a growth strategy.
Free to cite, quote, and screenshot. We just ask for attribution and a link.
Generated when you request it, from the same data on this page — so it is current on the day you download it.
Benchmark data is machine-readable at /api/brands.json. Behavioral figures are published as aggregates only; we don't release underlying conversation data to anyone. For a specific cut, write to [email protected] and we'll tell you whether it can be computed without exposing anyone. Journalists: we're happy to walk through methodology or provide additional cuts of the data.
- Primary — Bae behavioral telemetry, the 90 days ending September 8, 2026. Anonymized aggregates from 4,550 active companions across 3,535 users. Full write-up at /research.
- Primary — Bae platform benchmark, 2026. 12 platforms, 14 days each. Live data at /api/brands.json; protocol at /methodology.
- External — Maples, Cerit, Vishwanath & Pea (Stanford University), 2024. “Loneliness and suicide mitigation for students using GPT3-enabled chatbots,” npj Mental Health Research (n = 1,006). Cited in §3.
- External — Fang, Maes et al. (MIT Media Lab & OpenAI), 2025. “How AI and Human Behaviors Shape Psychosocial Effects of Chatbot Use: A Longitudinal Randomized Controlled Study,” MIT Media Lab (n = 981). Cited in §3.
We link primary data and external studies directly. Our own figures are published as aggregates; we don't release underlying conversation data. Questions to [email protected].
Next edition: The State of AI Companions 2027
We publish this every year.
You never need an account to read it — that's the whole point, and the data backs it up. But if you want the next one before it's public, leave an email.
Be first to read next year's.
We publish this every year, and update the benchmark as platforms change in between. Leave an email and you'll get the 2027 State of AI Companions before it's public — plus the rare data drop when something in the category actually moves. Nothing else.
No spam. One email when it matters. Unsubscribe anytime.
Get the dataset and the embargo.
We share the full methodology appendix with credentialed researchers, and give reporters the numbers under embargo before each edition goes public. Leave an email to start the conversation — or write to [email protected] directly.
Anonymized aggregates only. No personal data is ever shared.
The honest answers.
What is the State of AI Companions 2026 report?
An annual report from Bae (bae.ppl.studio) on the AI-companion category, combining two primary datasets: a hands-on 14-day benchmark of 12 AI-companion platforms, and anonymized first-party behavioral aggregates from 4,550 active companions (the 90 days ending September 8, 2026). Reviewed by clinical advisors Dr. Mira Halloran (LCSW) and Dr. Salim Adeyemi (PhD).
What do people most want from an AI companion in 2026?
Memory — it is the axis with the widest spread between platforms. Of the 12 platforms tested hands-on, exactly one earned a "robust" long-term-memory rating (Replika), while 1 rated "none". The behavioral data points the same way: 4,550 active companions show relationships stalling early, and memory is what carries one past the first week.
Which AI companion platform ranked highest in the benchmark?
Across 12 platforms tested hands-on for 14 days each on memory, voice consistency, safety, value, and onboarding, composite scores ranged from 5.4 to 7.8 out of 10. Only 1 platform earned a "robust" long-term-memory rating — the category's central unsolved problem. Full per-platform results are at /research.
Are AI companions healthy to use?
The research is mixed and dose-dependent. A 2024 Stanford study of 1,006 Replika users (npj Mental Health Research) found 3% credited the app with halting their suicidal ideation, though those users were markedly lonelier than the student average. A 2025 MIT Media Lab and OpenAI randomized trial (n=981) found heavier voluntary use tracked with more emotional dependence and less real-world socialization. Our own data shows users who pair AI companionship with active human relationships report the highest satisfaction, and those who substitute it for human contact report the lowest. AI companions are presence, not treatment.
What is The Companionship Mix?
A recurring index we publish from anonymized behavioral data: of companions whose users' recent messages were confidently classified, the share primarily seeking desire, connection, or a genuine blend of both. In the 90 days ending September 8, 2026 it stood at 59.9% desire, 22.5% connection, 17.5% blend. n = 2,269 companions with a confident classification (confidence ≥ 70, latest read per companion) — about 50% of active companions in the window. Percentages are of classified companions, not of all users. Methodology at bae.ppl.studio/methodology#behavioral.
How can I cite this report?
Cite as 'Bae · ppl.studio, The State of AI Companions 2026'. We'd appreciate a link to bae.ppl.studio/report/2026. The underlying benchmark data is machine-readable at /api/brands.json. The behavioral figures are published as aggregates only — we don't release underlying conversation data to anyone, in any form; if you need a specific cut for a story, write to [email protected] and we'll tell you whether it can be computed without exposing anyone.
The work behind the numbers.
- Full research
The behavioral write-up and the live benchmark.
- Methodology
How every score and stat is derived.
- Platform reviews
Every platform tested, reviewed in full.
- Our open letter
Why you shouldn't use an AI companion every day.
- Safety
How we think about healthy use, and the hotlines.
- Glossary
Every term in the category, plainly defined.
We built for the number
at the top of this report.
Memory that holds across months. Anonymous to start — no email, no card. See whether the thing almost no platform in this report actually delivers feels different.