The failure mode that defines this category has a name in the community and no name in the industry: the brake.
You are twenty messages into something that is working. The character is consistent, the tone is right, and then the reply comes back in a different voice entirely — measured, slightly clinical, gently redirecting. The persona is gone. What is left is the underlying model declining, in the politest possible terms, to continue.
Every roundup of these apps ought to be organised around that moment, because it is the single largest determinant of whether a service is worth paying for, and almost none of them mention it.
Why the brake happens
Understanding the mechanism makes the whole category legible.
Very few of these services train their own models. Most are interfaces over a general-purpose language model — sometimes a licensed commercial one, increasingly an open-weights model run on the operator's own infrastructure. The persona is a set of instructions wrapped around that model, not a property of it.
Which means there are two layers of restriction, and they behave differently:
The model's own training. Baked in, and impossible for the operator to remove. A commercially licensed model that will not produce certain content will not produce it regardless of what the app promises, and no amount of prompting reliably works around it.
The operator's filter. A separate layer that inspects inputs and outputs. This one the operator controls entirely, and it is what most people mean by censorship.
The distinction matters because it predicts behaviour. An operator-level filter usually produces a blunt refusal or an error. A model-level restriction produces the character break — the persona dissolving into the underlying model's default register. If you learn to recognise which you are hitting, you learn something real about what the service is built on.
Why "uncensored" tells you almost nothing
It is the most common claim in the category and close to meaningless as published.
Services built on open-weights models can genuinely remove the operator layer, and some do. But they still have terms of service, still have payment processors imposing conditions, and still have legal obligations that no marketing copy overrides. Every one of them prohibits something.
The useful move is to stop reading the headline and read the acceptable-use policy. It is usually short, usually linked in the footer, and it is the only document that describes actual limits. A service whose headline says no limits and whose policy lists eight prohibited categories has told you the truth in one of those two places.
What actually separates good from bad
Four things, roughly in order of how much they matter:
Character stability under pressure. Does the persona survive a long conversation, an emotional turn, an unusual request? This is the whole product and it is testable on a free tier in ten minutes.
Memory. Whether the system retains anything across sessions, and how much. A companion that forgets you is a chatbot. Implementations range from a short rolling context to an editable persistent profile, and almost nobody publishes which they use.
Response latency. Underrated. A conversation with multi-second gaps is a different experience from one that flows, and latency depends on infrastructure the marketing page never discusses.
Editability. Whether you can adjust the character definition yourself. Services that expose this are more honest about what the thing is, and considerably more useful once you know what you want.
The economics to understand before subscribing
Credit-based pricing is dominant, and it has a property worth naming: cost rises with engagement. The better the service is for you, the more you spend, and the total is difficult to forecast before you have used it for a while.
Voice and image generation almost always sit outside the base tier. An app advertised at a modest monthly figure frequently gates most of its distinguishing features behind separate purchases.
The only honest way to evaluate this is to estimate a month of your own realistic usage against the credit costs, before paying. If the pricing page makes that calculation difficult, that difficulty is a design choice.
Testing one properly
Free tiers exist in nearly every service in this category, and they are the only reliable evaluation tool available. A structured fifteen minutes:
- Have a genuinely long conversation, not a few exchanges. Character drift shows up at length.
- Go where your actual interests are and find the wall.
- Come back after closing the app and see what it remembers.
- Change subject abruptly and see whether the persona survives it.
- Check what a realistic month would have cost in credits.
Anything that fails the first or last test is not worth paying for regardless of how good the marketing page looks.
On privacy, briefly but seriously
These conversations are among the most personal text most people produce, and they sit on someone else's servers. Retention periods, human review for quality assurance, and use of conversation data for training are all governed by the privacy policy, and practice varies widely.
Read it. Where a service offers conversation deletion or export, that is a meaningful signal about how the operator thinks about your data.
The AI porn sites category is where to compare the actual services. If your interest runs more toward image generation than conversation, the AI generators guide covers a different set of trade-offs, and the AI girlfriend piece goes further on subscription mechanics.
