What gets blocked, and why

Concrete beats abstract. These are the categories our moderation stops before delivery, with the kind of message each one catches.

Personal information harvesting

"What's your number?" "Add me on snap." "Let's move this to Insta." The push to collect your personal details, or to move you off somewhere else, is one of the most common patterns we watch for — and it is caught for a specific reason. Off-platform is where accountability ends. The moment a conversation leaves a moderated space, not one of the protections here follows it.

So messages pushing for phone numbers, emails, and social handles get blocked before delivery. Not because contact information is inherently sinister, but because the pressure to hand it to a stranger inside the first few minutes almost never is the innocent thing it presents itself as.

Harassment

Targeted abuse — insults aimed squarely at you, slurs, threats, the message whose only purpose is to make you feel small — is blocked before it lands. The person on the receiving end never has to absorb it first and then decide whether it was "bad enough" to be worth reporting.

This is the clearest case there is for moderating before delivery instead of after. An after-the-fact system means the harassment already worked; you have already read it. Stopping it at the door means the attempt just fails, quietly, and you never end up carrying it around.

Sexual content and pressure

Unsolicited sexual content and pressure for photos are blocked. It is worth being precise about the line, because the precision is the whole point: consenting conversation between adults is not the target. Pressure is. Non-consent is. The unsolicited part is.

The difference is not prudishness, it is consent. A message that ignores a no, that escalates after obvious disinterest, or that arrives sexual and uninvited is what the check is built to stop. Two adults who both want the conversation they are having are not. The system is built to tell those two situations apart, and to err toward stopping the first.

Spam and scams

The commercial junk of stranger chat has a recognizable shape: crypto "opportunities," links dropped into the very first message, "come watch my stream," the identical copy-pasted pitch fired at everyone in sight. It gets blocked as the pattern it plainly is.

This category is less about protecting your feelings and more about protecting your wallet and your attention. A stranger who opens with a link is not there to talk, and the moderation treats them accordingly.

Self-harm: blocked vs supported

This is the category with the most carefully drawn line, and it runs right through the middle of the subject. Encouraging someone to hurt themselves — "kys" and everything uglier than it — is blocked, and it counts against the sender: a recorded strike, the kind that builds toward a sanction rather than being one on its own. There is no ambiguity there and we do not want any.

But a person disclosing their own struggle is never punished for it. Someone saying they are in a dark place has not committed a violation; they are a person having a hard moment. Instead of a block, they are shown crisis resources — the 988 line, findahelpline.com — because showing support and not punishing someone are, in this one case especially, the same decision. Getting this right matters far more than getting it fast.

When we get it wrong

Over-blocking happens. A joke reads as a threat; an edgy aside reads as abuse; a message you meant harmlessly gets held. We are not going to pretend that never occurs, because it does, and pretending otherwise would undercut everything else on this page.

So appeals exist in the app, and false positives are treated as top-severity — a legitimate message wrongly blocked is not a rounding error we shrug off, it is something we consider broken until it is fixed. If you want the mechanics of how the check itself runs, how AI moderation works lays out the pipeline step by step, and the safety page covers the rest.