Why AI replies sound like AI

6 min read

Everyone can spot a generated comment within a few words, but few people can say exactly what gave it away. The tells are consistent, and most of them are not fixable with a tone setting.

The tells are structural, not stylistic

Ask someone why a reply felt automated and they will usually say it was "too generic". That is the symptom rather than the cause. The actual giveaways are structural and they repeat across tools.

Generated replies restate the original post before responding to it. They open by validating ("Great point about..."), they hedge symmetrically so no sentence carries risk, and they resolve into a tidy summary nobody asked for. They are also, almost always, slightly too long for the thing they are replying to.

Real replies do none of that. People jump straight in, assume shared context, leave thoughts unfinished, and often contribute one specific fact rather than a balanced overview.

  • Restating the post before answering it
  • Opening with validation of the other person
  • Symmetrical hedging so nothing is actually claimed
  • A summary sentence that adds nothing
  • Length disproportionate to the original post

Why a tone dropdown cannot fix this

Most tools expose tone as a setting: professional, friendly, casual, witty. The problem is that tone is a global adjustment applied to output that is already structurally generic. You get the same skeleton wearing a different coat.

It also fails at scale in a way that is easy to miss. Tone presets are shared across every customer of that tool. If a hundred people pick "friendly", their replies rhyme with each other. Anyone who reads a lot of comments starts recognising the shape, and once you see it you cannot unsee it.

The deeper issue is that tone describes a register, and voice is made of much smaller things: which words a person reaches for, how long their sentences run, whether they use semicolons, whether they start with "honestly", whether they ever use an exclamation mark.

Voice is mostly the words a person would never use

This is the part that surprises people. Capturing someone’s voice is less about reproducing their favourite phrases and more about excluding the ones that would never appear.

Most people have a substantial anti vocabulary. Words like "leverage", "delve", "robust", "seamless", "game changer", "in today’s landscape". A model with no constraint reaches for these constantly because they are statistically common in the training data for professional writing. One of them in a reply is enough to break the illusion.

A voice model that records what a person avoids is doing more work than one that only records what they like.

Context is the other half

Even a perfectly matched voice sounds wrong if the reply does not belong in the conversation. The most common failure is replying to the topic rather than to the post: someone writes about a specific onboarding problem, and the reply is about onboarding in general.

Good replies are narrow. They pick one thread from what was said and pull it. That requires actually reading the thing, which is why tools that select a target and generate in one pass tend to produce comments that are technically relevant and obviously hollow.

How we approach it, and where the limits are

Quillen builds a persona from a person’s own writing rather than asking them to pick a register. It records phrasing, rhythm, and an explicit anti vocabulary, and each generated reply is checked against that persona before it reaches the queue. If a draft fails that check it is regenerated rather than shipped.

What that does not do is remove the human. Drafts wait for approval, and the realistic workflow is that you edit some of them. That is not a limitation we are hiding: a voice engine gets you from a blank box to something recognisably yours, and the last ten percent is judgement that belongs to the person whose name is on the account.

Anyone claiming their output never needs editing is describing a product that does not exist.

A test you can run in thirty seconds

Take any tool you are evaluating and generate two replies for two different people to the same post. If you can swap the names and nobody would notice, the tool has a tone setting rather than a voice model.

Then read the output aloud. Voice mismatches that are invisible on screen are obvious in the ear, because you already know what the person sounds like.