AI tell · grade D
Making ChatGPT sound less like ChatGPT: what a prompt can and cannot fix
Instructions for getting ChatGPT to sound less robotic fall into two categories, and only one of them tends to work. The first asks the model to avoid certain words and stock phrases. When we checked our own twelve-model logs for that exact vocabulary cluster, we found close to none of it: the classic connector words and inflated phrasing were already absent across every model we sampled. Telling a model to stop doing something it was not doing does not change the output. The second category is behavioral: opening with praise for the question, scaffolding any answer into headers and bullets, addressing the reader with heavy second-person coaching language. Those showed up in nine to twelve of the twelve models across every conversation topic in our sample, finance questions, scheduling questions, emotional ones, because they read as closer to a sampling default than a word choice, and a wording tweak has a much harder job overriding a near-universal default than it does swapping one word for another.
What it is
Split any given prompt instruction into which category it belongs to. Lexical instructions tell the model to avoid specific words or stock phrases; these are the most common instructions circulating, and our own logs show the words they target were already essentially unused across all twelve models we sampled, which means the instruction has nothing left to fix. Behavioral instructions ask the model to change how it opens, structures, or closes a reply. These target reflexes we measured at 9 to 12 out of 12 models, holding up across the full range of topics in our corpus, which makes them a far larger share of what actually reads as generated, and a far harder thing for a single line of prompt wording to override.
How often it shows up
The vocabulary cluster most style instructions target, connector words used to stack claims and phrasing that inflates plain statements into something grander, was searched for directly in our full corpus and found close to zero times across all twelve models. Meanwhile the validation opener appeared in all twelve models, structural scaffolding over casual exchanges appeared in all twelve, dense second-person address appeared in all twelve, a cheerleading closer appeared in ten of twelve, and a closing offer to do more or a forced-choice question appeared in nine of twelve, all recurring across every topic in the corpus rather than clustering on one kind of question.
Why readers flag it
A reader does not consciously catalogue vocabulary; they register rhythm and shape. A reply that dutifully avoids a banned word list but still opens by praising the question and closes with an offer to do more reads exactly as generated as one that never got the instruction, because the part a reader actually notices was never targeted. This is the honest reason the same prompt copypasta keeps circulating and keeps under-delivering: it treats a problem that mostly is not there while leaving the actual majority of the problem untouched.
The fix
Give the model a structural constraint it can actually obey rather than a mood. An instruction to answer in continuous prose with no headers and no bullets removes the entire scaffolding reflex far more reliably than an instruction to sound more natural. Cut the ends yourself after the fact: delete an opening sentence that grades your question and a closing paragraph that offers to do more, since our logs show that reflex firing in the large majority of models across every topic we logged, not just a narrow slice of question types. Where a reasoning-mode toggle exists, use it deliberately. In our logs, the same model's hedge-to-booster ratio moved from 4.08 in non-thinking mode to 0.63 in thinking mode. We have not tested how far a wording change alone moves that ratio, so lean on the toggle rather than assuming a phrasing tweak does the same job. Expect real limits from wording alone: an edit pass after the model has already written, which is the actual majority of what needs fixing, does more than a system-prompt line asking it to sound different.
Examples
That's a great question! Here's the thing: casual writing is just clear writing. - Keep sentences short. - Cut jargon. - Read it aloud. Let me know if you want more tips!
Casual writing is just clear writing. Keep sentences short. Cut jargon. Read it aloud.
You're absolutely right to ask. Getting started is simple. Set a plan, stay consistent, and track results. You've got this!
Getting started is simple: set a plan, stay consistent, and track results.
Questions
Is there a way to make AI sound less like AI?
Partly, and only if the instruction targets the right layer. Wording changes can shift vocabulary, which our logs suggest was rarely the actual problem. Reflexes like a validation opener or a forced structure are much harder to prompt away because they showed up in nearly every model we measured, across every topic in the corpus, not one narrow kind of question.
How do I get ChatGPT to not sound like ChatGPT?
Give it a concrete structural constraint (no headers, no bullets, no closing offer to do more) instead of a mood word like natural or casual. Then plan to cut the opener and closer yourself afterward, since both showed up across nearly every model in our logs, on almost every topic we sampled.
What makes ChatGPT sound like ChatGPT?
Mostly reflexes, not word choice: an opener that praises the question before answering it, an answer scaffolded into headers and bullets even for a casual topic, and a close that hands the exchange back with an offer to do more. All three appeared across nine to twelve of the twelve models we logged.
Why does ChatGPT sound robotic?
Not because of the words it picks; the classic AI-vocabulary cluster was close to absent in our own logs across every model. It reads as robotic because of near-universal structural and address-pattern reflexes that persist across every topic we logged, from finance questions to personal scheduling.
What is the human prompt for ChatGPT?
There is not one line that fixes this, because the biggest share of what reads as generated is behavioral rather than lexical. A structural constraint helps more than a mood word, and an edit pass after generation catches what the prompt could not.
underslop marks this tell and the rest of the consensus set in any draft you paste, using the same open lint that scores it here. The free tier caps at 500 words a run, and your text is held for 0 seconds.