magicshare
FeaturesPricingBlogGet startedLog in
Get started
FeaturesPricingBlogGet startedLog in
magicshare
FeaturesPricingBlogGet startedLog in
Get started
FeaturesPricingBlogGet startedLog in
magicshare
FeaturesPricingBlogGet startedLog in
Get started
FeaturesPricingBlogGet startedLog in

← blog

How to Make AI Content Sound Human: A 10-Minute Pass

2026-08-27·17 min readai-contenteditingcontent-strategybrand-voiceseo

Readers catch AI drafts on a small, stable set of tells, and almost none of them are individual words. In a 2025 study, five annotators who simply use LLMs every day read 300 non-fiction articles and, by majority vote, misclassified exactly one, outperforming most commercial detectors. What gives a draft away is shape: scaffolding intros, uniform sentence length, thin punctuation, vague attribution, no quoted humans. Each of those has a mechanical fix, and the whole pass fits in about ten minutes if you run it in the right order: structure, then rhythm, then sentences, then words.

That ordering matters more than any word list you have been handed. It is also the opposite of most advice on this subject, which starts with "delete the em dashes" and stops there.

Readers catch it first. Detectors show up later.

The clearest recent case is a publishing one. In January 2026, a self-identified book editor posted on Reddit that the novel Shy Girl showed hallmarks of having been written with a large language model. A 2.5-hour video essay by the reviewer frankie's shelf, titled "i'm pretty sure this book is ai slop," passed 1.2 million views by March 2026. Hachette then cancelled the U.S. release and discontinued the U.K. edition after what it described as a lengthy investigation. The author, Mia Ballard, told the New York Times that a freelance editor had added AI-generated content without her knowledge.

Notice the sequence. Pangram's founder Max Spero eventually found evidence that 78% of the book was AI-generated, but that verdict arrived after months of readers reacting to the prose. The detector confirmed a judgement the audience had already made.

That pattern repeats across the year's other cases. Poynter's round-up of 2026 AI-writing controversies found that what people actually flagged were "strange stylistic choices and clunky sentences", and that proving authorship afterwards was consistently hard. For a founder shipping a blog post, the practical translation is blunt: the audit happens in your reader's head, in the first two paragraphs, before anyone runs a tool. That is the audience your editing pass is competing against.

Being AI-drafted is not the problem. Reading unedited is.

It helps to be clear about what you are actually defending against, because it is not disclosure and it is not Google.

Ahrefs sampled 900,000 newly created English web pages in April 2025 and found that 74.2% contained AI-generated content: 2.5% pure AI, 25.8% pure human, and 71.7% some mix of the two. If three in four new pages have a model somewhere in their history, "this was AI-drafted" carries roughly as much signal as "this was typed." The reputational risk sits entirely on the other side of the line: publishing something that reads like nobody looked at it.

Your readers also have a word for that now. Merriam-Webster named "slop" its 2025 Word of the Year on 15 December 2025, defining it as "digital content of low quality that is produced usually in quantity by means of artificial intelligence". When a category gets a short, ugly, memorable name, people start applying it faster and with less charity. A reader who would once have thought "this is a bit dull" now has a label ready in one syllable.

What the 300-article study adds is how they apply it. The annotators leaned on "specific lexical clues ('AI vocabulary')" but also on "more complex phenomena within the text (e.g., formality, originality, clarity) that are challenging to assess for automatic detectors." Formality, originality and clarity are not vocabulary problems. They are structure, evidence and rhythm problems, which is exactly why a word-swap pass leaves a draft still sounding machine-made.

And one more boundary worth setting before the checklist: this is not a compliance exercise. Google's spam policies define scaled content abuse as generating many pages primarily to manipulate rankings, "no matter how it's created" — method-agnostic by design, and updated most recently on 15 May 2026 to also cover attempts to manipulate generative AI responses in Search. We have written separately about what Google's scaled content abuse policy really bans. The voice pass below is for readers, not for the algorithm.

The em-dash advice you have read has expired

Before the checklist, one piece of received wisdom needs retiring, because founders keep spending their editing minutes on it.

Em dashes were the 2024 tell. They are now closer to a model fingerprint than a species fingerprint. An arXiv preprint measured twelve instruction-tuned models across roughly 240,000 words and found rates ranging from 0.00 per 1,000 words for both Llama models to 10.62 for GPT-4.1, with Claude Opus 4.6 at 9.09 and Gemini 2.5 Flash at 1.28. The same paper put a human baseline at 3.23 per 1,000 words (median 3.83, range 0.33–17.12) across eight published essays. Two consequences for your edit: the target is human range, not zero, and the density of the tell depends on which model drafted the page.

The tell has also decayed in real time. OpenAI shipped an instruction-following fix on 14 November 2025; Sam Altman's own note was that "if you tell ChatGPT not to use em-dashes in your custom instructions, it finally does what it's supposed to do". By The Economist's July 2026 corpus study, only Claude used em dashes more than human writers; ChatGPT used markedly fewer than any human writer in their sample. Two credible measurements disagree because they caught different model versions at different moments, and that disagreement is the lesson: lexical tells expire in months, so do not build your editing habit on them.

What has not expired is the structural stuff, and The Economist's numbers are the most useful thing published on it. Across a 55,940-sentence corpus of its own journalism and model-written versions of it, the paper concluded that "AI prose is distinguishable by word and punctuation choice as well as sentence and paragraph structure." Specifically: models use fewer commas and semicolons than humans and almost no parentheses, lean on "and" as a connector, write longer sentences with less variation in length, favour polysyllables and nominalisations, and rarely quote a named expert. Punctuation scarcity, not punctuation excess. If you have been deleting dashes, you have been editing in the wrong direction.

The word list deserves the same scepticism. As The Economist's newsletter put it, "claiming that a text is by an LLM because it uses the word 'delve' is like claiming one is by Jane Austen because it uses 'imprudence'." Dr Karolina Rudnicka of the University of Gdańsk, quoted in the same piece, is blunter: "There is no single AI writing style, just as there is no single human writing style." There is also a fairness problem hiding in the blacklist. The most-cited explanation for "delve" is Alex Hern's reporting that RLHF annotation work was outsourced to workers in Nigeria and Kenya, where the word is ordinary formal English, and Nigerian writers pushed back publicly when the word was held up as bad writing. Banning other people's dialect is a poor proxy for editing.

So: structure first, words last. Here is the order.

The 10-minute pass, in order

Run it top to bottom on a printed timer if you like. The early minutes move the most.

Minutes 0–3: structure

Louis-François Bouchard, who has edited thousands of AI-assisted submissions, frames the division of labour well: "You own the structure; the model fills it." His fastest diagnostic is to reduce every paragraph to a one-line summary, read the summaries as an outline, and ask whether the shape collapses into definition → list → recap. If it does, no amount of sentence polish saves it.

Three deletions do most of the work here.

Delete the scaffolding intro. If your opening paragraph would fit above any article on any topic ("In today's fast-paced world," "In this article we'll explore"), cut it entirely and start at the first specific sentence. You will almost never miss it.

Delete the recap conclusion. A final section that restates your headings tells the reader nothing they did not just read. End where the argument ends.

Delete the "Challenges" block. Wikipedia's editor-maintained catalogue of AI writing signs names the formula precisely: "Despite its… faces several challenges…" followed by speculation, often under a heading like "Future Outlook." Either cut it or replace it with one named, sourced objection — which is a better section anyway.

Fix the formatting leakage the same catalogue flags while you are at this level: Title Case headings, boldface scattered across sentences nobody would skim, two-item vertical lists that were never parallel. Sentence-case the headings, unbold anything that is not a real scanning anchor, turn the short lists back into prose.

Minutes 3–6: rhythm and punctuation

Now read a page aloud in your head. You are listening for sameness.

Take each of your first three paragraphs and split its longest sentence in two. Then let one sentence in each paragraph run under eight words. That single operation attacks the length-uniformity signal The Economist measured, and it costs about ninety seconds.

Next, hunt the "and" chains. Where you find a sentence stitched together with two or three "ands", replace one with a comma, a semicolon, a colon, or a parenthetical aside. This is the counter-intuitive edit: you are adding punctuation variety, not stripping it, because the measured tell is scarcity. Parentheses in particular are near-absent from model prose and cost nothing to introduce.

Only now, count your em dashes. A rough count against the roughly 3.23-per-1,000-words human figure is enough; if a 3,000-word post has thirty of them, thin them toward a dozen and vary what replaces them. If it has none at all in 3,000 words, that is its own kind of flat.

Minutes 6–8: the three sentence tics

These three constructions account for a surprising share of the "something's off" feeling, and each takes seconds to fix.

Tic Example from the wild Mechanical fix
Negative parallelism "not only a meme — it's a celebration of grassroots car culture" Pick a side and assert it: "It's a celebration of grassroots car culture."
Rule of three "Construction and Renovation, Electrical and Plumbing, Hobby and Craft" Keep the strongest item. Two is fine; one is often better.
Trailing "-ing" clause "…creating a lively community within its borders, further enhancing its significance" Delete the clause. If the point matters, give it a subject and a full stop.

Every example above is quoted verbatim from Wikipedia's signs-of-AI-writing catalogue, which is the most complete tell list available anywhere and free to read. The Economist's study independently identified "not X but Y," "not only but also," and the rule of three as LLM favourites, which is about as much corroboration as this field offers.

Two smaller sweeps belong in the same two minutes. Restore the copula: "Gallery 825 serves as LAAA's exhibition space" becomes "Gallery 825 is LAAA's exhibition space." And strike undue significance — "marking a pivotal moment," "represented a significant shift," "enduring legacy" — replacing each claim of importance with the fact that would justify it, or with nothing.

Then run one Ctrl-F for leftover assistant chatter: "Certainly, here is," "I hope this helps," "As of my last knowledge update." It is rare in a supervised workflow and catastrophic when missed. Leftover chatter is one of the fingerprints behind English Wikipedia's speedy-deletion criterion for unreviewed LLM pages (WP:G15), which tells you how unambiguous a signal it is to a reader.

Minutes 8–10: evidence density

This is the highest-leverage part of the pass, and the part most editing guides skip because it takes actual knowledge.

Find every vague attribution — "industry reports suggest," "observers have cited," "experts argue" — and either name and link the source or cut the sentence. Vague attribution is simultaneously an AI tell and an uncited claim, so fixing it buys you two things at once.

Then find one adjective doing a number's job. "Robust," "seamless," "significant," "scalable": each of these is a placeholder where a figure belongs. Bouchard's version of the swap is "robust" → "handles 10k requests/second." You know your numbers; the model never did.

Finally, add one quoted human with a name attached. The absence of expert quotation was one of the distinguishing features The Economist found in AI prose, and it is the single edit a model cannot fake for you; it requires that you know someone, read someone, or have asked someone.

While you are here, un-nominalise. "The implementation of the new pipeline resulted in a reduction of costs" becomes "we shipped the new pipeline and costs fell." The Economist characterised default LLM style as "bland, pretentious prose lavished with Latinate words", and turning nouns back into verbs undoes most of it.

Vocabulary swaps come last, if at all. Delve, tapestry, underscore, pivotal, intricate: change them if they grate, but they buy the least of anything on this list.

If you only have three minutes

Some weeks ten minutes is not available. Do these five, in this order, and you will remove most of what a reader would notice:

  1. Delete the first paragraph if it could open any article.
  2. Delete the recap conclusion.
  3. Name and link one vague attribution.
  4. Replace one adjective with one number.
  5. Split the longest sentence in each of the first three paragraphs.

Three of the five are deletions, which is why the short version works at all. Most de-slopping is subtraction.

The same edits make you more citable

There is a pleasant coincidence here worth knowing about, because it changes how you feel about spending the ten minutes.

The GEO study, which tested optimisation methods against a 10,000-query benchmark, found that the highest-performing changes to a page were Quotation Addition (+27.8% visibility), Statistics Addition (+25.9%), Fluency Optimization (+25.1%) and Cite Sources (+24.9%), with top methods reaching a 30–40% relative improvement. Keyword stuffing came last among the tested methods, at +17.8%.

Read that list against minutes 8–10 above and it is the same list. Add a quotation, add a statistic, name and link your sources: those are the de-slop fixes and the citation fixes. The honest caveat is that GEO measures visibility in generated answers rather than reader trust; it is a proxy for the value of the pass, not proof of it, and nothing published tests whether a voice edit moves time on page or conversion directly. But when the reader-facing edit and the machine-facing edit point the same way, the decision gets easy. If you want the mechanics of that side specifically, our guide to what actually earns citations in ChatGPT and AI Overviews goes paragraph by paragraph.

What this pass does not do

Three boundaries, so you do not over-trust ten minutes of work.

It is not the fact-check. Voice and accuracy are separate passes with separate failure modes, and the accuracy one matters more. A beautifully rewritten sentence citing a URL that 404s is worse than a clunky one citing a live source. Run the facts pass on its own terms; our 20-minute link-verification checklist for AI drafts is the companion to this piece.

It is not detector evasion, and you should not treat it as one. Reported detector accuracy is all over the map: Turnitin claims a false-positive rate under 1%, while a small Washington Post test found roughly 50%, and Turnitin acknowledges missing roughly 15% of the AI-generated text in a document. Worse, GPT detectors "consistently misclassify non-native English writing samples as AI-generated" while scoring native samples correctly. Optimising for a scorer that unreliable is a bad use of your morning. The scores are increasingly in public view, though: Substack shipped a Pangram-powered AI detector in July 2026, to a mixed reception from writers. Edit for readers; just know the scores exist.

It cannot add expertise you did not put in. Google's helpful-content self-assessment asks whether content provides "original information, reporting, research, or analysis", whether it demonstrates first-hand expertise, and whether you are "mainly summarizing what others have to say without adding much value." No line edit answers those. They need material — your numbers, your customer conversations, your failed experiment from March.

One last caution: resist over-correction. Some novelists and essayists are now deliberately inverting AI defaults, leaving typos in and swapping em dashes for parentheses as winks to the reader. It is an interesting cultural moment and terrible advice for a business blog. Patterns chosen as anti-AI signals become AI-imitable within months; sloppiness stays sloppy.

Fix it upstream so the pass gets shorter

Everything above is remedial. The cheaper move is arriving at fewer of these edits in the first place.

Start with the draft-time instruction. That same em-dash study found that adding "write in flowing prose paragraphs only, no markdown formatting, headers, bullet points, bold text or lists" dropped Claude Opus 4.6 from 9.09 to 0.19 em dashes per 1,000 words and Gemini 2.5 Pro from 3.53 to 0.00. GPT-4.1 barely moved, from 10.62 to 9.10. The paper's proposed mechanism is that the em dash is markdown leaking into prose; Sean Goedecke offers a complementary theory, that models learned English partly from digitised late-1800s books, where em dashes were roughly 30% more common than in contemporary writing. Either way, a prose-style instruction is free and works on most models.

Then fix how you specify voice. The State of Brand's critique is the sharpest statement of why adjective lists fail: "A language model given a set of adjectives can only produce the statistical average of every company that ever wanted to sound that way, which is all of them." "Professional yet friendly" is an averaging instruction. What steers a model is mechanical: a maximum sentence length, a banned-construction list ("no 'not just X, but Y'", "no three-item adjective lists"), a rule that every section contains one number, and one first-person anecdote per post.

This is the part Magic Share is built around. The agent scans your site into a cached brand profile (product, audience, tone, keywords), so drafts start from how your existing pages actually read rather than from a template, and your mechanical constraints live in that profile instead of being retyped into a chat window every week. The vague-attribution tell gets handled at the source too: every cited URL is resolved before a draft is saved, and the verification report ships attached to the post, so "industry reports suggest" never reaches your inbox as an unsourced sentence.

None of that replaces your ten minutes. It is what makes ten minutes enough. Drafts wait for a human either way (nothing publishes on its own), and the point of the upstream work is that the approval step stays a coffee rather than an afternoon.

FAQ

How do I make AI content sound human without rewriting the whole thing? Work in the order structure → rhythm → sentences → words, and stop at the ten-minute mark. Deleting a scaffolding intro, deleting a recap conclusion, splitting three long sentences, naming one vague source and adding one number removes most of what a reader notices. Full rewrites are rarely the highest-value use of the time.

Should I delete every em dash? No. Aim for human range rather than zero. One measured sample put human writing at about 3.23 em dashes per 1,000 words, while models ranged from 0.00 to 10.62 depending on which one drafted the page. Zero em dashes in a long post is itself a pattern, and the tell has weakened since OpenAI's November 2025 instruction-following fix.

Do AI content detectors decide whether my post is a problem? They are not a reliable referee. Claimed false-positive rates range from under 1% to about 50% in a small newspaper test, the leading tool acknowledges missing roughly 15% of AI text in a document, and detectors systematically misflag non-native English writers. Readers, not detectors, are the audience that matters, and in the year's biggest cases readers noticed first anyway.

Will Google penalise my post for being AI-drafted? Google's spam policies are explicitly method-agnostic: scaled content abuse is about generating pages primarily to manipulate rankings, "no matter how it's created." The reason to run a voice pass is reader trust, not penalty avoidance. Value per page is the test, which is a different question from who typed it.

What is the difference between the voice pass and the fact-check? The voice pass fixes how the draft reads; the fact-check fixes whether it is true. They fail differently and should be run separately. Verify every cited URL resolves and every claim matches the sentence that supports it before you spend a minute on rhythm.

The short version

Slop is a prose problem before it is an ethics problem. Readers who use LLMs daily spot machine-written text almost perfectly, and they do it on structure and evidence rather than vocabulary. Every one of those signals has a fix that needs no judgement call: delete the scaffolding, vary the sentence lengths, add the punctuation back, name the source, add the number, quote the human. Ten minutes, in that order, gets you most of the way.

The rest is upstream. If you would rather your drafts arrive with the sources already named and resolved and your own site's voice already loaded, see how Magic Share drafts and verifies a post: your first three posts are free, and nothing ever publishes without you.

Want posts like this for your site?

Magic Share researches, writes and fact-checks posts like this for any site — point it at your URL and review your first draft today.

Plant your first post
FeaturesPricingBlogGet startedLog inSign upPrivacyTerms© 2026 magicshare