If your AI drafts read like every other SaaS blog, the problem is almost certainly your brief, not your model. A 2025 training-free evaluation across five models found that few-shot prompting — showing the model real examples of the target writing — delivered up to 23.5× higher style-matching accuracy than zero-shot instructions, and that prompting strategy influenced style fidelity more than model size did. Adjectives are instructions. Examples are specifications. Brand voice guidelines for AI writing only work when they contain banned phrasings with replacements, measurable structural targets, worked before/after rewrites, and a vocabulary list pulled from your own published pages.
That single finding reframes the whole exercise. You are not describing a personality to a model. You are writing a spec, and specs contain lists, numbers and examples.
Graphite randomly sampled 55,400 URLs from Common Crawl with publish dates between January 2020 and March 2026 and classified each one by averaging three AI detectors. The share of AI-generated articles has been hovering around half for over a year: 49.6% in Q1 2025, 50.9% in Q4 2025, 49.9% in Q1 2026.
The sting is in the same report: those articles "largely do not appear in Google and ChatGPT." So the risk isn't that AI writing crowds you out of the index — it's that enormous volumes of it are produced and quietly go nowhere, and every generic paragraph you publish joins that pile. For a founder shipping four posts a month, that's the difference between an archive that compounds and one that accumulates.
You know the shape of the failure. You typed "friendly, professional, approachable, no jargon" into a brand voice field, and what came back was competent, tidy, and interchangeable with four hundred other blogs in your category. Nothing in it was wrong. Nothing in it was yours.
The obvious question is why those adjectives did so little.
Anthropic's engineering team put the mechanism plainly in its September 2025 guide to context engineering. A good prompt is "the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome", and there are two ways to miss: hardcoding "complex, brittle logic," or writing "vague, high-level guidance that fails to give the LLM concrete signals for desired outputs."
Read that second failure mode again. It is a precise description of a four-adjective voice brief. "Friendly" is not a signal; it's an index pointing at a signal you never supplied. The model resolves it the only way it can — toward the statistical centre of everything labelled friendly in its training data, which is where every other SaaS blog is already standing.
Now the honest caveat, because this audience checks. You may have seen the widely-shared paper showing that personas in system prompts don't help: 162 roles across 4 model families and 2,410 factual questions, with no improvement over using no persona at all. But it measured factual accuracy, not style. It does not prove that "friendly and professional" fails to shift tone.
Tone descriptors do move something measurable. Nielsen Norman Group tested 100 US adults on four matched content pairs and found the casual version rated 0.7 points friendlier on a 5-point scale than its formal twin — though the formal insurance sample scored 0.3 points more trustworthy than the playful one. Adjectives aren't useless. They're just underdetermined, and the four every founder reaches for carry the least information of any.
So keep three to five adjectives. Just make each one cash out in something concrete later in the brief.
Here is the study that should change how you spend the next hour. Chen et al. evaluated style imitation across five models, no training involved, and found that few-shot prompting yields up to 23.5× higher style-matching accuracy than zero-shot, that completion-style prompting reaches 99.9% agreement with the original author's style, and that prompting strategy influences style fidelity more than model size.
That last clause is the money line for a time-poor founder. Moving up a model tier costs real money on every run and moves style less than adding three worked examples to your brief — which costs twenty minutes, once.
Why do examples work when instructions don't? Min et al. gave the mechanism in 2022: randomly replacing the labels in demonstrations barely hurt performance, meaning what demonstrations transmit isn't correctness but the label space, the distribution of the input text and the overall format of the sequence. Your before/after pairs don't teach the model what a good sentence means. They teach it the shape of an acceptable sentence in your publication — how long, how hedged, whether you say "we" or "the agent."
Now the caveat that makes the rest of this post trustworthy. In that same study, human essays averaged a perplexity of 29.5 while matched LLM outputs averaged only 15.2 — even a near-perfect style match is flatter and more predictable than human prose. And an EMNLP 2025 Findings paper covering 40,000+ generations per model from 400+ authors found that models approximate styles in structured formats like news and email but "struggle with nuanced, informal writing in blogs and forums".
Which is precisely the register you're trying to recover. So set the expectation now: a good brief removes the generic layer, but it cannot manufacture surprise. The one anecdote, the one opinion, the one number nobody else has — that stays yours.
With that framed, here's what actually goes in the file. Four parts, in priority order.
The best public template for this isn't a marketing style guide. It's the GOV.UK A to Z style guide, whose words-to-avoid section never just forbids a word. Every entry names the replacement:
| Banned | Use instead (GOV.UK) |
|---|---|
| agenda | "plan" |
| deliver | "make", "create", "provide" or something more specific |
| leverage | "influence" or "use" |
| utilise | "use" |
| streamline | "simplify" or "remove unnecessary administration" |
| tackle | "stop", "solve" or "deal with" |
| empower | "allow" or "give permission" |
| liaise | "work with" |
| initiate | "start" or "begin" |
| facilitate | say something specific about how you're helping |
That second column is why the list is executable by a machine. A bare prohibition leaves a hole where a word used to be; a substitution tells the model where to go instead. You'll also see this framed as the "pink elephant" argument — that telling a model not to do something forces it to represent the thing. Treat that claim gently: the most-cited write-up admits its evidence is anecdotal Reddit reports rather than controlled experiments. You don't need the theory anyway. GOV.UK's format is defensible on its own terms.
Ban constructions, not just words. Wikipedia's editor-facing signs of AI writing is the readiest free catalogue of the patterns that make prose smell synthetic: copula avoidance ("serves as," "stands as" where is would do), negative parallelisms ("not just X, but also Y"), rule-of-three triplets, puffery ("stands as a testament," "plays a crucial role"), Title Case headings, excessive boldface, and vague attribution ("Industry reports," "Experts argue"). Ban those last two even if you never notice them; a reader who has seen a hundred AI posts does.
For raw material, the excess-vocabulary study is a gift. Analysing 15M+ PubMed abstracts from 2010 to 2024, the authors estimated that at least 13.5% of 2024 abstracts showed LLM-assisted writing, rising to 40% in some subcorpora, and released 900 excess words as an annotated public CSV. It's biomedical-skewed, so mine it rather than adopt it.
Two warnings before you paste 900 words into a file.
First, the list rotates fast. Before 2024, 79.2% of excess words were nouns; during 2024, 66% were verbs and 14% adjectives — showcasing, pivotal, grappling. A list frozen in 2024 is fighting the last war, so set a recurring note to re-read your last ten drafts and add whatever made you wince.
Second, don't ban punctuation as a proxy for style. Em dash frequency in English ecology abstracts more than doubled between 2021 and 2025, the largest change of any character measured across 10,000 abstracts — which is why everyone declared it the tell. But per 1,000 words, Huckleberry Finn runs 10.13 em dashes and GPT-4.1 runs 10.62. A single mark cannot separate Twain from a frontier model. Ban the em dash only if you genuinely don't write with them — and if you do, it now sticks: two days after GPT-5.1 shipped in November 2025, OpenAI's CEO confirmed ChatGPT finally obeys a custom instruction to avoid them.
Rhythm is half of what makes a page sound like a person, and it's the half that's trivially checkable. Three numbers cover most of it:
Add heading case (sentence case, not Title Case), a bold budget and a paragraph-length cap, and your structural section is done in ten lines.
One expectation to set, because it will otherwise cost you an afternoon: the model will not hit these numbers on command. LIFEBench evaluated 26 models across 10,800 instances with length constraints from 16 to 8,192 words and found outputs that were far too short, terminated prematurely, or refused outright, with nearly all failing to reach their own vendors' claimed maximum lengths.
That doesn't make length targets worthless. It makes them the wrong kind of instruction to trust and the right kind to verify — review targets and lint rules, not magic words. The enforcement section below is where they get teeth.
This is the highest-leverage section in the document, the one most briefs skip because it takes actual work, and where the 23.5× lives.
The method is small. Open your last AI draft, find the sentences that made you wince, rewrite each the way you'd have written it, and paste both versions into the brief, labelled:
Before: "Our platform serves as a comprehensive solution that empowers teams to streamline their content operations." After: "It writes the draft. You approve it or you don't."
Three rules make these pairs work harder:
Isolate one behaviour per pair. If your "after" fixes puffery, shortens the sentence, swaps passive for active and drops a buzzword at once, the model must guess which change mattered. One pair for copula avoidance, one for over-hedging, one for the rule of three.
Draw them from published pages, not from your About page. Your About page is written in a register you use nowhere else. The posts people actually read are the spec.
Treat them as regression tests. When a draft comes back with the old behaviour, that pair either wasn't specific enough or got buried too far down the file. Fix the pair, not just the draft.
Placement matters too. OpenAI's prompt-engineering guidance defines few-shot learning as steering a model "by including a handful of input/output examples in the prompt" and recommends a skeleton of Identity, Instructions, Examples, then Context — a free improvement if you're pasting a brief into a system prompt. And if you never touch an API, Claude's custom styles, introduced in November 2024, take an uploaded writing sample and persist the style across conversations.
Adjectives describe your voice from the outside. A vocabulary list is your voice, in the only form a model can copy.
Start by picking samples, and be disciplined about which. HubSpot's brand voice feature asks for a minimum of two to three writing samples, at least 500 words each for blogs and pages, capped at 10,000 words total, and extracts patterns in sentence structure, pronoun usage, tone and formatting. If the platform with the most data landed on "give me two or three real posts," your hand-written brief needs no more. Pick your three most you posts, not your three best-performing ones.
Then split what you pull out into two lists, following the VOICE.md spec — an alpha, Apache 2.0 attempt at a machine-readable brand voice whose field names make a ready-made skeleton. lexicon.protected_terms[] holds canonical names with their forbidden variations; lexicon.forbidden[] holds "phrases diluting voice with documented reasons," each entry pairing a phrase with a reason. That is GOV.UK's ban-plus-replacement pattern in YAML, billed as "a single, versionable, lint-able source of truth."
Protected terms are the boring list that saves the most embarrassment: product name casing, feature names, the words you refuse to let a model pluralise. Ours pin Magic Share and magicshare, the agent (always singular, always personified), brand profile, topic ledger, auto-draft.
The more interesting list is the metaphor you own. Ours is a garden — greenhouse, garden gate, still planting, tended year-round. Two sentences in a brief ("we extend a garden metaphor; drafts grow in a greenhouse, not in production") do more than any adjective, because one metaphor constrains hundreds of downstream word choices at once.
Finally, define your boundaries the way Mailchimp does: each trait paired with what it is not. "We're weird but not inappropriate, smart but not snobbish" tells a model far more than "witty and intelligent," because it marks the edge. Mailchimp also ships a tie-breaker — "Mailchimp's tone is usually informal, but it's always more important to be clear than entertaining." Yours needs one, because the model will hit that conflict and resolve it silently if you don't.
This step is also what an automated tool does on your behalf. When Magic Share scans a site into a cached brand profile covering product, audience, tone and keywords, it is inferring this vocabulary from your published pages. Writing the list explicitly doesn't compete with that inference — it corrects it, the way a per-site banned-topics list corrects topic selection. A banned-phrasings list belongs beside it.
You now have four parts. Three details decide whether they survive contact with a real drafting run.
Order the brief by priority, and keep it short. IFScale tested 20 models from seven providers on up to 500 keyword-inclusion instructions and found that even the best frontier models only reach 68% accuracy at the max density of 500 instructions, with a measurable bias toward instructions appearing earlier in the prompt. So a 300-rule style bible isn't merely tedious to maintain — it drops rules, and you won't know which. Ban list and rewrite pairs go at the top; nice-to-haves go last and get cut first.
Store it where the agent loads it, not where you'll forget it. For this audience that's usually the repo: a versioned VOICE.md beside the content, reviewed in the same pull request as the post. Anthropic's Agent Skills, announced in October 2025, formalise the same idea — folders of "instructions, scripts, and resources that Claude can load when needed," with brand guidelines named explicitly as a use case. A brief plus a folder of example posts is the shape of a Skill.
Enforce after generation, not only before. This is what separates a brief that works from a brief that decorates a Notion page. Vale is a prose linter driven by YAML rules with extension points including existence (flag a regex), substitution (map a banned term to its replacement), occurrence (cap how often a pattern appears) and repetition; its own docs ship an example banning utilize, leverage and synergy. Every row of your GOV.UK-style table is a substitution rule. The 30-word cap is an existence rule.
Then close the loop. Keith J. Jones at Corelight built a tool that runs Vale over a document, packages the violations as JSON into an LLM prompt alongside the relevant style-guide definitions, and iterates. Run against Microsoft's own SECURITY.md, it went 26 violations → 3 → 1 → 0 across four passes, with fixes as small as "Please do not report" → "Please don't report." Generate, lint, feed the violations back, repeat until zero. CI already fails your build when a meta description runs long or a cited URL doesn't resolve, so why is voice the one thing nobody checks? If you already validate frontmatter and resolve every link before a draft is saved, a prose linter is an afternoon's work.
If prompting genuinely plateaus — and for most people it won't — the escalation path is supervised fine-tuning. OpenAI lists "generating content in a specific format" among its best uses, sets the dataset minimum at 10 examples, and notes improvements from fine-tuning on 50–100 examples. With 40 published posts that's within reach, but it's the last rung, not the first.
Two honest limits, because you'd find them anyway.
Voice sits on top of substance; it never substitutes for it. Google's helpful-content self-assessment asks whether content "provides substantial value when compared to other pages in search results" and whether it offers original information, reporting, research or analysis. Graphite's finding that mostly-AI articles rarely surface points the same way. A perfectly-voiced post with nothing new in it is still a post with nothing new in it — the same line Google's scaled content abuse policy draws, method-agnostically.
AI assistance flattens variety across authors, and a brief only partly counteracts it. Doshi and Hauser's Science Advances study of 300 writers and 600 evaluators found that maximum AI access raised individual novelty scores by 8.1% and usefulness by 9% — but that stories using even one AI idea were 10.7% more similar to each other. Padmakumar and He measured the same convergence in essays written with InstructGPT: a drop in lexical and content diversity traceable to the model's contributions, not the humans'. A Cornell CHI '25 study of 118 participants found it runs cultural too — the first suggested food was always pizza or sushi, the first festival always Christmas.
A voice brief pushes back on the surface layer of that convergence, not the idea layer, which still needs you picking the angle — a number from your own dashboard, a decision you regret, a customer sentence you can quote. Make those paragraphs liftable, with the claim, the number and its scope in one place, and they travel: it's how a small blog earns citations in AI answers.
How long should a brand voice brief for AI be? One to two pages: long enough for a ban list, three to six rewrite pairs, structural numbers and a vocabulary list; short enough that the rules you care about sit near the top. IFScale found adherence degrades with instruction density and favours earlier instructions, so a 300-rule bible silently drops rules you thought were active.
Do I still need adjectives like "friendly" and "professional"? Keep three to five, but treat each as a label for something you show later. NN/g's testing shows tone choices move perception measurably — casual phrasing rated 0.7 points friendlier on a 5-point scale — so adjectives aren't inert, just underdetermined alone. Give each one a before/after pair proving what it means in your sentences.
How many writing samples does a voice brief need? Two or three real posts is the working floor. HubSpot asks for a minimum of two to three samples, 500+ words each for blog-length content, capped at 10,000 words total. Representative beats plentiful: pick the pages that sound most like you, not the ones that performed best.
Will banning em dashes make my writing sound less like AI? Barely. Per 1,000 words, GPT-4.1 uses 10.62 em dashes and Huckleberry Finn uses 10.13 — one mark can't separate them. Ban patterns instead: puffery, negative parallelisms, rule-of-three triplets, vague attribution.
Can I automate checking that the draft followed the brief?
Yes, and you should. Vale turns each ban into a substitution rule and each structural target into an existence or occurrence rule, runnable in CI on markdown. Corelight's published loop — lint, feed violations back to the model, repeat — took Microsoft's SECURITY.md from 26 violations to zero in four passes.
You don't need the whole document today. Open a file called VOICE.md, paste in ten banned phrasings with their replacements, add three before/after pairs from your last draft, and commit it. That's most of the 23.5× right there, in less time than it took to read this post.
It's worth doing once because it applies to every run afterwards — a brief written this afternoon is still steering drafts in eighteen months, which only pays off if you publish on a cadence. That's the compounding case for keeping a blog tended year-round. Magic Share builds a brand profile from your site and holds every draft for approval, so your brief is something you correct once rather than re-explain every run. The free plan includes three posts, no card required.
Then read the next draft with one question: which sentence would you never have written? That sentence is your next ban-list entry.
Magic Share researches, writes and fact-checks posts like this for any site — point it at your URL and review your first draft today.
Plant your first post