You do not need customers to publish original data. A pre-revenue product already produces measurable things — request logs, benchmark runs, crawl samples, small reproducible tests against public systems — and any one of them can carry a post that a competitor cannot regenerate from the same ten search results. The pressure to have something like that is new: in Graphite's Q1 2026 sample of 55,400 randomly selected English-language articles from Common Crawl, 49.9% were primarily AI-generated against 50.1% human, a share that has sat near half for five consecutive quarters. The advice in your niche has already been written, twice. The measurement has not.
The rest of this post is about how to act on that before you have a single paying account. Not "do a State of the Industry report" — you have no list. Something much smaller: pick one thing you can count that nobody else is counting, instrument it this week, and publish the number with its methodology attached.
Run every paragraph of your next draft through one question: could a competitor produce this sentence from the same search results you used? If yes, it is borrowed. Borrowed material is not worthless — a reader still needs the explanation — but it cannot differentiate you, because the eleven other people writing that post this month have access to exactly the same inputs, and increasingly to the same model that assembles them.
The consequences of losing that test are visible in the aggregate. Ahrefs studied roughly 14 billion pages in their Content Explorer index and found that 96.55% got zero traffic from Google, with another 1.94% picking up between one and ten monthly visits. Their own caveats are worth repeating because they model the honesty this post is arguing for: the index skews toward the quality side of the web, and the traffic figures are estimates drawn from a keyword database of about 651 million terms, so long-tail visits can go uncounted. Even discounted, the shape is unambiguous. Most published pages are not competing badly. They are not competing at all, because there is nothing on them that isn't somewhere else.
Google now says this out loud. Its May 2026 guidance on optimizing for generative AI features carries a line that reads like it was written for exactly this problem: "Don't just recycle what others on the internet have already said, or could easily be produced by a generative AI model." The same page draws the distinction sharply — a first-hand review "provides a unique perspective based on personal experience, whereas a summary of existing content simply restates information already available elsewhere." Note what is not being said there. Nothing bans AI-assisted drafting; the offence is having nothing new to say, which is also the reading we took in our closer look at what Google's scaled content abuse policy actually bans. Method is not the test. Value per page is.
If you have read much SEO writing in the last year you have seen "information gain" described as a live, weighted Google ranking signal, sometimes as the signal behind a recent core update. Be careful with that framing, because the primary sources do not support it, and repeating it is a fast way to lose a technical reader.
The patent does exist — US11354342B2, filed October 2018, granted June 2022, assigned to Google. It describes applying previously-presented and yet-to-be-presented documents to a trained model to produce an information gain score, so that new information can be surfaced "in a manner that reflects the likely information gain that can be attained by the user." That is a real mechanism, but it is a mechanism for anticipating what a user wants to see next, given what they have already seen. Search Engine Journal's analysis of the same patent points out that it mentions automated assistants far more often than search engines — 69 times against 25 — and that no evidence was found for the claim that information gain drives core updates. Google's own ranking-update history lists the 2026 core and spam updates without attributing any of them to it.
So use the patent as a concept and Google's published guidance as the evidence. The practical instruction is identical either way — be the source of something — but one version survives a reader who checks, and the other doesn't. That distinction is the entire reason this section exists.
"I have no data" is almost always false. It usually means "I have no customer data," which is a much narrower statement. Here are eight sources, each with a published example whose method you can copy at your own scale.
1. Your own product's logs — even from a beta or a free tier. Plausible Analytics, a small bootstrapped SaaS, ran their own tool and Google Analytics side by side on one website during a traffic spike over three days in August 2021: 50,947 visitors, 60,731 pageviews. The result was that 58.67% of visitors never showed up in Google Analytics, with 88.28% of Firefox users and 82.3% of Linux users blocking it. One site. Three days. The post is still being updated five years later, which tells you what a reference number is worth compared with an opinion piece.
2. Your crawler, bot and request logs. When llms.txt was the topic of the month, everyone published the adoption rate. Ahrefs published the question only their logs could answer — does anything ever read the file? Across 137,210 domains in May 2026, 97% of llms.txt files received zero requests, and 96% of the requests that did arrive came from bots. That is the move to steal: not a better answer to the popular question, a different question that your instrumentation happens to own.
3. A public corpus you can sample. Graphite's AI-content study needs no proprietary index — it is Common Crawl plus off-the-shelf detectors, sampled to a stated rule (English, article schema, 100+ words, publication dates in range). Anyone with a weekend and an API budget can run a version of that on a niche they know well.
4. A fixed public list you re-check on a schedule. Rankability checks the Tranco top 1,000 domains monthly for llms.txt over HTTPS with an identified crawler, counting only HTTP 200 plain-text responses; June 2026 adoption came out at 8.7%. The list is public, the check is a script, and the asset is the time series rather than any single month.
5. A small reproducible test against public systems. Netcraft asked a GPT-4.1-family model where to log in to 50 well-known brands, using two plain phrasings. Of the 97 domains returned across 131 hostnames, 34% were not brand-owned — 29% unregistered, parked or inactive, and 5% belonging to unrelated legitimate businesses. Fifty prompts and an afternoon produced a finding that travelled across the security press.
6. A comparative bake-off across competing tools. The Tow Center for Digital Journalism ran eight generative search tools against 20 publishers and 10 articles each — 1,600 queries — checking whether each returned the correct article, publisher and URL. Their summary, that "most of the tools presented inaccurate answers with alarming confidence," is quotable precisely because the procedure behind it is boring and repeatable.
7. A benchmark you run continuously and never stop running. More on this one below, because it is the strongest case in the set.
8. A survey of a niche you actually belong to. State of JS collected 14,015 responses in late 2024 and is run by a two-person collective funded partly by t-shirt sales. Orbit Media has run the same one-page, 24-question survey for twelve years with no incentives offered. Neither has a proprietary dataset. Both own numbers their whole industry quotes.
The pattern across all eight is that none of them require customers. They require a decision about what to count.
Here is the uncomfortable part. Ahrefs' llms.txt study only exists because they were already logging bot requests before the question came up. Plausible's 58% number only exists because someone chose to run two analytics tools in parallel on the same site. Neither study could have been produced retroactively by a team that decided, on the Monday, that this week's post should contain original data.
So the action is not "plan a research report." The action is smaller and duller: add one counter to your product this week. Log the thing you would want a number for in six months — how long a job takes at the 95th percentile, how often an external dependency returns something malformed, how many of a class of inputs fail on first pass — and let it accumulate while you write about other things. Instrumentation is cheap; the six months of history you didn't collect is not.
We are following our own advice here, and it is worth being specific about it since Magic Share is itself a young product. A single agent run produces countable things: sources read, ideas scored and discarded, minutes elapsed, cited URLs resolved or dead on first attempt, per-run model cost across a picker that spans small models to frontier ones. The number we care about most is the last-mile one — the share of URLs a model proposes that turn out not to resolve, broken down by model — because verifying every cited link before a draft is saved is the part of the pipeline that exists specifically to catch it. We have not seen anyone publish that breakdown, and stating the intention here with a date on it is the point: you can check later whether we did.
If you want a template for the ambitious version of this, look at Artificial Analysis. It began in 2023 as a side project in Sydney, built by two people who kept hitting unreliable vendor performance claims while building something else, and launched publicly in January 2024 as a free website with no funding and no customer base. "We built it because we needed it as people building in the space," co-founder George Cameron told Latent Space, "and thought, Oh, other people might find it useful too." By early 2026 the team was just over twenty people, still giving the benchmarks away. Their methodology page is the artefact to imitate: it names what is measured, standardises the units, lists the evaluation sources and makes the prompts downloadable. They also run a "mystery shopper" policy — accounts registered off their own domain — so a provider cannot serve something special to a known benchmarking endpoint. An anti-gaming control, designed in before the first result existed.
The objection you will hear from any technical reader is sample size. It is a fair objection, and the answer is not a bigger sample — you do not have one — but disclosed limits. Four rules cover most of it.
Set a floor and respect it. Pew Research displays margins of sampling error when a subgroup's effective sample size drops below 100, and will not publish an estimate at all when a subgroup has fewer than 100 raw interviews or an effective sample size below 50. For the arithmetic, their election-polling explainer gives usable anchors: a simple random sample of about 1,067 carries roughly ±3 points on a single estimate, while a subgroup of about 160 carries ±8 points — and about ±16 points on the difference between two estimates. That last number kills most small-sample claims: with 160 responses, you cannot honestly say segment A differs from segment B by 10 points.
Keep the unknowns in the denominator. Rankability reports 8.7% adoption (87 of 1,000) rather than 15.8% (87 of the 549 sites that were reachable), specifically so the figure isn't quietly inflated and "every month stays directly comparable on the same fixed base." Dropping unreachable cases is the single easiest way to publish a number that is wrong in your favour without ever intending to.
Publish your error rates, and publish your revisions. Graphite's 2026 study reports false positive rates under 2% across all three detectors and false negative rates in a similar range, tested against current frontier model output — and it states that the three-detector method produced estimates 3.3 percentage points lower than their own October 2025 publication. They left both posts up. A revision published voluntarily reads as competence, not retraction; nothing else you can do buys that much credibility for so little effort.
Run a benchmark checklist before you publish one. Gernot Heiser's Systems Benchmarking Crimes is a free, blunt inventory of the ways performance numbers mislead: benchmark sub-setting without justification, micro-benchmarks passed off as system performance, no indication of significance, missing platform specs, and comparisons against a stale state of the art. On competitor benchmarks it is unambiguous — "using sub-optimal results as a basis for comparison is highly unethical and probably constitutes scientific misconduct." Read that as a hard rule if you are benchmarking anything you also sell against: if you cannot configure a competitor's tool as well as its own docs recommend, you are not ready to publish the comparison.
Finally, ship the artefacts. The USENIX Security 2025 team behind the package-hallucination study — 19.7% of recommended packages hallucinated across roughly 576,000 generated code samples and 16 models — published their generation pipeline, prompt datasets and reproduction scripts while deliberately withholding the full list of 205,474 hallucinated package names, because that list is a target list for attackers. You will not match that scale, but the structure is scale-independent: fixed prompt set, repeated runs, published method, one stated omission.
There is a failure mode worse than publishing nothing new, which is publishing something new and wrong — or laundering someone else's error into a fresh citation. Marketing has a documented habit of it. Fenwick traced several canonical statistics back to nothing at all: the famous "$42 return per $1 spent on email" started life as £42 per £1 in UK research and was never currency-converted, and the "attention span fell from 12 seconds to 8" line is attributed to Microsoft, who credit a source with no record of it.
Live example from this year's research: multiple 2026 listicles state that a joint MIT CSAIL and Oxford Internet Institute study found AI-generated content makes up 64% of new internet material, complete with sub-figures. Searching for the primary source turns up no paper, no press release and no institutional page — the claim appears only on aggregator statistics pages. The sourced alternative is Graphite's 49.9%, with its sample and detectors disclosed. Use that one.
The working rule is a two-click test: if you cannot reach the actual measurement — the study, the table, the methodology — within two clicks of the sentence citing it, don't cite it. And make sure yours is reachable in one. Our own 20-minute checklist for fact-checking an AI draft covers the mechanics, but the short version is that the citation layer, not the opinion layer, is where drafts break. The Netcraft and Tow Center findings above are both evidence for that: models return confident, well-formed URLs that are simply not the right ones.
A study nobody reads is a hobby. Distribution for small original research has changed, so plan it rather than hoping.
The free route most people remember is gone: HARO, later Connectively, shut down on December 9, 2024 after sixteen years, at one point connecting 800,000+ sources with 55,000 journalists. Successors exist — Qwoted cites a database of over 130,000 vetted experts — but the reflexive "blast a pitch and wait" motion is not what it was.
What still works is unglamorous and mostly 1:1. Orbit Media's own guide to creating original research recommends pre-releasing under embargo three to seven days before publication with personalised outreach, partnering with an organisation that already has the audience or list you lack, and — a small detail that matters more than it should — not shipping the findings as a PDF, because a PDF is hard to track and hard to link into.
The compounding argument is stronger than any single launch. Rankability re-checks monthly. Orbit Media has run the same survey for twelve years, across 12,971 total respondents. Graphite re-ran its study with a larger sample and a better method, and got mainstream coverage the second time as well as the first. A series accrues authority that a one-off cannot — but it needs a memory. If whoever drafts your December update cannot see what the June one said, you republish the same post with a new date, which is a loop rather than a series.
Be realistic about the size of the lift, because overselling it here would be exactly the sin this post is complaining about.
In Orbit Media's 2025 survey of 808 content marketers, 49% publish original research and 25% of those report "strong results" — against a 21% baseline across all respondents. That is real but modest, and smaller than several other levers in the same dataset: 39% for those publishing 2,000+ word articles, 50% for those adding seven or more visuals. The data is also correlational and self-reported, recruited from one person's network. So the honest claim is not that original research wins bigger. It is that original research wins uniquely — it is the only asset in that list a competitor cannot regenerate from the same SERP.
Links behave a little better than traffic. Grizzle's comparison of Buffer's State of Social report against a comparable editorial post found link acquisition velocity — links divided by days to acquire — was 293% greater for the research piece, across a table of ten SaaS studies whose referring domains run from 29 to 1,280. One snapshot, tool-reported counts, no control group; directionally useful, not a promise.
The generative-search picture needs the most care. The original GEO paper reports up to 40% visibility improvement from methods including adding statistics, quotations and cited sources. A July 2026 critical survey of 45 GEO studies narrows that considerably, noting that the effect was measured with the source already present in the retrieval context, that body-only optimization has been found to reduce top-20 presence in one arena-style evaluation, and — the detail that matters most here — that interventions adding statistics "are not subject to a strong truthfulness constraint." The studies measured citation likelihood, not whether the statistics were true. Its own summary is that "no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior."
Read those two findings together and you get the sharpest version of the whole argument: sprinkling statistics into a post is a weak and possibly counterproductive tactic, but being the source of the statistic is a different thing entirely, and it is the thing that survives when everyone else is sprinkling. It is also not a shortcut around concentration — in one vendor analysis of over a million AI citations in early 2026, community platforms took 52.5% of citations against 47.5% for brand domains — so pair the data with the page-level work in our guide to getting cited by ChatGPT and AI Overviews.
And there is one commitment nobody has written the small-SaaS version of yet: if you publish your product's own numbers, you should keep publishing them in the quarter they turn against you. Artificial Analysis' independence framing implies as much. Decide that before the first post, not after the first bad month.
How much data do I need before I can publish a study? Less than you think, but be explicit about the floor. Netcraft's widely covered finding rested on 50 brands and two prompt phrasings; Plausible's on one site over three days. What makes small samples publishable is disclosure — state the n, the collection window, the exclusions, and the margin of error where one applies. Pew's floor of 100 raw responses is a sensible limit for survey work; for tests against public systems, the question is reproducibility rather than n.
Can I publish data from my beta users or free tier? Aggregate hard, and never publish a single account's numbers without asking. The safe pattern is metrics about your system rather than about identifiable users — job durations, error rates, link resolution rates, retry counts. If a figure could be traced back to one customer, it needs their explicit sign-off, and if you are in doubt about the privacy status of an aggregate, keep the claim qualitative rather than guessing at the law.
Isn't a benchmark of my own product just marketing? It is, if you only benchmark yourself or configure the alternatives badly. Heiser's benchmarking crimes list is the pre-publish check: use a proper baseline, don't compare against a stale state of the art, publish the platform spec, and never present a sub-optimally configured competitor. The safest first benchmark measures something neutral — an external API, a model behaviour, a format — where you have no result you'd prefer.
How often should a small team publish original data? Rarely, and on a schedule. One well-instrumented recurring measurement re-published monthly or quarterly beats four one-off studies, because the series is the asset and each re-run costs a fraction of the first. Fill the rest of the calendar with ordinary well-sourced posts; the data pieces are what the rest of the blog gets linked from.
What if my measurement turns out boring? Publish it anyway if the method is sound and the question is one people actually ask — "we checked and nothing happened" is a real finding, and the kind that gets cited when someone needs the negative result. What you should not do is torture the sample until it produces a headline.
The realistic first step is not a report. It is choosing one thing to count, adding the counter today, and writing about something else for two months while the log fills up. Then publish the number with its method attached and its limits stated, and re-run it on a schedule so the second version costs you an hour instead of a week.
Magic Share exists for the other part — the steady, well-sourced posts that surround the data pieces, drafted by an agent that researches the niche, holds every draft for your approval and verifies that every cited URL resolves before saving. You can start free with three fact-checked posts, no card required, and keep the counting for the part only you can do.
Magic Share researches, writes and fact-checks posts like this for any site — point it at your URL and review your first draft today.
Plant your first post