An AI content experiment compares content made with AI against content made without it, on the same site or channel, over the same months, with the same metrics. Most "results" posts online skip at least one of those conditions, which is why they disagree. This article gives you a 3-month template to run the test yourself. The numbers in the title are yours: the tables below are a template, left blank on purpose so you fill them with your own results instead of trusting someone else's screenshot.
We wrote it this way because the honest answer to "does AI content work?" is "it depends on how much a human adds, and you have to measure it". Below are the questions people keep asking on Reddit, answered one by one, then the full method.
Why do AI content results online contradict each other?
Because most of them only show one side of the sample. People post the big win or the big crash, rarely the full set of pieces they published. Without the failures and the boring middle, you cannot tell whether AI helped.
A user on r/SEO described this exact gap:
"AI content gets a quick win, then around month 3, or at the next core update, it falls off. I've seen plenty of screenshots of the collapses. What I never see is the other half of the sample"
Source: r/SEO thread
That quote also explains why the template below runs for three months. A test that stops after a few weeks measures the early bump and misses any later drop.
Use it now: before reading any AI results post, check whether it reports every piece published or only the best ones.
Does Google penalise AI-generated content?
Google says it does not penalise content for being AI-generated; it penalises low-value content made at scale to manipulate rankings. The line is about purpose and quality, not the tool.
Google's own guidance on AI-generated content says it rewards helpful, original content however it is produced. Its spam policies name "scaled content abuse": producing many pages mainly to manipulate search rankings, whether by automation, people or both.
The risk is real when the goal becomes volume. One SEO on r/SEO described a manual action after this instruction from their boss:
"My boss strongly insisted that I use AI to generate content and cut the editorial/audit time as much as possible. He even demanded that I publish at least 10-50 articles a week."
Source: r/SEO thread
That is why "edit %" is a column in the template. It is the variable that separates assisted content from scaled content.
Does YouTube flag AI content?
YouTube's monetization rules target content that is mass-produced or repetitive, not AI as such. Channels that add clear human input, such as their own voice, commentary or editing, are on safer ground than channels that publish near-identical AI videos.
The relevant rules sit in the YouTube channel monetization policies. A creator on r/youtubers described what happened when the platform disagreed with their own view of the work:
"the YouTube algorithm completely demonetized/flagged my channel for "Inauthentic Content / Mass-Produced"."
Source: r/youtubers thread
We cover the voice question in detail in Does AI Voiceover Hurt a YouTube Channel?. For the experiment, the takeaway is the same as for search: record how much of each piece was human work, so you can tell which level of AI use is safe for your channel.
How do I set up a fair AI content experiment?
Split your next batch of content into two groups that are as alike as possible, change only how much AI is used, and publish them over the same weeks. Decide the metric and the length before you start, and do not stop early.
Here is the setup:
- Pick the unit. Articles for a blog, videos or shorts for a channel. Do not mix units in one test.
- Pick topics of similar demand. Use topics with similar search volume, or for video, ideas with similar outlier scores in your niche. Otherwise topic strength hides the AI effect.
- Split into two groups. Group A is AI-assisted (AI drafts, a human edits). Group B is human-only. Aim for at least 10 pieces per group; fewer than that and one lucky piece decides the result.
- Alternate publishing. Publish A, B, A, B over the same weeks, so seasonality hits both groups equally.
- Log edit %. For each AI-assisted piece, record roughly what share of the words a human changed. A diff tool or a word count before and after is enough.
- Fix the metric now. For a blog: clicks from search at the end of each month. For a channel: views at 30 days, or the video's outlier score against its own channel.
- Run three months. Then compare month by month, not just the final total.
A skill that checks pages for search and AI-answer readiness helps keep the two groups at the same quality bar before publishing.
Use it now: AI SEO skill on AgentAlley
How do I measure both sides?
Measure both groups with the same metric at the same checkpoints, month 1, month 2 and month 3, and report every piece, including the ones that got nothing. Plot both groups on one chart so a late drop is visible.
For video, raw views are a weak comparison because they depend on how big your channel was that week. A better metric is the outlier score: the video's views divided by the channel's median views. We explain it in Outlier score là gì?. Track the score for every test video, then compare it with the outliers other channels in your niche are getting, so you know whether a "win" is big or only big for you.
Use it now: Track your test videos against niche outliers in govrl
What should the results table look like?
One row per group, one column per checkpoint, plus the number of pieces and the average edit %. Add a "difference" row so the comparison is explicit. This is a TEMPLATE: every cell is blank until you fill it with your own numbers.
Results template (TEMPLATE, fill in your own numbers, one line per group):
- AI-assisted: pieces ___ · avg edit % ___ · month 1 ___ · month 2 ___ · month 3 ___ · avg outlier score or clicks ___
- Human-only: pieces ___ · avg edit % ___ · month 1 ___ · month 2 ___ · month 3 ___ · avg outlier score or clicks ___
- Difference (A minus B): month 1 ___ · month 2 ___ · month 3 ___ · avg ___
Add a second table that lists every piece with its result, sorted from worst to best. That is the "other half of the sample" the Reddit thread said nobody shows. If you publish your results, publish this table too.
What skills matter most when working with AI on content?
Three skills decide the outcome more than the model you pick: checking evidence yourself, editing hard, and measuring honestly. The threads above all point to the same pattern: AI helps when a human stays in charge of judgement.
A creator on r/NewTubers described the trap of letting AI read your results for you:
"Every time you get a completely misaligned read from an AI, it distorts what's actually happening on your channel."
Source: r/NewTubers thread
Use AI to collect and draft, then check the numbers yourself. If you write for search, two more skills help with structure and citations:
- SEO GEO: audit one page at a time for search and generative engines.
- AI Search SEO Citation Strategies: structure pages so AI assistants quote them.
What do I do with the results?
Keep the level of AI use that matched or beat human-only content after three months, and drop the level that faded. If AI-assisted pieces with high edit % held up and low edit % pieces dropped, your rule is simple: keep editing. Then run the test again next quarter, because models, platforms and policies all change.
Whatever you find, share it with the full table. The internet has plenty of screenshots of wins and crashes. It has very few complete samples.
Related reading
- Outlier score là gì? (Vietnamese)
- Does AI Voiceover Hurt a YouTube Channel?
- How to Find Outlier Videos in Your Niche
Sources
- Google Search Central, Google Search's guidance about AI-generated content (2023)
- Google Search Central, Spam policies for Google web search (scaled content abuse)
- YouTube Help, YouTube channel monetization policies
- Reddit threads quoted above, linked inline (r/SEO, r/youtubers, r/NewTubers)
