The question is asked backwards, which is why the answers are so unsatisfying.
Google does not penalize content for being AI-generated. Its position, stated plainly in its guidance, is that it rewards high-quality content however it is produced, and that using automation to generate content primarily to manipulate rankings is a spam policy violation. Both halves matter. The tool is not the problem; the intent and the result are.
That distinction is unusually easy to test on ecommerce category pages, because the failure is so visible. A store with four hundred categories generates four hundred descriptions from the category name alone and publishes four hundred pages of fluent, interchangeable text. Nothing is penalized in the manual-action sense. The pages simply never rank, because there is nothing in them worth ranking.
This is the practical version of the question: how do you generate category copy at scale and end up with pages that hold?
What Google's position actually says
Two documents carry the substance.
Google's guidance on AI-generated content states that its focus is on the quality of content rather than how it is produced, and that using automation — including generative AI — to produce content with the primary purpose of manipulating rankings is a violation of its spam policies.
Google's spam policies define scaled content abuse as generating many pages primarily to manipulate rankings rather than to help people, regardless of whether the process is automated, human, or a combination.
Read together, the standard is about purpose and outcome, not method. The awkward part for ecommerce is that "generate a description for every category" is a scaling exercise by definition — so whether it counts as abuse depends entirely on whether the output helps anyone.

Why the thin-category failure happens
The mechanism is the same one that makes supplier product copy invisible, one step removed.
When the only input is the category name, the model has nothing specific to say. It produces something grammatical and generic — "our collection of running shoes offers quality options for every runner" — and because every category gets the same treatment, the pages become interchangeable. A search engine choosing between four hundred pages that say nothing distinctive has no reason to prefer any of them.
There is no penalty event. There is a filtering outcome, which looks identical from the outside and is harder to diagnose because nothing in Search Console flags it.
The inputs that make generated copy hold
The fix is not writing everything by hand. It is giving the model something real per category.
The products actually in the collection. Which sub-types, which price band, which brands. A model that can see the set can describe it; one working from the name is guessing.
How the sub-types differ. The single most useful sentence a category page can contain is the one that tells a shopper which sub-type they want. That requires knowing the difference — road-to-trail versus aggressive lug, merino versus synthetic.
The questions buyers ask about this category. Your support inbox and your Search Console query list both hold them. Sizing that runs small across a brand, care requirements, compatibility.
What the category is not for. The limit that saves a return.
Your own filters and structure. "Filter by drop and cushioning below" is specific, useful, and impossible to generate without knowing the store.

A workflow that scales without going thin
1. Segment the catalog by value. The categories carrying revenue get written or heavily edited by a human. That is usually 10–15% of categories and the majority of sales. 2. Build a structured input per category — product set summary, sub-type differences, top three buyer questions, price band. This is the step that decides quality, and it is mostly a data exercise rather than a writing one. 3. Use a template that forces specifics. A prompt that asks for prose produces prose. A prompt that asks the model to fill named fields — who it is for, how the sub-types differ, sizing note, what it is not for — produces something usable. 4. Generate, then sample. Ten random drafts from a batch of two hundred will tell you whether the batch is publishable. Reading zero of them is how stores end up with four hundred identical pages. 5. Check for interchangeability. Take two similar categories and read their copy side by side. If you could swap them without noticing, the inputs were too thin. This is the whole quality test, and it takes a minute. 6. Publish in batches, not all at once. A hundred at a time lets you measure and adjust. Four hundred at once leaves you unable to attribute anything. 7. Never publish unread claims. If the model invented a certification, a measurement or a material, that is now a claim your store is making.
What to avoid
- Generating from the category name alone. The single cause of the thin-page outcome.
- Publishing at full catalog scale in one release. It removes your ability to learn anything from the result.
- Word-count targets. They produce padding, and padding is what makes a page look automated.
- Using AI to write the pages that matter most. Your top categories are where a human is cheapest relative to the revenue at stake.
- Treating the output as finished. Generated copy is a draft with good grammar.
How to tell whether it worked
Because there is no penalty signal to watch for, you are measuring quality indirectly.
- Query count per page. A category page with real copy starts appearing for attribute phrasings — "trail shoes for wide feet" — that reveal which filtered views deserve their own page. Generic copy picks up nothing new.
- Impressions on the category term, tracked per category rather than as a site average.
- A control group. Leave a matched set of categories untouched. Seasonality moves everything; the control is what separates your work from the season.
- Six to eight weeks before judging. Reindexing takes time.
- Time on page and scroll depth on the categories you rewrote. If nobody reads the copy, it is not helping anyone regardless of how it was produced.
The checklist
- Top-revenue categories written or heavily edited by a human
- Structured input built per category: product set, sub-type differences, buyer questions, price band
- Prompt template asks for named fields, not free prose
- No generation from the category name alone
- Ten drafts per batch read before publishing
- Interchangeability test run on two similar categories
- No word-count targets
- No invented certifications, measurements or materials
- Published in batches of ~100, not the whole catalog
- Control group of untouched categories retained
- Baseline impressions and query counts recorded per category
- Results reviewed after six to eight weeks
Sources
- Google Search and AI-Generated Content — Google Search Central Blog
- Spam Policies for Google Web Search — Google Search Central
- Creating Helpful, Reliable, People-First Content — Google Search Central
- Google Search Essentials — Google Search Central
- SEO Best Practices for Ecommerce Sites — Google Search Central
Frequently Asked Questions
Want this run against your store? Book a call with The Reach Bureau.