Methodology

How we score a product

Katapic gives every product two scores out of 100, on two axes that are kept apart because they fail for different reasons. This page lists all fifteen signals, what each one measures and exactly how much it weighs. We publish it because a score you cannot audit is a number, not a diagnosis.

Search visibility, 8 signals

Whether a classic search engine can tell what the product is and match it to a query.

SignalWeightWhat it measures
Readability10%Sentence and word complexity, on a scale calibrated for the language of the listing: eight languages use a published index, Japanese and Chinese a sentence-length proxy we declare as such.
Keyword density15%How often the target keyword appears. Full marks between 1% and 2.5%; stuffing is penalised.
Keyword in the opening15%Whether the target keyword appears in the first 200 characters.
Length vs platform norm15%Description length against what works on the platform the product is sold on.
Vocabulary uniqueness10%Distinct words over total words. Catches copy-paste boilerplate across a catalogue.
Opening length10%Whether the first block lands in the 120 to 180 character window a meta description uses.
Structured hints in the text15%Presence of brand, price, dimensions, model and category cues a parser can pick up.
Sentence variety10%Variation in sentence length. Flat, identical rhythm reads as generated.

AI search visibility, 7 signals

Whether an assistant can lift a usable, attributable statement out of the listing. The rubric follows the peer-reviewed GEO study below, adapted to product pages.

SignalWeightWhat it measures
Citation density15%Links and attribution phrases, recognised in all ten supported languages.
Quantitative claims16%Concrete numbers: sizes, weights, capacities, quantities. Assistants quote specifics.
Factual clarity15%Share of short declarative sentences, at most 22 words. Long clauses are hard to lift.
Question and answer structure14%Questions, FAQ-shaped phrasing and lists an assistant can map onto a user question.
Authority signals16%Brand mentions, years, certifications: what makes a claim traceable to someone.
Markup depth14%Richness of the HTML structure: headings, lists, tables rather than one flat block.
Source attribution10%Whether key claims are attributed to a real source, counted in absolute terms.

Each axis is a weighted sum of its own signals. The weights on each axis add up to exactly 1.00, so a score is always out of 100 and no signal can quietly dominate.

How we score a category page

A category page answers a different question than a product does. Not "what is this" but "which of these do I pick", so it is scored with its own rubric. Two of the fifteen product signals are missing here on purpose: they look for a price, a measurement and an article code, and a category has none of those and should not. In their place are two signals that would make no sense on a product.

Search visibility, 8 signals

The same text signals as a product, with one difference: the length band is shorter, because under a category description sits the product grid, and that is what the visitor came for.

SignalWeightWhat it measures
Readability10%Sentence and word complexity, on a scale calibrated for the language of the listing: eight languages use a published index, Japanese and Chinese a sentence-length proxy we declare as such.
Keyword density15%How often the target keyword appears. Full marks between 1% and 2.5%; stuffing is penalised.
Keyword in the opening15%Whether the target keyword appears in the first 200 characters.
Length vs platform norm15%Description length against what works on the platform the product is sold on.
Vocabulary uniqueness15%Distinct words over total words. Catches copy-paste boilerplate across a catalogue.
Opening length10%Whether the first block lands in the 120 to 180 character window a meta description uses.
Range coverage10%Whether the text names what is actually inside: the types, formats, sizes or materials present, rather than a paragraph that would fit any category in the shop.
Sentence variety10%Variation in sentence length. Flat, identical rhythm reads as generated.

AI search visibility, 6 signals

Whether an assistant can lift a usable answer to "which one should I get". Sources weigh less here than on a product: a category page states almost nothing that needs backing, and rewarding citations would only push toward inventing one.

SignalWeightWhat it measures
Selection criteria30%Whether the text explains what to choose on: use, material, size, budget. It is the heaviest signal of the axis, because it is the whole job of a category page.
Question and answer structure20%Questions, FAQ-shaped phrasing and lists an assistant can map onto a user question.
Factual clarity20%Share of short declarative sentences, at most 22 words. Long clauses are hard to lift.
Quantitative ranges15%Ranges rather than single figures ("36 to 42", "3 to 12 litres"): they describe the set, where a product describes one item.
Source attribution10%Whether key claims are attributed to a real source, counted in absolute terms.
Citation density5%Links and attribution phrases, recognised in all ten supported languages.

The AI-search rubric is adapted from this peer-reviewed study: GEO: Generative Engine Optimization, KDD 2024

From score to grade

The letter beside each score is a plain band, not a curve. It does not move with how other stores are doing.

A

85+

B+

75 – 84

B

70 – 74

C+

60 – 69

C

55 – 59

D

40 – 54

F

0 – 39

The third check: structured fields

Text is not everything. Google AI Mode groups listings into product clusters keyed on the barcode, and the OpenAI product feed specification asks for category, brand, material, dimensions and condition. A listing can score well on all fifteen text signals and still be unmatchable, because the machine reads the fields, not the paragraph. Katapic checks those separately, validating barcodes with the GS1 check digit and categories against the published Google taxonomy.

What the score never does

  • It never generates an identifier. A made up barcode drops a product out of the cluster exactly like a missing one, so inventing them actively harms you. We detect, validate and report; we never synthesise.
  • It never rewards invented facts. A product legitimately without a barcode is not a defect, it is a different finding.
  • It never compares you to other stores. The bands are fixed, so your score means the same thing in January and in June.

See these signals on your own products

A score on both axes for every product in your WooCommerce catalogue. No signup, no card.

Scan free