What a similarity score measures
Feed a patent claim and a product description to a modern model and ask how similar they are, and you will get a confident number. That number is real, and it measures something real: how much vocabulary, structure and concept the two texts share.
What it does not measure is whether the product practises the claim.
Those are not close to the same question. A claim is a conjunctive list of limitations — a legal sentence in which every element must be present in the accused thing. A product description is marketing prose written to sound impressive. High overlap between them tells you the two documents are about the same technical area. That is a fine reason to look closer. It is not a finding.
Why this particular error is so easy to make
Because the output looks like a finding.
A screening tool that returns "87" against a claim produces something that slots neatly into a spreadsheet, sorts, and gets forwarded. Nothing in the number's presentation signals that it is a similarity measure rather than an infringement assessment. Two steps later it is in a deck, and by then nobody remembers it was a screening heuristic.
The specific failure we guard against: a score computed against generic claim elements — a processor, a memory, a network interface — will be high for almost any product in the category, because those elements genuinely are present in almost any product in the category. The score is not wrong. It is answering a question nobody should be asking.
What actually has to be shown
For each limitation of the asserted claim, independently:
1. The element, as written. Not paraphrased — claim language is load-bearing and paraphrase quietly broadens or narrows it. 2. A specific citation in the accused product showing that element present. A datasheet page, a manual section, a configuration guide, an exhibit. 3. An honest verdict per element — present, absent, or uncertain. Uncertain is a real answer and must be available, because the alternative is a system that guesses.
And then the conjunction: if any limitation is absent, the claim is not practised, regardless of how compelling the other elements looked.
That last rule is what makes a claim chart a claim chart rather than a pile of suggestive quotes.
Uncertain has to be a permitted answer
A screening system that can only output present or absent will produce false positives under pressure, because on a genuinely ambiguous element it must pick one.
The useful signal is often "the corpus does not settle this" — which tells you precisely where to spend the next hour, and is far more honest than a confident answer built on a document that does not actually say what was needed.
What we do with scores
We use them for exactly one thing: deciding what to look at next. A score can promote a candidate into an analysis queue. It can never, in this system, write a conclusion, and it never appears as support for one.
Anything that reaches a terminal verdict has been through element-by-element analysis with citations per element — and the tooling is built so that a score physically cannot be the thing that advanced it.
Why we are telling you this
Because it is the question to ask any vendor selling automated infringement analysis: **show me a finding and tell me which parts of it are evidence and which parts are resemblance.**
If they cannot separate those two things in their own output, the number they are selling you is a starting point being priced as a conclusion.