Patents miss a lot of AI activity

The paper by Carolina Castaldi, Fulvio Castellacci, Andrea Fronzetti Colladon, Ludovica Segneri, and Francesco Venturini makes a simple point that is easy to forget in AI measurement work. Patents mostly capture technical invention. They are better suited to firms building new AI methods, models, hardware, or systems than to firms putting AI into a product or service that reaches the market.

That matters because most firms will not invent frontier AI. They will adapt it. A software vendor may wrap a language model into a workflow tool. A company may add AI features to an existing service. Another may offer analytics built around large language models. None of that has to produce a patent.

So patent counts lean toward large firms, specialized firms, high-tech sectors, and more formal R&D. They are still useful, but they tell us about one part of the process: the creation of protectable technical knowledge. They are weaker for the next step, where AI gets packaged into things customers can buy. That is the gap this paper tries to fill.

What an AI trademark is meant to capture

A trademark is filed when a firm protects a brand, product name, or service identity in the market. The authors use that feature as the economic signal. An AI trademark is not meant to prove that the firm invented AI. It is meant to show that the firm is claiming an AI-linked good or service in a public filing.

That distinction is useful. The signal is closer to commercialization than invention. It points to product development, product upgrading, and market-facing use of AI. The key information is not only the mark itself, but the filing's goods and services description. That text tells us what the company says the product or service does.

This gives trademarks a different role from patents. Patents ask, roughly, whether the firm created a new technical solution. AI trademarks ask whether the firm is offering something in the market where AI is part of the product story.

There are limits. A trademark filing does not tell us revenue, adoption, product quality, or whether the AI component is meaningful. Some filings may be defensive. Some may never become important products. But the timing and coverage are attractive. Trademark records are public, frequent, and available across sectors, including services.

How the authors identify AI-linked filings

The method is text based. The authors start from prior work on AI trademark identification, then refine and expand the dictionary used to detect AI in trademark records. They search the goods and services descriptions for AI-related language. A filing is treated as AI-linked when the description indicates that the claimed product or service uses, applies, or is built around AI technologies.

This sounds plain, but the paper's practical contribution is in making the process systematic and repeatable. AI is a messy term. Some filings use explicit language such as artificial intelligence, machine learning, neural networks, or large language models. Others may describe related functions in looser terms. A dictionary approach will never be perfect, but it is transparent. Researchers can inspect it, revise it, and apply it across trademark offices.

The authors also frame AI filings as rare events. That is sensible. Even in a period of heavy AI attention, only a small share of all trademarks will be clearly AI-related. Any firm-level model using these filings has to deal with sparsity. The paper does not pretend that a raw count is enough in every setting.

This is also where the work fits with the patent measurement debate. Trademarks will have their own classification errors, but they add an independent view. The point is not to replace patents. It is to measure a different mechanism.

What the Italian firm evidence adds

The empirical section uses trademark filings linked to Italian firms. The main finding is the one practitioners should care about: AI trademark filers are not simply the same firms as AI patent filers.

They are more numerous. They show up in different sectors. Services have more representation than they do in AI patent data. This is what we would expect if trademarks are catching product and service commercialization rather than frontier invention.

That result is the paper's strongest evidence. If AI trademarks merely replicated AI patents with extra noise, they would be less interesting. Instead, the two datasets appear to identify different groups of firms engaged with AI in different ways. Patents find firms closer to technical development. Trademarks find firms that are putting AI into market offerings.

For investors, that distinction matters, but it should not be oversold. A company does not need to invent a new model architecture for AI to matter to its product strategy. The commercial story may come from bundling AI into an existing product line, repricing software seats, automating part of a service, or entering a niche faster than peers. Trademark data may catch some of that activity before it is obvious in standard financial reporting. That is a research hypothesis, not a result established by this paper.

Where the signal could matter for investors

A tradable signal would not be "AI trademark equals buy." That is too blunt. The more plausible use is as a dated event or slow-moving firm characteristic.

A few examples make sense:

The last split is interesting. A patent-heavy firm may be building core technology. A trademark-heavy firm may be packaging AI into products. The possible financial channel is different. For the first group, value might come through licensing, technical leadership, or future platform control. For the second, the channel could be nearer-term product refresh, pricing, cross-sell, or market share. Those links would need to be tested rather than assumed.

The signal could also help sector analysts avoid a common mistake: treating AI adoption as an ICT-only story. If trademark filings show AI-linked services in insurance, education, health care, logistics, or industrial software, that is evidence of diffusion into business lines that patent data may undercount.

Why we could not backtest it yet

We could not test this signal on our platform data. The reason is straightforward. The core input is trademark filing data, including the goods and services description text, the filing date, and the owner name. We do not have that in the available data catalog.

A faithful backtest also needs a point-in-time mapping from trademark owner to listed company. That is harder than it sounds. Filings can be made by subsidiaries, local entities, acquired firms, holding companies, or product units. A backtest has to know, as of the filing date, which public parent owned the applicant. Otherwise it risks lookahead bias or missed links.

The paper's setting is Italian firms. Our usual test universe is centered on US-listed equities, ETFs, and crypto. Even if the trademark data were available, we would still need clean issuer matching and survivorship-safe security identifiers. The paper also treats patents as a complement, and we do not have the required patent data in-platform either.

That does not weaken the paper's contribution. It just means the research signal is not plug-and-play for our current stack. The next concrete step would be a dated trademark feed with applicant names, full goods and services text, and a maintained applicant-to-public-parent link. Without that, we would be testing a proxy for the paper rather than the paper's actual mechanism.