AI

Google’s European Search Dataset Licensing Program: “Anonymized” by the Buyer’s Own Controls

Aerial view of a vast crowd in wine monochrome, individuals indistinguishable in the mass

Google has published the page the Digital Markets Act forced it to build: a license for its European search data, aimed at rival search engines and qualifying AI chatbots. Trade press spotted it August 31. The page calls what recipients receive personal data, then relays the Commission’s view that the recipients’ own controls can achieve the DMA’s anonymization standard: the final guarantee sits on the buyer’s side of the transaction.

What Is the European Search Dataset?

The European Search Dataset Licensing Program covers a dataset that, per Google’s own page, “comprises ranking, query, click, and view data from Google Search in the European Economic Area (EEA).” It covers EEA search behavior only, not global Search. Only providers that meet the European Commission’s eligibility bar for online search engines under Article 2(6) of the DMA, a category the Commission’s measures extend to AI chatbots with search-engine functionality, can apply to license it.

Who Actually Qualifies

The eligibility list on Google’s page is narrow. An applicant has to run a search engine directed at EEA users, sit outside the control of non-EEA state actors, and not answer to anyone under EU sanctions. It also needs either at least two consecutive years operating in the EU or, for younger companies, €50 million-plus in capital investment, plus at least 50,000 monthly average users in the EU over the past year. Google can request supporting documents and responds to an expression of interest within seven calendar days.

None of that describes an SEO tool, an agency, or a researcher. The recipient has to already be a search engine, or a chatbot with search built in, and it still has to clear an independent assurance audit before Sample C or the full dataset.

See also  Search Revolution: ChatGPT Usage Triples as Google Market Share Drops to 66.9%

Personal Data, Anonymized by Whose Hand

The interesting sentence sits in the page’s summary of the Measures: “The Measures require Google to provide personal data to recipients and only then rely on the recipients’ internal data segregation as a guarantee that individuals won’t be re-identified.” The next line hands the anonymization claim to someone else: “In the European Commission’s view, such internal data segregation on the recipients’ side can achieve the anonymization standard required by the DMA.” Google is not asserting the data is anonymous. It is relaying the Commission’s view that the recipients’ own segregation can achieve that standard, and noting where the machinery meant to achieve it sits.

The Commission’s own specification decision describes the stripping that happens before any of this reaches a recipient: usernames and IP addresses removed, queries with rare terms or sensitive content suppressed, and every user folded into a cluster of at least 1,000 people sharing location, device type, and language. 95% of users land in clusters of 25,000 or more. Access to Sample C and the full dataset then depends on an independent audit before delivery and further monitoring reports after, the same instinct that shows up whenever a name alone can’t be trusted and someone has to check the behavior instead of the label.

What Ships, and When

Three samples exist, at three different depths, per Google’s page:

Sample Contents Cost
Sample A 1,000 rows Free
Sample B Synthetic dataset, up to 10 million queries Fee
Sample C 5% sample of the full Search Dataset Fee, audit required first

Fees, the page states, are “determined on FRAND terms, which the Measures limit to the incremental costs of making the data available, together with a specified rate of return” — a rate of return the Commission’s decision ties to Alphabet’s own weighted average cost of capital, with a narrow exception for an additional margin. No figure has been published for any of it.

See also  ChatGPT Citation Shift: Reddit and Wikipedia Dominate as Referral Traffic Drops 52% in September

The Commission’s measures attach conditions that shape what a licensee ends up holding: data arrives no sooner than seven days after the query happened, the sharing arrangement runs up to five years per beneficiary, invalid traffic is excluded, and the environment is ringfenced against re-identification attempts and onward sharing. Use is restricted to improving the licensee’s own search technology; the measures directly forbid training general-purpose AI models, building consumer profiles or non-search products, and copying Google’s results instead of building independent ranking.

A Timeline Written by Someone Else

Google’s two headline dates are not choices; the Commission’s July 16, 2026 decision set the clock. On the Commission’s milestone list, template licence agreements and test data samples sit at the two-month mark, the finalized anonymized dataset at four months, and a pricing offer at around six, in January 2027. Google’s page dates the licensing agreement from September 17 and says the data samples “will be available as of November 16, 2026.”

Ahrefs’ AI Adjusted Volume metric, covered here in late August, already showed what an estimate built on top of Google’s numbers looks like once you open the formula. This is the same input at a different layer: a licensed dataset described by its own publisher as personal data, anonymized by a standard whose final guarantee sits on the recipient’s side of the fence, arriving on a schedule someone else wrote.

Sources: Google, European Search Dataset Licensing Program; European Commission, Alphabet specification proceedings on sharing Google search data.