← All posts

One Year of EU GPAI Rules: How AI Vendors Stack Up

One Year of EU GPAI Rules: How AI Vendors Stack Up
TL;DR

The EU AI Act's GPAI obligations have been in force since August 2025. Anthropic leads on contractual data protections but like all three vendors has not published the required training data summary. Google is transparent about training on user data but lacks structured Annex XII compliance documentation. OpenAI's public documentation is the hardest to assess. Enterprise buyers should verify four specific items before integrating any GPAI model into products that may fall under the Act's high-risk obligations, which take effect August 2, 2026.

The GPAI Clock Has Been Ticking

The EU AI Act's General-Purpose AI obligations entered into force on 2 August 2025. That means every provider of a GPAI model offered in the EU market has been legally required to comply with Article 53 for a full year now. And on 2 August 2026, two days from now, the remainder of the Act kicks in, bringing high-risk AI system obligations into force and raising the stakes for everyone downstream of these models.

Enterprise buyers integrating AI into their products need to understand which foundation model providers have their compliance house in order. A vendor's public documentation tells you a lot about whether they have taken GPAI seriously or are still figuring it out. Here is a clause-by-clause comparison of what three major providers , Anthropic, Google, and OpenAI , currently disclose.

What Article 53 Actually Requires

Article 53 of the EU AI Act (full text here) imposes four core obligations on providers of general-purpose AI models:

  • Technical documentation (Annex XI): Training methodology, data sources, compute resources used, known limitations, and evaluation results. Must be provided to the AI Office and national authorities on request.
  • Transparency to downstream providers (Annex XII): Information enabling system integrators to understand capabilities and limitations, so they can comply with their own obligations under the Act.
  • Copyright policy: A publicly available policy to comply with EU copyright law, including a mechanism to respect the opt-out from text and data mining under Article 4(3) of the DSM Directive.
  • Public summary of training data: A sufficiently detailed summary of the content used for training, made publicly available.

For models designated as having systemic risk under Article 51 , which almost certainly includes the frontier models from all three vendors , Article 55 adds model evaluation, systemic risk mitigation, serious incident reporting, and cybersecurity requirements. Penalties under Article 101 reach 3% of annual worldwide turnover or €15 million, whichever is higher.

Anthropic: The Compliance Leader by Design

Anthropic's commercial terms (effective June 2025) read like someone had Article 53 open in another tab while drafting them.

Data usage: The commercial terms state unequivocally: "Anthropic may not train models on Customer Content from Services." That is not an opt-out. It is a hard prohibition embedded in the contract. For enterprise buyers subject to GDPR and the AI Act's downstream obligations, this is the gold standard , your data is not training data, period.

EU establishment: Anthropic Ireland, Limited serves EEA, Swiss, and UK customers. Having a legal entity in the EU is not explicitly required by Article 53, but it signals operational commitment to EU regulatory compliance and gives the AI Office a direct enforcement target.

DPA: A Data Processing Addendum is incorporated by reference into the commercial terms. This matters because Article 53's transparency obligations intersect with GDPR's data processing requirements , downstream providers need contractual clarity about what happens to their data.

The consumer gap: Anthropic's consumer terms for Claude.ai tell a different story. They state Anthropic "may use Materials to provide, maintain, and improve the Services and to develop other products and services, including training our models, unless you opt out." Consumer users can opt out through account settings, but the default is training-on. And even after opt-out, data flagged for safety review is still used for training. This split , commercial is locked down, consumer is opt-out , is worth noting if your team uses both Claude API and Claude.ai.

What is missing: Anthropic has not published a dedicated Article 53 technical documentation package. Their research page includes model cards and system prompts, but nothing labeled as Annex XI or Annex XII compliance documentation. The training data summary is also absent from public view , the most significant gap in an otherwise strong compliance posture.

Google Gemini: Transparently Training on Your Data

Google's Gemini Apps Privacy Hub (updated July 2026) is remarkably direct about data usage. The contrast with Anthropic's commercial terms could not be sharper.

Data usage for training: Google states plainly that your data is used to "Maintain and improve our services" and "Develop new services," and that "These uses extend to the generative AI models and other machine-learning technologies powering our services." The Privacy Hub even warns: "Please don't enter confidential information that you wouldn't want a reviewer to see or Google to use to improve our services, including machine-learning technologies." This is not a bug in Google's documentation. It is the business model.

Human review: Google confirms that "Human reviewers (including trained reviewers from our service providers) review some of the data we collect for these purposes." Reviewer-annotated data is retained separately from user-deletable activity and is not deleted when you delete your Gemini history.

Copyright policy: Google has not published a standalone EU AI Act copyright compliance policy. Their general copyright policy exists in the broader Google Terms of Service, but there is no specific mechanism for the DSM Directive Article 4(3) opt-out that Article 53 requires GPAI providers to respect.

Training data summary: Not publicly available. Google publishes research papers describing training methodologies, but a "sufficiently detailed summary" of training data content as required by Article 53(1)(d) is not accessible to the public.

The enterprise question: For businesses integrating Gemini into products subject to the AI Act's high-risk requirements (which kick in August 2, 2026), Google's training-on-default posture creates a chain of liability. If your system is high-risk, you need the model provider's Annex XII transparency documentation. Google has not published it. That gap should concern any enterprise buyer.

OpenAI: The Documentation Black Box

OpenAI's legal pages are client-side rendered JavaScript applications, which makes scraping them impossible , and, relevantly, makes them harder for regulators, auditors, and enterprise compliance teams to parse programmatically. This is not a compliance strategy. It is an accessibility problem that doubles as a transparency problem.

What we know from public sources: OpenAI's Enterprise API terms prohibit training on customer data, similar to Anthropic's commercial stance. But OpenAI's consumer products (ChatGPT) use data for training by default, with an opt-out available. Their GPT-4o system card represents the closest thing to Article 53 technical documentation , it includes evaluation results, safety testing, and risk categorisation. But it does not map to Annex XI or Annex XII structure, and it is published as a research artifact rather than a compliance document.

Training data summary: Not published. OpenAI has moved away from detailing training data sources in recent model releases, making this arguably the weakest area across all three vendors examined.

EU presence: OpenAI has an EU entity (OpenAI Ireland Limited) and has engaged with the AI Office. But there is no publicly available, structured compliance package mapping OpenAI's frontier models to Article 53 and Article 55 requirements.

The Comparison at a Glance

Here is how the three vendors stack up against the four Article 53 obligations based on publicly available documentation as of July 2026:

RequirementAnthropicGoogleOpenAI
No-training guarantee (commercial)Yes (contractual)NoYes (enterprise API)
EU legal entityYes (Ireland)Yes (Ireland)Yes (Ireland)
DPA availableYes (incorporated)AvailableAvailable (enterprise)
Public training data summaryNoNoNo
Annex XI/XII documentationPartial (model cards)Partial (research papers)Partial (system cards)
Copyright policy with DSM opt-outNo standalone policyNo standalone policyNo standalone policy

Nobody gets a perfect score. The most striking finding: not one of the three major GPAI providers has published a "sufficiently detailed summary" of their training data as required by Article 53(1)(d). This is not a niche compliance checkbox. It is one of only four obligations in the article. A full year after the rules took effect, it remains universally unaddressed.

What Enterprise Buyers Should Verify

If you are integrating a GPAI model into a product (talk to us about an independent vendor assessment) that may fall under the high-risk AI system obligations , which apply from August 2, 2026 , here is what to ask every vendor:

  1. Show me your Annex XII transparency documentation. Not a blog post. Not a research paper. The structured documentation the AI Act requires for downstream providers. If they cannot produce it, you are accepting compliance risk blind.
  2. Where is your training data summary? Article 53(1)(d) applies to all GPAI providers, not just those with systemic risk. If the vendor cannot point to a public URL with a sufficiently detailed summary, they are non-compliant with a core obligation.
  3. What is your copyright opt-out mechanism? The DSM Directive's Article 4(3) requires GPAI providers to respect rightsholder opt-outs. Ask for the mechanism. If the answer is unclear, the training data's legal basis may not hold up under regulatory scrutiny.
  4. Is my data used for training? Do not accept "we have a privacy policy" as an answer. Get it in contractual language. Anthropic's commercial terms set the bar. Anything less should be priced into your risk assessment. For help evaluating vendor compliance, see our pricing or get in touch.

The Bottom Line

Anthropic leads on contractual data protections but shares the industry-wide failure to publish a training data summary. Google is transparent about training on your data but has not published the structured compliance documentation the AI Act requires. OpenAI's compliance posture is the hardest to assess from public documentation, which is itself a transparency problem.

The EU AI Office has clear enforcement powers under Article 88 through 94, including the ability to request documentation, conduct evaluations, and require corrective measures. Providers of GPAI models placed on the market before August 2, 2025 have until August 2, 2027 to bring existing models into compliance under Article 111(3). But for models placed on the market after that date, the clock started ticking immediately. A year in, the documentation gaps are real (see what a proper audit looks like) , and they create real downstream liability.

Written by David Swan, reviewed and fact-checked against primary regulatory sources. AI-assisted but human-directed.

Frequently asked questions

When did the EU AI Act's GPAI obligations take effect?

The General-Purpose AI obligations under Chapter V of the EU AI Act entered into force on 2 August 2025. Providers of GPAI models placed on the market after that date were required to comply immediately. Models already on the market before that date have until 2 August 2027 to achieve compliance under Article 111(3).

What are the four core GPAI obligations under Article 53?

Article 53 requires GPAI providers to: (1) prepare and maintain technical documentation including training methodology and evaluation results; (2) provide transparency information to downstream providers who integrate the model; (3) implement a policy to comply with EU copyright law including the DSM Directive opt-out; and (4) publish a sufficiently detailed summary of training data content.

Which AI vendor has the strongest data protection for enterprise customers?

Anthropic's commercial terms include a contractual prohibition on training models using customer data, with no opt-out required. Google explicitly uses Gemini user data for model training and improvement. OpenAI's enterprise API terms prohibit training on customer data, but consumer ChatGPT data is used for training by default with an opt-out available.

Have any major GPAI providers published their training data summary?

No. As of July 2026, none of the three major GPAI providers (Anthropic, Google, or OpenAI) has published a sufficiently detailed public summary of their training data content as required by Article 53(1)(d) of the EU AI Act.

What should enterprise buyers verify before integrating a GPAI model?

Buyers should verify four things: (1) whether the vendor provides structured Annex XII transparency documentation, not just research papers; (2) whether a public training data summary exists; (3) the vendor's copyright opt-out mechanism under the DSM Directive; and (4) whether customer data is used for model training, with contractual guarantees.

What penalties can GPAI providers face for non-compliance?

Under Article 101 of the EU AI Act, providers of general-purpose AI models can face fines of up to 3% of their annual worldwide turnover or €15 million, whichever is higher, for infringements of their obligations under the Act.