03.03.2026

HERE'S WHAT YOU CAN DO TO PROTECT YOUR CONTENT

Introduction: Generative AI – a blessing for everyone?


I would say no, particularly when it comes to copyright infringements. After all, AI is not only trained using its own data. It is also trained using works created by others. With articles, books, reports, analyses, specialist articles and editorial content. In other words, with the intellectual property of creators, publishers and authors, who are neither automatically consulted nor fairly compensated for this.

So this is certainly no boon for creators whose content is used to train AI systems without their explicit consent.

It therefore comes as little surprise to me that there are now lawsuits and class actions on this issue. Essentially, the point is that copyright-protected content cannot simply be used as a basis for training, at least not without the consent of the creators or appropriate compensation. At present, there is still a lack of clear, uniform regulations in this area. This is precisely why there are growing calls to establish binding frameworks that clearly regulate the handling of intellectual property in the context of AI and provide better protection for creators.

With AI-generated answers and formats such as Google AI Overviews, many publishers and authors are losing clicks on their websites and, as a result, often also reach, reader loyalty and, in some cases, revenue. The original source therefore continues to provide the content, but in reality it is only Google that benefits from this.

Good blog posts, specialist articles or guides are, after all, meant to be found. They are intended to build trust, demonstrate expertise and, ideally, generate qualified enquiries. This is precisely where the conflict of objectives lies: protecting content can also mean that it loses visibility and, in some circumstances, no longer appears in search results as intended. Neither in Google Search nor in the responses from AI systems, which are, after all, extremely valuable for generating enquiries.


With Google Extended, Google has already taken a first step towards an opt-out for AI use. At the same time, the discussion shows that many companies and publishers feel this does not go far enough – particularly because transparency, control and accountability are still lacking.

So whilst there is still a great deal of uncertainty surrounding this issue, we’ll take a closer look at what we can do in the meantime to protect our content. Here, we distinguish between openly accessible content and content that requires comprehensive protection.

As already mentioned, these are particularly useful when the aim is to build trust and raise visibility. Examples include:

  • Blog articles offering in-depth expertise
  • Specialist articles
  • helpful explanatory texts
  • Content for positioning
  • Clearly articulated viewpoints on relevant topics
  • Case studies and client testimonials

It is precisely this kind of content that is absolutely vital for building reach and demonstrating our expertise. And this is becoming increasingly important in the age of AI. Enquiries generated by AI systems are particularly valuable. After all, when an AI recommends us, it immediately gives us a significant trust advantage. In fact, there is a huge difference between the quality of leads generated purely via search engines and those coming via AI systems. I see this time and again in practice. That is why we also talk about the shift from SEO to GEO. In other words, optimising content not only for search engines but also for language models such as ChatGPT, Gemini, Claude, etc. The key here is, above all, to create regular, high-quality content that is extremely relevant to the target audience (see my other posts under ‘News’ for more on this).

First and foremost, it is important to note that there is currently no such thing as complete protection. Anyone who publishes content online cannot prevent it from being copied, adapted or used in other ways with 100 per cent certainty. However, there are measures that can be taken to raise the barriers and regain at least some control.

 Well-known AI crawlers can, at least in theory, be blocked via the robots.txt file. These include, for example, GPTBot and Google-Extended.

This is not a perfect safeguard. Anyone who disregards it can still access the content. But it sends a clear signal: this content is not authorised for unrestricted use by AI.

A brief digression: Understanding AI crawlers: for example, Google-Extended and GPTBotfhin appear in standard Google searches.

GPTBot is OpenAI’s web crawler, which collects content to train and improve AI models.

  • It collects publicly available content from the web
  • can also be excluded via robots.txt
  • is part of the infrastructure behind models such as ChatGPT

The same applies here:
Blocking is a signal, but not technically absolute protection. Both systems illustrate how the web is changing:

  • Content is no longer collected solely for search engines
  • but is also processed and used by AI models

For website operators, this means:

They must increasingly decide who is allowed to use their content – and for what purpose.
As
Isaid, you have to weigh up what’s more important: relevance and visibility, or protecting the content.

My rule of thumb is that the content I publish is primarily intended to inform and help. This content may – and indeed should – be used by AI systems. It is precisely through this that I have already generated valuable enquiries. I’m also simply delighted when I receive positive customer feedback on it.  Documents such as my comprehensive playbooks, which contain in-depth knowledge, and workshop materials, on the other hand, are strictly protected and are generally only accessible locally.

Although it is often underestimated, clear legal notices are useful. Anyone who clearly states on their website that content must not be scraped automatically or used for AI purposes is at least laying the groundwork. This prevents some people from simply copying articles such as this one, plugging them into ChatGPT, having them reworded, and then thinking they’re particularly clever.

Some will ignore this nonetheless, but anyone who is genuinely reputable will respect it.

So you can (and should): 

  • Include copyright notices
  • Clearly set out terms of use
  • Explicitly highlight training restrictions

That’s definitely my top tip. For me, authentic – or unique – content offers the best protection. So make sure your content isn’t easily copied, for example:

  • your own data / use cases / references
  • personal experiences with specific practical examples
  • specific perspectives
  • your own language

And please don’t copy content yourself or simply let AI generate it. AI can tell pretty quickly whether your content is original or not. The more original it is, the higher your content will be ranked. Particularly with specialist articles, your own contribution is essential. AI can help, but it doesn’t automatically replace the intellectual effort behind a truly strong piece of writing. For me, that’s actually rule number one. It’s also, of course, a prerequisite for being able to include copyright notices.

So do make an effort; don’t just copy and paste it into the AI. Check your sources and only write about topics you really know about.

The situation is, of course, different for content whose value lies not primarily in its reach, but in exclusive access. This includes, for example:

  • exclusive data or analyses
  • in-house methods, frameworks or processes
  • paid content
  • sensitive knowledge repositories such as playbooks, workshop handouts and course materials
  • content that is directly monetised

a) Technically shield sensitive content

If content is to be properly protected, stricter access restrictions are the only effective solution. These include, for example:

  • Login areas
  • genuine paywalls
  • restricted members’ areas
  • internal documentation outside the public website

This is significantly more effective than a mere warning for bots. At the same time, however, it also reduces visibility. This is precisely why one should make a very deliberate decision as to which content really needs this protection.
 

b) Controlling bot access

Technical solutions such as Cloudflare can help to detect automated scraping and better manage access. This at least makes certain forms of mass data extraction more difficult. We have been using Cloudflare for our web and software projects for quite some time now (more on this in my next expert talk with our security expert Andreas).

Personal views on the current situation:

AI may be a tool for progress, but it must not be a free pass to exploit intellectual property for commercial gain without permission.

It is unacceptable for the major AI providers simply to help themselves to the work of others. Progress based on ‘data theft’ is, in the end, no real progress at all. And does it not actually do more harm than good to humanity?

Behind every text, every photo and every piece of research lies time, expertise and genuine passion. Simply labelling this as ‘training material’ devalues human labour. On top of that, AI is set to transform the world of work on a massive scale. The fact that these systems are now being fed, of all things, with the content produced by the very people whose jobs they threaten is almost cynical. That is why I believe we need clear regulations and boundaries: greater transparency, fair choices and effective protection for intellectual property.

Sources:

https://developers.openai.com/api/docs/bots

https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers?hl=de

https://www.eff.org/deeplinks/2023/12/no-robotstxt-how-ask-chatgpt-and-google-bard-not-use-your-website-training

https://www.boerse-express.com/news/articles/google-oeffnet-ki-opt-out-fuer-verlage-neue-regeln-fuer-sichtbarkeit-881525

https://www.cloudflare.com/the-net/building-cyber-resilience/regain-control-ai-crawlers

https://www.heise.de/news/Googles-KI-Zusammenfassungen-Opt-out-fuer-britische-Medienhaeuser-angekuendigt-11219699.html

https://blog.google/company-news/inside-google/around-the-globe/google-europe/cma-response

https://www.eff.org/deeplinks/2023/12/no-robotstxt-how-ask-chatgpt-and-google-bard-not-use-your-website-training?

 

AI transparency notice: AI assists me in creating the AI Snacks. However, the content is based on reliable sources, my experience from real-world projects and questions that clients repeatedly ask me on specific topics. Ultimately, each post contains a significant amount of my own input.

To improve readability, we have omitted gender-neutral phrasing in the text – naturally, the text always refers to everyone.

The content on this website is provided for information purposes and may be read and linked to. It is not permitted to copy the content in whole or in substantial parts without consent, to input it into AI systems, to process it automatically or to use it as the basis for your own content.