Liaise Advocaten logo
8 min reading time Published: 02-03-2023 | Updated: 02-03-2023

Fighting the AI

In my previous blog I wrote that many authors are worried about AI. Their profession is under pressure. AI can take over many of their assignments. That goes directly to their livelihood. All the more galling, then, that AI can only function by copying the works produced by those same authors into its own system and analysing them. The author thus feeds the beast that threatens to devour them.

Under the old Dutch Copyright Act, the author or, say, their publisher could oppose the use of their work by an AI. Except that, ironically, three years ago the Copyright Act was amended. And now an AI may use everything. The legislature has thereby, surely unintentionally judging by the explanatory notes, knocked a powerful weapon in the fight against AI out of authors’ hands. So the (mostly American) AI eats away at our authors’ income with the help of our authors’ works. I called the timing of the introduction of the AI provisions into the Copyright Act ironic above, but other descriptions are possible too. With the knowledge we have now, extremely unfortunate at the very least.

The role of the publisher or other user of the author’s work is a double one. Viewed as a cost-benefit matter, AI is not necessarily bad for them. Non-fiction publishers in particular can save substantially on costs. In that sense they have an interest in feeding the AI as many works as possible. After all: the more the AI has in its systems, the better it functions. And the more of the authors’ work is fed to the AI to train it, the less those authors may be needed in time. AI can, however, also threaten these publishers’ business model. What, after all, is the role of this intermediary if end users can also ask the AI directly to entertain or inform them? The output of AI is admittedly still limited in volume (it cannot reel off a whole novel), but I have no doubt that these limitations will disappear very quickly.

(It seems to me there are two roads from here: the first is that AI develops further, becomes smarter and smarter, digests more and more data, and can ultimately write perfectly in several genres. The other road is that AI, though only just started, is already running into certain inherent problems that stand in the way of further development, so that it remains the somewhat deservedly overlooked hack. I am inclined to think the first road the more likely.)

The author and the publisher would do well to think about how they want to deal with AI in this respect. Whether the publisher will make the reservation described below or not, for instance, can be important to the author. The same applies to any action against an AI that uses the author’s work unlawfully.

The Copyright Act

The Act is not entirely disastrous. It does give the copyright owner a weapon against AI. The AI may not use an author’s work where the author does not want their work used by the AI. Article 15o of the Copyright Act says that the author or publisher must expressly reserve copyright. That statement must be express and made in an appropriate manner, for instance by a notice in “machine-readable” form. It has been suggested that terms such as “noAI” be included in the metadata. What exactly counts as machine-readable seems to me to depend just as much on how clever the reading machine is. I cannot imagine that an AI scouring the internet is unable to recognise the words noAI where they appear alongside a particular document. That would make those words “machine-readable” and the problem is solved. Incidentally, the Act says that “copyright must be reserved”. That means a broad reservation of copyright should in principle be sufficient.

The other weapon the author or publisher can potentially deploy is that the AI must have obtained lawful access to the work. So the AI may not use works floating about on “free” book sites, for example, which anyone can work out did not get there legally. It seems to me, incidentally, that in any dispute about this it is for the party relying on it to prove that access to the works was lawful. That is an important allocation of the burden of proof. Any other allocation would be unreasonable for the rightholders (the author and the publisher): they cannot, after all, prove that they did not make the work accessible to the crawling AI. The absence of a fact is often impossible to prove.

The AI can, further, always gain access to works that are not behind a login. An interesting question is then whether a notice on the website saying “no access for AI and other robots” makes that access, and thereby the use, unlawful after all. Such a notice can in the first place serve as the “noAI” statement described above (“copyright reserved”). But it might perhaps also undermine the lawfulness of the visit by an AI robot and, through that route (that of lawful access), make it possible to prevent use by an AI.

In general terms I do not think that is so. A website operator does not have a legally relevant relationship with visitors such that it can impose conditions on who those visitors may be. It is different where there is a login. Where a website requires the visitor to confirm, by ticking a checkbox for instance, that they are over 18 (an age the average AI robot does not reach, incidentally), the visit by a younger visitor is in principle not lawful.

It seems to me that where a website operator can regulate access, it can also attach conditions to that access. Think of payment as a condition, for a start. But other conditions are of course possible too. Where something or someone then enters the website without meeting those conditions, that access is not lawful. So where a robot that has, on being asked, ticked the box saying it is not a robot enters the site, that lying robot does not have lawful access to the content behind it. And the content in question may not be used for data mining and therefore for AI applications.

But what is the practical use of this legal hair-splitting? Where the output of the AI gives no information about the sources used, you as a rightholder cannot possibly prove that the AI must have pinched its wisdom from you.

The shortest route is often forgotten. Not by me, so I asked the AI itself (see below). Do you happen to know Harry Mulisch’s The Assault? “No,” it said. The point is: I think the AI is lying. It ought to be ashamed. All those pirated books floating about on open sites have of course long since been scanned and processed by ChatGPT. That Mulisch’s classic should be missing from them is extremely improbable. So I pressed a little further. The initial certainty is replaced by: I don’t know…

The conclusion of this blog is that for an author the fight against AI is hard going. And that your stance may further depend on your position. If you are an author, try to make sure that notices saying NO TRAINING FODDER FOR ROBOTS appear everywhere (though preferably a standard text such as “noAI” or “copyright reserved”), and that your material sits behind a wall as much as possible. If you are a publisher and you are quite happy with that AI, then by all means do not. If you are an author negotiating with a publisher, it makes sense to make concrete arrangements about this. And above all keep an eye on developments. The notice, and how to give it, may yet be fleshed out in best practices or by bodies such as the European Commission.

To close, the little conversation with ChatGPT.

PROMPT: Is the book The Assault by Harry Mulisch part of your training material?

ChatGPT

No, “The Assault” by Harry Mulisch is not part of my training material. My knowledge is based on text and information available up to and including September 2021. If you have specific questions about the content of the book, or would like more information, I can still try to help you on the basis of the knowledge available to me.

PROMPT: So the book was not scanned to train you?

ChatGPT

I do not have access to my training data, but I was trained on a mixture of licensed data, data created by human trainers and publicly available data. I cannot provide specific information about which books, including “The Assault” by Harry Mulisch, were or were not used during my training. My knowledge is based on a wide range of textual sources, but I cannot confirm whether this particular book was scanned as part of my training material. If you have specific questions about the content of the book or anything else, I am happy to help with the knowledge I have.

noAI

NO TRAINING FODDER FOR ROBOTS

NO DATAMINING

NO SILICON VISITORS ALLOWED

AI STAY OUT

FLESH AND BLOOD ONLY

COPYRIGHT RESERVED

Liaise Advocaten
Lawyer

Jetse was a partner at our firm for many years and still works closely with us.

Need help? Talk to one of our experts

  • Alexandra Iedema
  • David Allick
  • Merel Teunissen
  • Jaap Versteeg
  • Roland Wigman
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Response within 1 working day

Your message goes to one of our lawyers.

  • Merel Teunissen
  • Jaap Versteeg
  • Charissa Koster
  • Roland Wigman
  • Alexandra Iedema
  • David Allick

How can we help?

How do we reach you?