How does an author prove they are not an AI? Two tips.
Introduction
New works of art by painters long dead, tracks in the style and with the timbre of artist A sung by a computer, summaries produced in under a minute that pass the teachers’ critical scrutiny with flying colours: what can artificial intelligence (AI) not do these days? One press of a button and AI has already done it, spitting out an enormous volume of AI creations along the way. Creations that at first sight differ little from what painters, composers and people of flesh and blood make.
The question is whether copyright exists to protect all these AI creations. The answer is no. Copyright exists to protect the human author, by protecting their creation. That is driven by economic motives: authors are important to our society, so we must give them the chance to earn money from their work. That is why we must protect their work against free-riders. Moral and ethical motives play a part too: it is not right to profit from the work of authors without paying them for it, and for that reason as well we must protect their work. Neither reason for protecting the author is a reason to protect AI. And so it is not protected. The work of an AI is in principle unprotected.
The problem
The closeness of AI creations to human creations raises a quite different problem, however. How does an author defend themselves against the accusation that the work was made not by them but by an AI?
The search, and attribution
This blog describes a rather frustrating search for the possibilities of such a defence. Before we begin that search, it is worth pointing out how relative the need to mount such a defence actually is.
It comes to this: it is not the author who must prove that a creation bearing their name was not made by an AI; it is the person alleging that the creation was not made by the author who will have to prove it. And that is extremely difficult, if not impossible. Except, of course, where the author says in so many words that they are not the one who made the work.
Article 4 of the Dutch Copyright Act plays an important role here. That article creates what is known as a “presumption of authorship” for the author. It means that under the Copyright Act, the person named on a creation as its author is deemed to be the author. Strictly speaking that does not help the author, because the question here is not whose the creation is, but whether there is a creation at all.
We are inclined to read the article as meaning that where a human being is named as the author, that person is deemed to be the author of the creation concerned, and that the creation is thereby, in the same breath, deemed not to have been produced by an AI.
Another element (developed further below) is that a human author can edit the output of the AI. Where they do so in a way that is sufficiently creative in copyright terms, that author acquires copyright in those changes, changes often intimately bound up with the AI work. Copying the edited AI work will therefore amount to a copyright infringement against the author who edited it.
ChatGPT
Now back to our search for an answer to the question how to give the author, apart from attribution, a better evidential position.
An obvious step would be for the author to record the various steps in the creative process, in sketches or notes. That would allow the creative choices originating from the author themselves to be demonstrated. But it can be objected that it is questionable whether the author really made those creative choices. A fake author can of course simply ask the AI to produce such sketches or notes. So this route does not appear to get us very far.
A next idea was to deploy AI to prove that AI was not involved in the creation. We confined ourselves mainly to texts, since as writing professionals those are what we can handle most easily.
ChatGPT
When we asked ChatGPT whether it had written a text (which it had written itself), it answered:
‘Yes, I wrote the text you have just quoted. As an AI model I am able to generate text on the basis of the input I receive. The text I write is based on my training data and on algorithms that compose text in a coherent and relevant way.’
For an author, of course, what matters is a negative answer to that question. On uploading a text we had written ourselves, ChatGPT answered:
“No, the text you have quoted was not written by me. If you have any other questions or need help, do not hesitate to ask.”
This looks like a solution to the problem, but unfortunately the AI is not really reliable. It often just talks nonsense, to put it unkindly. Texts written by ChatGPT are first characterised as written by it, but the very next day labelled ‘human’. The reverse happens too. ChatGPT has no memory it can dig into, so the output, including its answer to this question, is generated afresh each time. And since an AI is language-based, there is a good chance it is simply saying whatever sounds right. That does not mean a negative answer is worthless. It only means there is no complete certainty that ChatGPT did not write the text itself.
As ChatGPT itself writes:
‘No, ChatGPT has no memory in the sense that it can remember earlier interactions or information from one conversation to the next. Every interaction with ChatGPT stands on its own and is not stored for future use. The model does not remember personal data or information about individual users between conversations. Each time you ask a question or make a remark, the response is generated on the basis of the immediate context of that conversation and the textual input provided. ChatGPT was designed with privacy and data security in mind, and it does not store personal information or data.’
So asking whether an AI is the author or not does appear to have (some) point. It at any rate adds some evidence to the proposition that the author, and not the AI, is the author. And thereby that the text or design is a creation protected by copyright. That this evidence is not conclusive does not mean it is not valuable.
Unfortunately, our further tests below (in connection with AI texts edited by a human) show that the force of this evidence is even lower than “not conclusive”. But first a remark about the so-called AI detectors.
AI detectors
The programmers have not been idle either. Particularly in connection with detecting fake news and AI-written essays and the like, several AI detectors have been developed. These include tools such as Turnitin, Copyleaks and OpenAI’s Classifier.[1] They are not 100% effective either, however. Turnitin says it is 99% accurate where texts are concerned, as does Copyleaks. That would be a fine score, but it may be doubted. We checked an AI-written text which Copyleaks assessed as 64.1% likely to have been generated by AI. Whereas if the starting point is 99% accuracy, you would have expected a higher percentage than 64.1% (that may, incidentally, be because the text was on the short side). But quite apart from that, it is of course striking that two AI detectors promise exactly the same score (99%). And 99% is an implausible figure, in the sense that 99% is a number often used to convey a small margin of error without specifying it. We cannot escape the impression that that is what has happened here. The real success rate may well be lower, a factor being that AI is still in its infancy. AI can and will develop further, and there seems every chance that, learning as it goes, it will close the recognisable gap with the human mind, so that the accuracy of these programmes may fall sharply. Put differently: even if the AI detector works well today, that does not mean it will still work well tomorrow. Bearing in mind, of course, that AI detection is also still in its infancy. In that sense there is a continuous arms race between the AI that will increasingly want to pass for human and the AI detection service that wants to unmask the machine.[2]
(Programmes for detecting AI-generated images are not watertight either. Too low a resolution, too small a part of the image, or a light edit can already render the detectors wholly unreliable.[3] The detectors are also often trained on images from one specific AI, so that pictures from other systems go unrecognised.[4])
The author edits AI output
A final issue is that an author can edit the output of an AI and thus lead the AI, and Copyleaks and the rest, to think that the input they have edited is the result of a human creation, while a part of it, small or very large, does indeed come from a computer.
In copyright terms things now get complicated: an adaptation of a text can itself be protected by copyright, without that affecting the protection of the original, unedited text.
Where an author sets to work on someone else’s text, someone else’s work, there are three variants, in ascending order of how far-reaching the adaptation is:
- The adaptation is minimal and is not itself “creative”: the adapter acquires no copyright, and only the original author has copyright in respect of the “adapted” work.
- The adaptation is more far-reaching and is itself “creative” too. Think of a translation or a thorough rewrite. The original author’s copyright still subsists in the “adapted” work, but copyright also subsists in the adaptation itself: the adaptation (the translation), that is. There are therefore two copyright owners in respect of the adaptation: the original author and the adapter.
- The adaptation is so far-reaching, and itself so “creative”, that an entirely new work comes into being. Only the adapter then has copyright in respect of the adaptation and is, in copyright terms, no longer an adapter but an author in their own right.
With that in mind, we first tested a substantial adaptation (somewhere between 2 and 3) and then a minimal one.
We did so as follows. First we asked the AI to write an essay about copyright and photographs on the internet. The following paragraph comes from that essay:
ChatGPT_: Copyright, in essence, protects the intellectual property of authors. This applies to photographs too, whether they were made in analogue or digital form. When someone takes a photograph, they are in many cases automatically the holder of the copyright. This means others cannot simply use the photograph without the author’s permission. Putting a photograph on the internet changes nothing about this basic principle._
Copyleaks recognised this passage as 64.1% likely to be AI-generated. ChatGPT also told us that this text could have been generated by it. By way of experiment we rewrote this passage substantially, without however altering the intent of the sentences or their order. This is the result:
At its core, copyright exists to protect the intellectual property of authors. That applies to photographs too, and it makes no difference whether they are analogue or digital. Where a person takes a photograph, they will in most cases acquire the copyright in that photograph automatically. If others then wish to use that photograph, they will have to ask that person’s permission. Where the photograph is used on the internet, that continues to apply.
We asked ChatGPT whether this was its text, and the answer was no:
No, the text you have quoted was not written by me, although it does deal with the same subject I discussed in the earlier essay: copyright in photographs and the protection of authors’ intellectual property.
Copyleaks assessed our rewritten text as 97.7% likely to be human.
That means that where this text is an adaptation without an entirely new work coming into being, the author (the present writer) now holds the copyright in his adaptation but not in the text generated by the AI.
We then rewrote the text minimally. So minimally that our adaptation probably does not meet the necessary threshold of creativity. This is the lightly edited text:
Copyright in essence protects the intellectual property of authors. This applies to digital or analogue photographs too. When someone takes a photograph, they are in many cases automatically the holder of the copyright. Others may therefore not simply use the photograph without the author’s permission. Putting a photograph on the internet changes nothing about this basic principle.
Asked about this text, ChatGPT says: “No, I did not write the text you have provided.” Copyleaks is all but certain that this text is human. With 96.4% probability this is a human text, according to Copyleaks.
This means a further erosion of the status of AI and AI detectors as an oracle on authorship. The genuine author will, after all, always be open to the accusation of having lightly altered an AI text to create the illusion of a work of their own making. So not only do AI and the AI detectors offer no 100% certainty about the authorship of a living author; given the ease with which an author can pass off an AI text as their own, they are of almost no help at all. For the time being.
What now?
Can the author do nothing, then? The point is for the author to lay down a credible creative trail of their work. That helps in defending against the allegation that the computer, not the author, made the work, even though a malicious author could lay down a comparable trail.
We would therefore say: keep the original works (texts and designs) in all their variants. Available registration tools or techniques such as blockchain may be of use here. A work with a date attached to it that can be verified in an external database will expose an AI imitation immediately.
And beyond that, a genuine author remains well protected for the time being as long as that author keeps putting their name to their work as author. An author with fewer scruples has a great deal to gain. Such an author can simply let the AI do the work, then quickly adjust the text a little and put their name to it as author, and then they are the author, because for the time being it is not possible to prove otherwise. That allows such an author to work many times faster and to build up an enormous body of work in a short time. Conduct of that kind cannot yet be detected. Whether you would sleep easily as a celebrated “fake” author is another matter, and doubtful. Technology marches on, and there is certainly a chance that this gap will be closed technically.
Conclusion
For the time being, AI itself and AI detectors offer some help only with unaltered texts. Given the relative ease with which that detective power can be circumvented by adding small changes, that help is very limited. For the moment, therefore, we see no way for the human author to prove conclusively that they, and not AI, made the work.
We recommend keeping texts (and designs) in their various stages, with dates. That way the author creates a reasonably credible trail. Registration techniques with a reliable date stamp could help.
But above all we recommend that the author states their name and their capacity as author with all their work.
With thanks to Bloeme de Boer
Afterword: this article was rewritten in the light of developments in the field of AI. An earlier version still attached importance to putting questions to AI.
[1] New Tool Can Tell If Something Is AI-Written With 99% Accuracy (forbes.com)
[2] Another Side of the AI Boom: Detecting What AI Makes – The New York Times
[3] How Easy Is It to Fool A.I.-Detection Tools? – The New York Times
[4] Hoe zie je of een foto gegenereerd is door AI? | EOS Wetenschap (in Dutch)