The fight over AI training data
All AI has to be trained. ChatGPT and its ilk scrape their training data together by grazing the internet, but for a great deal of AI that does not work. Those systems need specific data.
If an AI wants to tell a pupil which topics they should practise further, the collected works of Goethe are of no use to it. Nor are they if an AI wants to draw up a charging schedule for your electric car, calculate the fastest route from A to B, or interpret the output of cameras. In those cases specific data are needed. The AI has to learn to make good decisions. Sometimes the training data remain important even after the first training session. Particularly where an AI is updated, it may be necessary to run the whole training sequence again. AI looks clever, but at unexpected moments it can be rather stupid.
Many of these training data are personal data. They are, after all, often data about human activity entered into a system. With pupils it is a matter of results linked, in the Netherlands at least, to a pseudonymised account. That is to say that the account is unique and can be traced back to the pupil concerned via third parties, but that the supplier usually cannot do so itself. Pseudonymised data are personal data too, however. Charging data from cars can often be linked to a driver as well, and the same goes for the location data needed to determine traffic flow and thus the best route.
In almost all contract negotiations where training data play a part, the parties end up fighting over those data. It goes like this: the supplier says it needs usage data to train its systems, and the customer does not want the data used for that. The customer says it owns the intellectual property in the data and therefore does not have to share them, and that the privacy rules, the GDPR, prohibit sharing in any event.
The heat of the discussion is fuelled by the focus on data in many organisations. Data is the new gold, after all. That leads to situations in which, even where no added value whatsoever can be discovered in holding the data, parties are still very keen to keep them exclusive.
Viewed from a distance, you would say the parties’ interests coincide nicely. It is of course in the customer’s interest for the AI to be well trained. The point is, though, that for the intelligence of the AI it usually makes no difference whether the customer’s data are used or a competitor’s (not always, mind; some training is customer-specific). Where that is so, the customer says: take the competitor’s data instead, then they can have the associated privacy hassle.
To which the supplier replies that it hears the same story from the competitor, and therefore applies, and must apply, an unbending policy of treating everyone alike. There is also often a degree of mistrust. The thought is then that the data will be used not so much to improve the products as to improve the marketing. Particularly where data about vulnerable groups such as children are concerned, that provokes strong resistance, and not entirely without reason.
Solutions? As far as the intellectual property angle is concerned, it is not that difficult in my view. Within intellectual property, a lawyer can indicate fairly precisely who may do what, and where the centre of gravity of that ownership ought to lie, although it is highly questionable whether any statutory intellectual property in these data exists at all. Where data are concerned, the result is always a patchwork of legal workarounds in which “grip” on the data is arranged so as to construct a kind of ownership. The tools for that are confidentiality clauses and agreements about who actually controls the data and on whose servers they sit.
Then privacy. Contrary to what the customer often says, using data to train an AI is frequently perfectly permissible. Generally speaking, training AI is a legitimate interest of the supplier. And of the customer too, for that matter: the customer also has an interest in an AI that makes fewer mistakes. The supplier may therefore use those data for that purpose, unless the interference with the privacy of the data subjects is too great. With driver data and pupil data, the interference with the privacy of the drivers and pupils seems very manageable.
It can be different in other cases, so parties must always look carefully at precisely which data are used, how to anonymise them as far as possible, and what risk the data subjects (still) run. With medical data, for instance, the rules on medical professional secrecy play a complicating role as well. Once the parties have established that it is permitted, the data subjects must also be told that it is happening and what rights they have. Who has to do that depends on the precise data flow.
There are also several possibilities as regards the division of roles between supplier and customer. The obvious arrangement is for the supplier to be the controller for the processing for training purposes. But from the customer’s point of view, the customer can be the controller too: where the assignment consists of making the AI smarter, for instance. It is most practical for controllership to lie with the supplier, since as controller it can then continue to have the data at its disposal even once the contract with the customer comes to an end.
In short: there is a good deal of fighting over the data, but ultimately an open conversation will leave room for all the interests to be aired. Customer and supplier almost always find a workable compromise. Certainly with our help.