

0·
6 days agoPeople learn from books and existing code. Does that means that anything we write is derived work?
An LLM knowledge matrix is a neural network, not a database. It does not contains an exact copy of any given text. It may remember it, but human brain can do the same.
You have to prove LLM code is derived work in the same way you prove human work is derived work: check the difference between the original work and the derived work.
You can make the case that an LLM is not a person, you cannot make the case that LLM work is not creative. The grounding justification of copyright lacked a foundational example of a non human creative process. If each single token emitted by the LLM is the result of math on the entire knowledge matrix, then each single token is derived from all the copyrighted documents used in training in a percentage you cannot even estimate. The mechanical transformation of an LLM is unknowable, unquantifiable and non deterministic. The copyright claim makes no sense in this case. A new shared mechanism must be put into place if we want to translate copyright through such mechanism. Lot of the training data used by LLM is synthetic so it pass through multiple layers of unknowable, unquantifiable, non deterministic “mechanisms”.