نوع مقاله : مقاله پژوهشی
عنوان مقاله English
نویسنده English
This study aims to examine the legal status of the use of copyrighted works in the training of large language models (LLMs), with particular emphasis on the Books3 dataset and the LLaMA2 model. Through a comparative analysis of the legal frameworks of Iran, the United States, the European Union, France, and Germany, the study seeks to identify the legal bases of liability for the various actors involved and to evaluate the adequacy of existing copyright rules in addressing the challenges posed by the training of LLMs. The research adopts a qualitative approach based on documentary and comparative legal analysis. The data were collected from legislation and regulations, international legal instruments, judicial decisions, technical documentation relating to Books3 and LLaMA2, and relevant academic literature. The materials were analysed using a functional comparative method, examining how each legal system addresses a common set of issues rather than comparing statutory wording alone. The Books3 dataset and the LLaMA2 model were employed as an empirical reference point for testing the legal analysis.
The findings indicate that the principal challenge extends beyond the mere legality of using training data. Rather, the identification of the provenance of training datasets, the ability to establish whether copyrighted works have been used, and the allocation of legal responsibility among the various actors involved in the development and deployment of LLMs are of fundamental importance. The comparative analysis demonstrates that, in the United States, the lawfulness of such use is primarily assessed under the doctrine of fair use on a case-by-case basis. By contrast, the European Union has established a harmonised text and data mining exception under Directive (EU) 2019/790, which France and Germany, as member states, have transposed into national law with certain variations; lawful use under this framework depends on conditions such as lawful access, compliance with statutory requirements, and, in certain cases, the right of rightholders to opt out. In Iran, however, no specific legal framework governs the use of copyrighted works for training LLMs, resulting in significant uncertainty regarding the liability of developers and other actors involved in the AI lifecycle.
Conclusion: The findings suggest that the responsible development of LLMs requires a transparent legal framework governing the use of training data. Accordingly, the study recommends revising national copyright legislation, introducing explicit legal provisions concerning the training of AI models, establishing statutory obligations for the identification and documentation of training datasets, developing collective licensing mechanisms for the lawful use of copyrighted works, and defining clear legal limits on the use of protected materials for AI training. These measures would contribute to achieving an appropriate balance between the protection of copyright and the continued development of AI technologies.
کلیدواژهها English