پژوهشنامه پردازش و مدیریت اطلاعات

پژوهشنامه پردازش و مدیریت اطلاعات

ملاحظات حقوقی بهره‌برداری از داده‌های آموزشی دارای حق مؤلف در مدل‌های زبانی بزرگ: با تاکید بر مجموعه ‌دادۀ Books3 در مدل LLaMA2

نوع مقاله : مقاله پژوهشی

نویسنده
استادیار، گروه پژوهشی مدیریت اطلاعات و سازماندهی دانش، پژوهشکده اسناد، سازمان اسناد و کتابخانه ملی ج.ا.ا، تهران.
10.22034/jipm.2026.2072896.2195
چکیده
پژوهش حاضر با هدف بررسی وضعیت حقوقی استفاده از آثار دارای حق مؤلف در آموزش مدل‌های زبانی بزرگ، با تمرکز بر مجموعه‌داده Books3 و مدل LLaMA2، انجام شده است. پژوهش با تحلیل تطبیقی نظام‌های حقوقی ایران، ایالات متحده، اتحادیه اروپا، فرانسه و آلمان، در پی تبیین مبانی مسئولیت حقوقی بازیگران مختلف و ارزیابی کفایت قواعد موجود حقوق مؤلف در مواجهه با آموزش مدل‌های زبانی بزرگ است.
پژوهش کنونی با رویکرد کیفی، مبتنی بر تحلیل محتوای اسناد و تحلیل تطبیقی نظام‌های حقوقی انجام شده است. داده‌ها از طریق بررسی قوانین و مقررات، اسناد بین‌المللی، آرای قضایی، مستندات فنی مربوط به Books3 و LLaMA2 و مطالعات علمی گردآوری و با روش تحلیل تطبیقی حقوقی مبتنی بر رویکرد کارکردی تحلیل شده‌اند. مجموعه‌داده Books3 و مدل LLaMA2 به‌عنوان نمونه‌ای تجربی برای آزمون تحلیل‌های حقوقی مورد استفاده قرار گرفته‌اند.
یافته‌های پژوهش نشان داد که چالش اصلی در این حوزه صرفاً مشروعیت استفاده از داده‌های آموزشی نیست، بلکه احراز منشأ داده‌ها، قابلیت اثبات استفاده از منابع دارای حق مؤلف و تعیین مسئولیت بازیگران مختلف فرایند آموزش و بهره‌برداری از مدل‌های زبانی اهمیت اساسی دارد. بررسی تطبیقی نشان داد که در ایالات متحده مشروعیت استفاده عمدتاً بر پایه دکترین استفاده منصفانه و ارزیابی موردی انجام می‌شود، در حالی که اتحادیه اروپا در قالب دستورالعمل ۲۰۱۹/۷۹۰ چارچوب استثنای متن‌کاوی و داده‌کاوی را تعیین کرده و فرانسه و آلمان، به‌عنوان کشورهای عضو، این چارچوب را با تفاوت‌هایی در قوانین ملی خود اجرا کرده‌اند؛ شرایطی مانند دسترسی قانونی، رعایت الزامات قانونی و در برخی موارد امکان اعمال حق انصراف در هر دو سطح اتحادیه‌ای و ملی مطرح است. در مقابل، در حقوق ایران چارچوب مستقلی برای استفاده از منابع دارای حق مؤلف در آموزش مدل‌های زبانی بزرگ پیش‌بینی نشده و این موضوع موجب ابهام در تعیین مسئولیت توسعه‌دهندگان و سایر بازیگران شده است.
نتایج پژوهش بیانگر آن است که توسعه مسئولانه مدل‌های زبانی بزرگ مستلزم ایجاد چارچوبی شفاف برای استفاده از داده‌های آموزشی است. بر این اساس، بازنگری در مقررات مالکیت ادبی و هنری، پیش‌بینی قواعد مشخص درباره آموزش مدل‌های هوش مصنوعی، ایجاد الزامات قانونی برای شناسایی و مستندسازی داده‌های آموزشی، توسعه سازوکارهای مجوزدهی جمعی برای بهره‌برداری از منابع و تعیین حدود قانونی استفاده از داده‌های دارای حق مؤلف می‌تواند زمینه ایجاد تعادل میان حمایت از حقوق پدیدآورندگان و توسعه فناوری‌های هوش مصنوعی را فراهم سازد.
کلیدواژه‌ها

عنوان مقاله English

Legal Considerations in the Use of Copyright-Protected Training Data in Large Language Models: Insights from the Books3 Dataset and the LLaMA2 Model

نویسنده English

Zeinab Papi
Assistant Professor, Department of Information Management and Knowledge Organization, Research Center for Documents, National Library and Archives of the Islamic Republic of Iran, Tehran
چکیده English

This study aims to examine the legal status of the use of copyrighted works in the training of large language models (LLMs), with particular emphasis on the Books3 dataset and the LLaMA2 model. Through a comparative analysis of the legal frameworks of Iran, the United States, the European Union, France, and Germany, the study seeks to identify the legal bases of liability for the various actors involved and to evaluate the adequacy of existing copyright rules in addressing the challenges posed by the training of LLMs. The research adopts a qualitative approach based on documentary and comparative legal analysis. The data were collected from legislation and regulations, international legal instruments, judicial decisions, technical documentation relating to Books3 and LLaMA2, and relevant academic literature. The materials were analysed using a functional comparative method, examining how each legal system addresses a common set of issues rather than comparing statutory wording alone. The Books3 dataset and the LLaMA2 model were employed as an empirical reference point for testing the legal analysis.
The findings indicate that the principal challenge extends beyond the mere legality of using training data. Rather, the identification of the provenance of training datasets, the ability to establish whether copyrighted works have been used, and the allocation of legal responsibility among the various actors involved in the development and deployment of LLMs are of fundamental importance. The comparative analysis demonstrates that, in the United States, the lawfulness of such use is primarily assessed under the doctrine of fair use on a case-by-case basis. By contrast, the European Union has established a harmonised text and data mining exception under Directive (EU) 2019/790, which France and Germany, as member states, have transposed into national law with certain variations; lawful use under this framework depends on conditions such as lawful access, compliance with statutory requirements, and, in certain cases, the right of rightholders to opt out. In Iran, however, no specific legal framework governs the use of copyrighted works for training LLMs, resulting in significant uncertainty regarding the liability of developers and other actors involved in the AI lifecycle.
Conclusion: The findings suggest that the responsible development of LLMs requires a transparent legal framework governing the use of training data. Accordingly, the study recommends revising national copyright legislation, introducing explicit legal provisions concerning the training of AI models, establishing statutory obligations for the identification and documentation of training datasets, developing collective licensing mechanisms for the lawful use of copyrighted works, and defining clear legal limits on the use of protected materials for AI training. These measures would contribute to achieving an appropriate balance between the protection of copyright and the continued development of AI technologies.

کلیدواژه‌ها English

training data
copyright
large language models
fair use
LLaMA2
Books3
Iranian legal system

مقالات آماده انتشار، پذیرفته شده
انتشار آنلاین از 11 مهر 1405

  • تاریخ دریافت 22 بهمن 1404
  • تاریخ بازنگری 28 شهریور 1405
  • تاریخ پذیرش 07 مهر 1405