latest NewsNational

Delhi High Court Says OpenAI’s AI Training on Copyrighted Content Is ‘Fair Dealing’, Declines Relief to ANI

New Delhi: In a significant ruling that could shape the future of artificial intelligence and copyright law in India, the Delhi High Court on Friday held that OpenAI’s use of lawfully accessed copyrighted material for training its large language models (LLMs) falls within the ambit of “fair dealing” under the Copyright Act and, at this stage, does not constitute copyright infringement.

The decision came in a lawsuit filed by ANI Media Pvt. Ltd. in 2024 against OpenAI Inc. and OpenAI OpCo LLC.

The news agency had alleged that OpenAI unlawfully used its copyrighted news content to train the artificial intelligence models powering ChatGPT without obtaining permission.

The ruling marks the first time an Indian constitutional court has delivered a judicial opinion on whether the use of copyrighted material for training generative artificial intelligence models can qualify as fair dealing under Indian copyright law.

Pronouncing the order in open court, Justice Amit Bansal observed that, on a prima facie assessment, OpenAI’s storage and use of ANI’s original publicly retrieved material for training the large language models underlying ChatGPT was protected under Section 52(1)(a) of the Copyright Act.

The court held that such use, at this stage, could not be regarded as copyright infringement.

The High Court also examined the responses generated by ChatGPT through Retrieval-Augmented Generation (RAG) technology.

Justice Bansal observed that the AI-generated outputs were not substantially similar to ANI’s original material and therefore did not amount to infringement under the Copyright Act.

According to the court, ANI was unable to demonstrate that ChatGPT had memorised, reproduced or “regurgitated” its original copyrighted content while responding to user queries.

The judge held that there was no prima facie evidence showing that the AI model was reproducing ANI’s protected works in its responses.

On that basis, the court concluded that ANI had failed to establish a prima facie case warranting the grant of an interim injunction against OpenAI.

The High Court further observed that the balance of convenience favoured OpenAI.

It reasoned that restraining the company through an interim order at this stage could cause irreparable harm not only to OpenAI but also to the wider public, considering the broader implications for technological innovation and public access to AI-based services.

OpenAI, headquartered in San Francisco, United States, is the artificial intelligence research company behind ChatGPT and several other generative AI models used globally.

During the proceedings, ANI had argued that OpenAI used its publicly available copyrighted news reports to train its AI systems without authorisation.

The agency also alleged that ChatGPT could reproduce portions of its copyrighted material verbatim and occasionally generate inaccurate or fabricated information while wrongly attributing it to ANI.

OpenAI, however, maintained before the court that training an AI model was fundamentally different from reproducing copyrighted content.

The company argued that merely processing publicly available information for machine learning purposes could not be equated with copying or commercially exploiting the underlying work, drawing an analogy with reading a book without infringing its copyright.

The company also informed the court that the pre-training of its AI models was carried out outside India, with training datasets stored on servers located overseas.

It further submitted that its systems had been progressively refined to minimise the possibility of reproducing source material and that, after the completion of the training phase, the models no longer retained direct access to the original datasets.

During earlier proceedings in November 2024, OpenAI had informed the High Court that ANI’s website had already been placed on its “blocklist”, ensuring that the agency’s domain would not be used for future AI training exercises.

Apart from ANI, several media organisations, publishing bodies and industry associations had also joined the proceedings, expressing concerns over the use of copyrighted material for artificial intelligence training.

These included the Federation of Indian Publishers, the Digital News Publishers Association and the Indian Music Industry, all of whom raised broader issues relating to copyright protection in the age of generative AI.

दिल्ली हाई कोर्ट का बड़ा फैसला: OpenAI द्वारा कॉपीराइट सामग्री से AI प्रशिक्षण ‘फेयर डीलिंग’, ANI को अंतरिम राहत नहीं

नई दिल्ली: कृत्रिम बुद्धिमत्ता (एआई) और कॉपीराइट कानून के क्षेत्र में एक महत्वपूर्ण फैसला सुनाते हुए दिल्ली हाई कोर्ट ने शुक्रवार को कहा कि ओपनएआई द्वारा वैध रूप से उपलब्ध कॉपीराइट सामग्री का उपयोग अपने बड़े भाषा मॉडल (एलएलएम) के प्रशिक्षण के लिए करना प्रथम दृष्टया भारतीय कॉपीराइट कानून के तहत “फेयर डीलिंग” की श्रेणी में आता है और इसे फिलहाल कॉपीराइट उल्लंघन नहीं माना जा सकता।

यह आदेश एएनआई मीडिया प्राइवेट लिमिटेड द्वारा वर्ष 2024 में ओपनएआई इंक. और ओपनएआई ओपको एलएलसी के खिलाफ दायर मुकदमे पर सुनाया गया। समाचार एजेंसी ने आरोप लगाया था कि उसकी कॉपीराइट संरक्षित समाचार सामग्री का उपयोग बिना अनुमति के चैटजीपीटी जैसे एआई मॉडल के प्रशिक्षण में किया गया।

यह फैसला इसलिए भी ऐतिहासिक माना जा रहा है क्योंकि किसी भारतीय संवैधानिक न्यायालय ने पहली बार यह विचार व्यक्त किया है कि जनरेटिव आर्टिफिशियल इंटेलिजेंस मॉडल के प्रशिक्षण में कॉपीराइट सामग्री का उपयोग भारतीय कॉपीराइट कानून के तहत “फेयर डीलिंग” माना जा सकता है या नहीं।

खुले न्यायालय में आदेश सुनाते हुए न्यायमूर्ति अमित बंसल ने कहा कि प्रथम दृष्टया यह प्रतीत होता है कि चैटजीपीटी के अंतर्निहित बड़े भाषा मॉडल के प्रशिक्षण के लिए एएनआई की सार्वजनिक रूप से उपलब्ध मूल सामग्री का संग्रह और उपयोग कॉपीराइट अधिनियम की धारा 52(1)(a) के अंतर्गत संरक्षित है।

अदालत ने कहा कि उपलब्ध तथ्यों के आधार पर इसे कॉपीराइट उल्लंघन नहीं माना जा सकता।

न्यायालय ने रिट्रीवल-ऑगमेंटेड जेनरेशन (आरएजी) तकनीक के माध्यम से चैटजीपीटी द्वारा तैयार किए गए उत्तरों की भी समीक्षा की।

अदालत ने कहा कि इन उत्तरों और एएनआई की मूल सामग्री के बीच पर्याप्त समानता नहीं पाई गई, इसलिए इन्हें भी कॉपीराइट उल्लंघन की श्रेणी में नहीं रखा जा सकता।

अपने आदेश में न्यायालय ने कहा कि एएनआई यह साबित करने में असफल रही कि चैटजीपीटी ने उसकी मूल समाचार सामग्री को याद रखा, हूबहू दोहराया या उपयोगकर्ताओं के प्रश्नों के उत्तर में उसी का पुनरुत्पादन किया।

अदालत के अनुसार इस संबंध में प्रथम दृष्टया कोई विश्वसनीय साक्ष्य प्रस्तुत नहीं किया गया।

इन्हीं निष्कर्षों के आधार पर हाई कोर्ट ने कहा कि एएनआई अंतरिम निषेधाज्ञा प्राप्त करने के लिए आवश्यक प्रथम दृष्टया मामला स्थापित नहीं कर सकी है।

इसलिए इस स्तर पर ओपनएआई के विरुद्ध अंतरिम रोक लगाने का कोई आधार नहीं बनता।

न्यायालय ने यह भी कहा कि सुविधा का संतुलन ओपनएआई के पक्ष में है। अदालत के अनुसार यदि इस चरण पर अंतरिम रोक लगा दी जाती है तो इससे केवल ओपनएआई ही नहीं, बल्कि व्यापक जनहित और कृत्रिम बुद्धिमत्ता आधारित तकनीकी नवाचारों को भी अपूरणीय क्षति पहुंच सकती है।

अमेरिका के सैन फ्रांसिस्को मुख्यालय वाली ओपनएआई विश्व की प्रमुख कृत्रिम बुद्धिमत्ता अनुसंधान संस्था है, जिसने चैटजीपीटी सहित कई जनरेटिव एआई मॉडल विकसित किए हैं, जिनका वैश्विक स्तर पर व्यापक उपयोग किया जा रहा है।

मुकदमे की सुनवाई के दौरान एएनआई ने तर्क दिया था कि ओपनएआई ने उसकी सार्वजनिक रूप से उपलब्ध समाचार सामग्री का उपयोग बिना अनुमति के अपने एआई मॉडल के प्रशिक्षण में किया।

एजेंसी ने यह भी आरोप लगाया कि चैटजीपीटी कई बार उसकी कॉपीराइट सामग्री को शब्दशः प्रस्तुत करता है तथा कभी-कभी गलत या काल्पनिक जानकारी भी एएनआई के नाम से जोड़कर प्रदर्शित करता है।

दूसरी ओर ओपनएआई ने अदालत को बताया कि एआई मॉडल का प्रशिक्षण किसी कॉपीराइट सामग्री की प्रतिलिपि तैयार करने के समान नहीं है।

कंपनी का कहना था कि सार्वजनिक रूप से उपलब्ध जानकारी का मशीन लर्निंग के उद्देश्य से विश्लेषण करना उसी प्रकार है जैसे किसी पुस्तक को पढ़ना, जिसे उसके कॉपीराइट का व्यावसायिक उपयोग नहीं माना जा सकता।

ओपनएआई ने यह भी कहा कि उसके मॉडल का प्री-ट्रेनिंग चरण भारत के बाहर पूरा किया जाता है तथा प्रशिक्षण से संबंधित डेटा विदेशी सर्वरों पर सुरक्षित रखा जाता है।

कंपनी ने बताया कि समय के साथ उसके मॉडल में ऐसे सुधार किए गए हैं जिससे मूल सामग्री के पुनरुत्पादन की संभावना न्यूनतम हो गई है।

साथ ही प्रशिक्षण पूरा होने के बाद मॉडल के पास मूल प्रशिक्षण डेटा तक सीधी पहुंच नहीं रहती।

नवंबर 2024 की सुनवाई के दौरान ओपनएआई ने अदालत को यह भी बताया था कि उसने एएनआई के डोमेन को पहले ही अपनी “ब्लॉकलिस्ट” में शामिल कर लिया है।

इसका अर्थ है कि भविष्य में एएनआई की वेबसाइट की सामग्री का उपयोग उसके एआई मॉडल के प्रशिक्षण के लिए नहीं किया जाएगा।

इस मुकदमे में केवल एएनआई ही नहीं, बल्कि कई मीडिया संस्थान, प्रकाशन संगठनों और उद्योग निकायों ने भी पक्षकार के रूप में भाग लिया।

इनमें फेडरेशन ऑफ इंडियन पब्लिशर्स, डिजिटल न्यूज़ पब्लिशर्स एसोसिएशन और इंडियन म्यूज़िक इंडस्ट्री जैसे संगठन शामिल हैं, जिन्होंने एआई प्रशिक्षण में कॉपीराइट सामग्री के उपयोग को लेकर अपनी गंभीर चिंताएं न्यायालय के समक्ष रखी थीं।

Leave a Reply

Your email address will not be published. Required fields are marked *