[By Vashmath Potluri & Shubhranshu]
The authors are students of NALSAR, Hyderabad.
Introduction
In December 2024, Asian News International (ANI) filed a copyright infringement suit against OpenAI, alleging that its large language models (“LLMs”) had reproduced ANI’s news content without consent. While the case is widely perceived as a test of India’s copyright regime, it reveals a deeper and more systemic competition law issue: OpenAI, backed by Microsoft’s infrastructural and financial resources, enjoys exclusive control over high-quality training data through Reddit and Stack Overflow, and privileged integration into Microsoft’s software and cloud ecosystems as its models power AI features across Microsoft 365 Copilot and GitHub Copilot via the Azure OpenAI Service.. In 2024, over 60% of professionals in India reported using tools like ChatGPT and Microsoft Copilot, signalling the rapid mainstreaming of generative AI. Yet Indian AI startups attracted only $92 million in funding that year, compared to $13 billion in the United States, an imbalance that highlights the structural disadvantage domestic firms face in accessing essential AI inputs like data, computing, and distribution.
This article uses the ANI v. OpenAI dispute as a lens to introduce a competition law perspective that has been largely overlooked. It argues that OpenAI’s conduct may amount to abuse of dominance under Section 4 of the Competition Act, 2002 (“The Act”), and calls for a suo motu investigation by the Competition Commission of India (“CCI”). This argument proceeds in two parts. First, it defines the relevant market as upstream and downstream, comprising access to training data and development of Application Programming Interface (“API”) which are software tools that enable developers to integrate AI models into their products and services, to establish a denial of market access and vertical leveraging which refers to the use of dominance in one market, such as access to data, to gain an unfair advantage in another, such as enterprise-facing AI services. Secondly, it highlights doctrinal gaps in the Act to deal with non-pricing exclusionary conduct, which conduct restricts rival participation not through higher prices, but through control over essential inputs, technical lock-ins, and bundling arrangements. Accordingly, it proposes a two-pronged reform drawing inspiration from the EU, UK, and USA to strengthen the ex-ante framework of India’s competition regime in the era of AI.
Delineating the Relevant Market: A Layered Approach
A scrutiny of potential abuse of dominance by the CCI against OpenAI should begin with defining the relevant market under Sections 2(r), 2(s), and 2(t) r/w 19(7) of the Act. These sections together provide for traditional factors like substitutability, price sensitivity, and consumer choices, but they become inadequate in the context of generative AI, where market power is determined by access to high-quality training data and compute infrastructure.
To address this, the article adopts a layered market structure: an upstream market for access to training data and a downstream market for API’s. This approach finds support in Shamsher Kataria v. Honda Siel Cars India Ltd., where the CCI held that the aftermarket for spare parts and services was distinct from the primary car market due to structural lock-in, limited alternatives, and information asymmetry. Though linked technologically, the markets were considered economically independent by the CCI.
This reasoning directly applies here. In Upstream, access to high-value training data is scarce and non-replicable, granting early movers like OpenAI a durable edge. In Downstream, developers integrating proprietary models via APIs face high switching costs, limited interoperability, and opaque performance metrics. These conditions result in technical and contractual lock-ins, which create sustained reliance on the dominant provider’s ecosystem. Together, these combined features justify treating the training and deployment layers as separate, yet interrelated, markets.
Upstream Market: Denial of Access to Training Data
Section 4(2)(c) of the Competition Act, 2002 prohibits a dominant company from doing anything that leads to the denial of market access “in any manner.” In Umar Javeed, Sukarma Thapar, Aaqib Javeed v. Google LLC & Ors, this phrase was interpreted broadly to hold that unfair conditions, creation of technical barriers, or lack of transparency constitute denial of access. This applies directly to OpenAI’s conduct in the upstream market of generative AI, which involves access to large, high-quality datasets like news articles, coding forums, and online discussions. These data sources are essential for training LLM’s. However, the real advantage lies not in general access to data, but in control over high-quality, non-replicable datasets that significantly improve model performance.
OpenAI’s exclusive deals with platforms like Reddit and Stack Overflow give it early access to high-quality conversational data crucial for fine-tuning large language models. Indian developers, however, face major hurdles: India’s copyright law lacks a text and data mining (TDM) exception, annotated datasets in local languages are scarce, and compute access remains prohibitively expensive. While there’s no formal refusal of access, these combined legal and infrastructural barriers make it commercially unviable for domestic firms to compete. This amounts to a constructive denial of market access, which constitutes an abuse under Section 4(2)(c) of the Act where exclusion occurs not through outright refusal but through systemic disadvantage.
This pattern of exclusion is reinforced by OpenAI’s own GPT-4 Technical Report, which acknowledges the use of a mix of public and licensed data, highlighting the importance of access to curated datasets. Similarly, the UK Competition and Markets Authority (CMA) has cautioned that exclusive control over high-quality, non-public datasets can give certain firms an undue competitive advantage and restrict market competition. In its recent assessments of foundation models, the CMA has treated such data as a core input and highlighted the risk of market foreclosure resulting from closed data ecosystems. This framework provides a valuable basis for the CCI in applying Section 19(4) of the Act, which focuses on factors such as control over key inputs, barriers to entry, and the ability to operate independently of competitive constraints. Recognising data as infrastructure allows the CCI to identify exclusionary conduct even in markets where price or output manipulation is absent. OpenAI reflects this structure: it controls critical training data, is deeply integrated into Microsoft’s infrastructure, and faces no effective domestic competitors. Its market power arises not only from early innovation but from its capacity to prevent others from entering and scaling within the market. This raises a clear case for regulatory scrutiny under Indian competition law.
Downstream Market: Leveraging of Data Dominance
Section 4(2)(e) of the Act prohibits a dominant enterprise from using its position in “one relevant market to enter or protect another.” This provision targets vertical leveraging, where dominance in the upstream segment is extended to the downstream market in a way that distorts competition. In Google Android, the CCI held that Google’s dominance in the mobile OS market was unlawfully used to impose pre-installation conditions on device manufacturers, thereby restricting rival app developers and limiting consumer choice. Therefore, the test is to determine (1) dominance in the primary market; and (2) its strategic use to exclude competition in the secondary market.
Applying this to OpenAI, its upstream dominance in foundational model development, which refers to the process of training large-scale, general-purpose AI models on vast and diverse datasets to enable their adaptation across multiple applications, derived from exclusive access to high-quality training data and compute infrastructure, is being used to consolidate control in downstream application markets which includes AI-powered productivity tools, coding assistants, and enterprise APIs, all of which are closely integrated into Microsoft’s software ecosystem through Microsoft 365, GitHub Copilot, and Azure OpenAI Services. Microsoft’s strategic investment of $3 billion in India’s AI and cloud sectors, alongside extensive partnerships across public and private domains, further amplifies OpenAI’s reach.
OpenAI’s GPT models are not offered as modular or interoperable services. Instead, they are tightly bundled into enterprise platforms, which limits substitutability and gives users little flexibility in choosing alternative providers. This bundling creates what are known as switching costs, which are barriers that make it difficult and costly for users to move to competing services. In enterprise settings, switching costs arise from the need to rebuild software integrations, retrain staff, adapt internal workflows, and potentially lose access to widely used Microsoft tools and ecosystems. For Indian content creators like ANI, this creates a double exclusion: not only is their content mined without compensation, but the resulting AI tools now embedded into Microsoft’s productivity suite are monetised through closed enterprise channels to which they have no meaningful access. The CCI’s ruling in Matrimony v. Google reinforces this concern by recognising that persistent user lock-in, opaque backend processes, and lack of viable alternatives can entrench dominance and deter market responsiveness.
In OpenAI’s case, the lock-in is reinforced by high integration costs, closed APIs, and a lack of transparency around pricing and model performance. These design choices limit user mobility and suppress innovation by raising artificial barriers for rivals. Crucially, the downstream entrenchment strengthens OpenAI’s upstream dominance. As more users interact with GPT-powered tools via Microsoft’s platforms, OpenAI gains access to valuable behavioural data used to further optimise its models. This creates a feedback loop where more usage leads to better models, which in turn makes switching less attractive, reinforcing OpenAI’s dominance at both layers.
Such conduct falls within the scope of Section 4(2)(e) of the Act, which prohibits a dominant enterprise from using its position in one relevant market to protect or enter another. By bundling services, restricting access through exclusive arrangements, and creating technical barriers, OpenAI limits opportunities for Indian developers, reduces consumer choice, and stifles innovation. To prevent long-term structural harm to the Indian AI ecosystem, the CCI must examine these practices and consider targeted regulatory measures to preserve openness, interoperability, and fair competition in the generative AI space.
Way Forward: Building a Data-Aware Competition Law Framework
Despite a 3.6-fold increase in India’s generative AI startup base from over 66 in the first half of 2023 to more than 240 in the first half of 2024 funding has only seen a modest 1.25x rise, largely concentrated in early-stage rounds. As a result, Indian companies continue to face challenges in accessing critical AI infrastructure due to limited financial opportunities and the lack of access to high-quality datasets that make it difficult for them to compete with global giants like OpenAI, which benefit from early and exclusive access to such resources. To address these systemic barriers and ensure that India’s AI ecosystem remains open and innovation-friendly, this article proposes a two-pronged reform agenda.
1. Codifying the Essential Facilities Doctrine
Regulatory authorities are beginning to recognise that having access to high-quality data and using it behaviourally has turned into an essential market advantage for a business. This move is clear from Article 6(12) of European Union Digital Markets Act (DMA) which requires “gatekeepers,, meaning large digital platforms that occupy a central position in the online ecosystem and act as critical intermediaries between users and businesses, to share key datasets on Fair, Reasonable, and Non-Discriminatory (FRAND) terms. This highlights a growing understanding that in AI-driven digital markets, exclusion typically arises not from pricing strategies, but from the strategic control of critical data inputs. However, CCI continues to rely on traditional price and product-based approaches, leaving it ill-equipped to tackle exclusion rooted in data dominance. As the generative AI sector shows, exclusive access to non-substitutable training datasets acts as a formidable entry barrier.
To fill this gap, the Act must be amended to reflect the infrastructural nature of data. First, Section 19(4) should be amended to treat exclusive control over essential datasets as an indicator of dominance. Alongside, Section 4(2) should prohibit unjustified refusals to license such data where it leads to foreclosure of competition. This amendment must be operationalised through the Essential Facilities Doctrine, which determines when private control over a key input constitutes structural exclusion. Applied in Microsoft v. European Commission., the Commission found Microsoft’s refusal to provide interoperability data to rival server developers unlawfully foreclosed downstream markets. Generative AI presents a similar risk: high-quality training datasets are non-substitutable inputs, and control over them allows incumbents to lock out challengers.
Accordingly, the test must must apply four conditions to determine whether such a refusal constitutes abuse of dominance: (i) the input must be indispensable for competition; (ii) replication must be unfeasible in practice; (iii) access must not undermine the dominant firm’s legitimate operations; and (iv) the refusal must lack objective justification. These criteria offer a clear standard to assess when data withholding becomes exclusionary.
2. Embedding Interoperability and Open Access
In the EU and UK, regulators now require dominant tech firms to ensure interoperability, allowing users and developers to switch or integrate across platforms, thereby enhancing contestability. India, however, lacks such safeguards. This enables ecosystem lock-in, as seen in OpenAI’s bundling of its models with Microsoft Azure, which restricts rivals’ access and raises switching costs for developers.
To address this, India should amend the Act to empower the CCI to prevent interoperability restrictions. A new clause under Section 4(2) could prohibit dominant firms from using technical or contractual means to tie AI models to specific cloud platforms. Correspondingly, Section 27 should also be expanded to allow structural and behavioural remedies such as mandating open APIs or prohibiting exclusive bundling where such practices foreclose competition. Alongside this, India must legislate open access to public sector datasets such as court rulings, census data, and health records should be made available in machine-readable formats under standard licences. Following the EU’s Open Data Directive (Directive EU 2019/1024), such reform would equip Indian developers with critical training inputs and reduce dependence on proprietary datasets. Together, interoperability mandates and open data access are essential to dismantle structural barriers and ensure a more inclusive, innovation-friendly AI ecosystem in India.
