AI, Copyright, and Research: A Guide

Columbia University Libraries provides its research community with access to millions of copyrighted and licensed resources. While AI tools offer novel ways to analyze information, their use must align with U.S. copyright law, university data policies, and Columbia’s specific licensing agreements with publishers. 

The content on this page is for informational purposes and is not legal advice. Key contacts for different topics mentioned in this guide are listed at the bottom of this page.

Different AI tools may be used for different purposes. CUIT’s Data Classification Table should be referred to if there is any question about what tool is appropriate to use with sensitive, confidential, internal, or public information.

CUIT makes available several AI services within secure, education-focused environments. The current and upcoming tools are detailed CUIT's AI Services page, while CUIT's Data Classification Table details which tools are appropriate for which purposes. CUIMC classifies several CUIT-managed tools as being HIPAA-compliant. CUIMC users also have access to HIPAA-compliant Microsoft Copilot. Instructions for Copilot access are detailed on the AI and Generative Technology Use at CUIMC page.

Best Practices for Responsible AI Use focus on actions, rather than on tools. Notably, CUIT advises against entering the following into public or unapproved AI systems:

  • Personally identifiable information for students, faculty, or staff (including email addresses)
  • Health or clinical data
  • Financial, legal, or confidential institutional records
  • Unpublished research or proprietary materials

In general, personal accounts - paid and free - with large commercial companies do not prioritize user privacy - at least not from the company itself. An example of this is how Claude or ChatGPT might remember a user’s past chats and bring up topics from other threads in newer chats. User queries and the resulting responses are tagged, cataloged, and retrievable for future use, including the training of future AI models.

This depends on the terms outlined by the creator(s) of the dataset and/or the terms of the repository or publisher making the dataset available. Roper Center Artificial Intelligence (AI) Policy​​​​​​​ and ICPSR Policy on the Use of Large Language Models are examples of such terms. Generative Artificial Intelligence and Open Data: Guidelines and Best Practices from the US Department of Commerce is another excellent resource. When working with a dataset, locate either the License, or Terms of Use. If neither are available, contact the affiliated creator(s) or organization(s).

You should exercise extreme caution when working with works that are within copyright and not openly-licensed. Refer to CUIT’s Data Classification Table if you have questions about what tool is appropriate to use with sensitive, confidential, internal, or public information. Personal AI tools not licensed by Columbia University do not have any protections for users and their data, and as safe best practice in-copyrighted texts (including unpublished work) should never be uploaded into these tools.

If using a personal AI tool that has not been licensed by Columbia University, the following types of works are generally appropriate for upload based on licensing and/or copyright status:

Sensitive data (as defined by the University’s Data Classification Policy) may only be used with AI tools specifically vetted and approved by Columbia. 

Uploading a document to summarize it or create a custom chatbot (like a Gemini Gem) uses Retrieval-Augmented Generation (RAG) rather than foundational model "training." However, importantly, this still involves copying and processing protected intellectual property on a vendor's server. While this may not be considered “training,” and therefore not in violation of licensing agreements in that specific way, it may still be a violation of copyright.

Private AIs (LLMs run locally on a user’s computer) process data entirely on your local machine, meaning no data is transmitted to a cloud vendor. This eliminates the risk of third-party data retention and can be a useful strategy for keeping confidential information strictly local. However, derivative works created from copyrighted materials may still constitute infringement. 

Stanford’s Foundation Model Transparency Index provides rankings for open models released by businesses. WhatLLM.org helps users compare 100+ large language models across price, performance, speed, and quality using the Artificial Analysis Intelligence Index. Olmo from Ai2 is an example of an open model released by a non-profit research organization, the Allen Institute.

Text and Data Mining (TDM) is the process of deriving information from (often large amounts of) text through the use of computers.

Beyond these resources, researchers should consider whether their TDM activities can be supported by a Fair Use argument. The chapter Copyright from the 2021 book Building Legal Literacies for Text Data Mining provides a good starting point. More information about the publication can be found on the Authors Alliance announcement. In 2015, the Association of Research Libraries argued that “As long as the researcher is not bound by a contract that forfeits her fair use rights, she may proceed with TDM so long as her results do not make the full text, or substantial portions, of the underlying articles publicly available.”

A key thing to remember here is that the agreements between scholarly publishers and Columbia Libraries may explicitly prohibit TDM activities, and so it is best to ask for more information about these agreements if there is any question.

MIT Libraries provides an excellent guide, Citing AI, that gives users practical information about how to track and disclose their use of AI tools. 

The AI declaration statement template authored by the Open Library of Humanities (OLH) provides a nuanced framework within which authors can explain their use of generative AI. 

The Artificial Intelligence Disclosure (AID) – Statement Builder offers a granular template for citing AI tool usage by specific task.

For a more utilitarian approach to citing AI tools, see these resources related to specific style guides:

More information about citation and reference management at Columbia Libraries can be found on the Reference & Citation Management page.

Websites and tools, including those provided by CUIT, often have terms and conditions. These can include granting the owner of a website or tool copyright ownership - or a license to use - the work that users upload. The Cornell Lab of Ornithology Terms of Use can be referred to as an example.