Does Your Legal AI Tool Train on Client Data? What the Policies Actually Say

What CoCounsel, Everlaw, Relativity, Clio, Kira, and Lexis+ actually promise on data training and retention, plus the 2026 ruling on AI and privilege.

By Dana Whitfield18 min read

On February 17, 2026, a federal judge in the Southern District of New York became the first in the country to rule on a question every lawyer using AI has quietly wondered about. If you feed information from your legal matter into an AI chatbot, is what comes out still privileged? In United States v. Heppner, Judge Jed Rakoff said no. A criminal defendant, after retaining counsel, used the free consumer version of Anthropic's Claude on his own to research his case. He generated strategy memos based on what his lawyers had told him. When prosecutors sought those AI-generated documents, Rakoff ruled they were protected by neither attorney-client privilege nor the work-product doctrine. This was in part because Claude is not a lawyer, and in part because Claude's own consumer terms of service reserved the right to log the exchange, use it to train future models, and disclose it to regulators or other third parties. A tool with those terms, the court found, is not a place a person can reasonably expect privacy.

Heppner did not happen to a lawyer, but it is the clearest evidence yet that "what does this AI tool do with my input" has stopped being an abstract ethics-CLE question and become a live litigation fact pattern. Three weeks later, Nippon Life Insurance Company of America sued OpenAI in the Northern District of Illinois (filed March 4, 2026) after a claimant fed her own attorney's confidential correspondence into ChatGPT. She asked the tool whether she was being misled, and based on the answer, she fired her lawyer and moved to reopen a settled case. This was a different fact pattern, but it had the same root cause: privileged material was handed to a consumer AI product whose data practices the user did not understand. This is happening against a backdrop most lawyers have not fully absorbed. The AI tools built specifically for legal work generally publish very different data commitments than the free chatbots their clients and junior staff already use casually.

This piece does the vendor-by-vendor and bar-opinion-by-bar-opinion homework. It covers what separates a consumer AI product from an enterprise one, what the leading legal AI vendors actually commit to in their own published security and privacy documentation, what zero-retention and data-residency options exist, what state bars beyond the ones this site has already covered say about confidentiality duties under Rule 1.6, and a concrete checklist for verifying a vendor's claims before a single client document goes in.

Consumer AI and enterprise AI are not the same product, even with the same name

Many lawyers do not realize that "trains by default" and "does not train by default" are genuinely different products from the same technology provider. Using a personal account under consumer terms introduces a high AI confidentiality law firm risk. Using an enterprise account under business terms provides a much safer foundation for legal AI data confidentiality.

OpenAI data practices

By default, OpenAI does not use content submitted through ChatGPT Business, ChatGPT Enterprise, ChatGPT Edu, or the API platform to train or improve its models. This includes both user inputs and AI outputs, unless a customer explicitly chooses to opt in. This protection is outlined on the OpenAI Enterprise Privacy page. OpenAI can also execute a formal legal AI DPA (Data Processing Addendum) for ChatGPT Business, ChatGPT Enterprise, and API customers to support compliance with global privacy regulations.

Personal ChatGPT accounts are different. For Free, Go, Plus, and Pro accounts, training on a user's conversations is turned on by default. According to the OpenAI Data Usage FAQ, a user must actively turn training off in their account settings under Data Controls, or use a Temporary Chat, to prevent their conversations from being used to improve future models.

Anthropic data practices

By default, Anthropic does not use inputs or outputs from its commercial products to train its models. This policy applies to Claude for Work, the Anthropic API, and Claude Gov. Business and Enterprise customers receive Commercial or Enterprise Terms of Service that prohibit training without exception. This commitment is detailed in the Anthropic Privacy Center FAQ.

Consumer accounts are different. Claude Free, Pro, and Max subscriptions default to having model training turned on. According to the Anthropic Consumer Terms, a user can manually disable this setting. However, if left on, Anthropic may retain these conversations for up to five years to improve its models.

This is the exact distinction Judge Rakoff's Heppner opinion turned on. The defendant used a consumer product, not an enterprise deployment. The consumer terms of service were a major reason why the court found no reasonable expectation of privacy.

Practical translation

Saying "I use Claude" or "I use ChatGPT" is not a complete answer to an ethics inquiry. Which tier you use, and under what terms, matters as much as the company providing the tool.

A firm-wide policy should block personal or consumer AI accounts for any work involving client information. Steering staff toward a paid enterprise deployment with a signed DPA addresses the single biggest and most preventable version of this risk.

What the purpose-built legal AI vendors say about your data

Purpose-built legal software vendors generally offer much tighter controls. Here is what the leading platforms publish about their legal AI data privacy commitments. These are the vendors' own stated policies, not independent audits by this publication.

CoCounsel Legal (Thomson Reuters)

According to the Thomson Reuters CoCounsel Security FAQ, user content and prompts submitted to CoCounsel Legal "are not used to train or improve CoCounsel and associated products or LLMs." This data is not used in third-party output, and it is not stored by the underlying model providers like OpenAI or Google.

Thomson Reuters reports that its third-party model partners are contractually prohibited from using any customer data to train their models. The company also disables its third-party providers' abuse-monitoring tools. This ensures no human at those external providers can access customer data. This practice is detailed in the Thomson Reuters article, "The Two Non-Negotiables of Secure Legal AI."

Harvey

Harvey states on its security page that it contractually prohibits its model providers from training on customer data. It does not use inputs, outputs, or uploaded documents to train its underlying models.

According to its security policy, Harvey requires zero data retention from its model providers. Some customers choose a same-day retention setting so their data is deleted automatically after a prompt or output is produced. Harvey also advertises regional data residency choices, allowing firms to keep data within Europe, Switzerland, the United States, or Australia.

Lexis+ with Protégé

LexisNexis states in its security documentation that it "never uses customer data to train AI models" and that prompts and uploaded documents are "never used to train external AI models." This policy is detailed on the LexisNexis blog, "7 Key Facts about Legal AI Security and Privacy."

The company reports that client data is encrypted with AES-256 at rest and TLS 1.2+ in transit. It also follows the RELX Responsible AI Framework. This information is published in the Lexis+ AI Security Information document for law schools. If you want to compare how this tool operates alongside its main competitor, you can read our comparison of CoCounsel vs. Lexis+ with Protege for Legal Research (2026).

Everlaw

In the Everlaw AI Governance Framework, the vendor states that "none of Everlaw's LLMs will train or fine-tune their models or services on customer data." None of the large language models it uses retain that data once processing is complete.

Everlaw is SOC 2 Type 2 certified across security, availability, confidentiality, and privacy. It also maintains an ISO 27001-aligned information security management system. These certifications are documented on the Everlaw Trust Center. For firms evaluating other options in this space, see our guide on Everlaw Alternatives for Law Firms (2026).

RelativityOne / Relativity aiR

According to a Microsoft customer story on Relativity's deployment, Relativity's infrastructure partner (Microsoft, via Azure) is contractually required to never train models on customer data. Any customer data used for aiR's analysis is not retained by either Relativity or Microsoft once processing is complete.

RelativityOne is built on Microsoft Azure and meets more than 90 international and industry compliance standards, including FedRAMP. Relativity also offers Relativity Server as an on-premise deployment distinct from the RelativityOne cloud product. This is detailed on the Relativity compliance and privacy page.

Clio Duo / Manage AI

Clio states in its data privacy resources that data processed through Clio Manage AI is "never used to train external AI models." Manage AI strictly follows a firm's existing Clio Manage permissions. It will not surface documents or matters that a given user is not authorized to see.

Clio notes that depending on a firm's location, a query may be processed on servers outside the firm's home jurisdiction. However, the resulting output is stored back within the firm's region under Clio's standard data-handling practices. These policies are detailed in Clio's "AI Data Privacy for Law Firms" guide. To see how this compares to other practice tools, check out Clio Manage AI vs. MyCase IQ (2026).

Kira

Kira states in its vendor documentation that it does not use customer data for model training. Organizations can train their own custom extraction models on their own data without that data leaving their secure environment. Kira offers data residency options across the US, Canada, Europe, and Asia-Pacific. If you are reviewing other systems for transactional work, you can read our breakdown of Spellbook vs. Kira for AI Contract Work (2026).

Not every vendor in our 43-tool tracker publishes a specific, citable training or retention statement. The absence of a clear statement is something a buyer should ask about directly before sending any client data.

Zero retention, data residency, and on-prem: what to actually ask for

To secure your client documents, you must know what technical protections to request from a vendor. Here are the core terms to bring to the negotiating table.

Zero data retention legal AI

Under a standard contract, an AI vendor may store your queries and documents for a set period to monitor for system abuse. A zero data retention configuration means the vendor and its underlying model providers delete your data immediately after the prompt and output cycle is complete. Harvey offers this as a configurable setting. You should ask any legal AI vendor whether a zero-retention or short-retention configuration is available and if it requires an extra fee.

Data residency

Firms with cross-border clients or strict regulatory environments must know where their data resides. Kira offers residency across four regions: the US, Canada, Europe, and Asia-Pacific. Clio notes that temporary processing might happen on servers outside your home country, even if the final stored data returns to your region. Do not accept a general verbal claim that a vendor is global. Get their regional storage and processing commitments in writing.

On-premise and private infrastructure

Most modern legal AI tools are cloud-only. Relativity is a notable exception, offering Relativity Server as a fully separate on-premise deployment. This is ideal for firms or matters that cannot send data to any third-party cloud. If you use cloud tools, look for enterprise-grade security certifications like SOC 2 Type II to verify their data protection standards.

Subprocessors and model providers

A legal AI tool is often an interface built on top of external models like OpenAI, Google Gemini, or Anthropic Claude. Even if your primary vendor promises not to train on your data, their model providers might. You must ask whether the underlying subprocessors are contractually bound by the exact same no-training and retention commitments as the primary vendor.

What state bars actually require on confidentiality: beyond the baseline

State bars are moving past general warnings. They are issuing highly specific requirements for lawyers handling privileged documents AI tools.

We have already detailed the high-level Rule 1.6 framework in our guide on what the bar actually requires when you use AI. That piece covers ABA Formal Opinion 512, which requires lawyers to get informed consent before inputting client data into generative AI tools. It also covers guidelines from Florida, California, and New York City.

Several other states have since issued direct, detailed guidance.

Texas

The State Bar of Texas Professional Ethics Committee released Opinion 705 in February 2025. Developed by the Taskforce for Responsible AI in the Law, Texas Ethics Opinion 705 states that a lawyer "must not knowingly reveal client confidential information to any person other than those who are permitted to receive the information."

The committee clarified that this duty extends to both privileged information and all other information relating to a client, whether or not it is separately privileged. The opinion warns that AI tools often require detailed inputs, which creates a high risk of inadvertent disclosure. This is especially true for systems that share information with third parties. It directs lawyers not to input confidential or sensitive client data into public or unvetted AI platforms like consumer chatbots.

District of Columbia

The D.C. Bar published Ethics Opinion 388 in October 2024. This opinion frames the confidentiality analysis around two concrete questions. First, will the client's information be visible to third parties? Second, will the interaction affect the answers the tool gives to future users?

The opinion is direct about consumer tools. It warns that many free AI products "collect and use inputs for training and future improvement." It notes that their privacy policies effectively treat user inputs as an asset "to be exploited and sold to others."

Pennsylvania and Philadelphia

The Pennsylvania Bar Association and Philadelphia Bar Association Joint Formal Opinion 2024-200, issued in May 2024, is highly direct. It states that "a lawyer must not input any confidential information of a client into AI that lacks adequate confidentiality and security protections." While this is an advisory opinion that is not legally binding on the state's Disciplinary Board, it shows where regulators are aligning.

California's proposed rule change

The State Bar of California Committee on Professional Responsibility and Conduct (COPRAC) went a step further. At its March 13, 2026 meeting, COPRAC approved proposed amendments to Rule 1.6 itself.

The proposed text defines "reveal" to include:

"exposing confidential information to technological systems, including artificial intelligence tools, where such exposure creates a material risk that the information may be accessed, retained, or used... in a manner inconsistent with the lawyer's duty of confidentiality"

Under this proposed rule, simply inputting client information into an AI tool with weak data-handling terms could constitute an ethical violation. This would be true even if no actual leak or public disclosure occurs, as long as there is a material risk of misuse by the system or another user. Public comment on this package closed on May 4, 2026. It remains a proposal and is not yet adopted, but it represents the most aggressive regulatory approach to date.

The trend is clear. You must know exactly how a tool handles your input before you use it on a client matter. Treating "I didn't know" as a defense is no longer viable.

The Heppner ruling and why it matters beyond one criminal case

The Southern District of New York's decision in United States v. Heppner provides a clear warning about using unsecured AI tools. It is the nation's first concrete judicial ruling on AI and attorney-client privilege.

The facts of the case

After being subpoenaed and retaining counsel, the defendant, Heppner, used the free consumer version of Anthropic's Claude. He did this on his own, without his attorneys' direction or knowledge. He used the tool to research legal issues connected to the government's investigation. He also input information he had learned from his lawyers, which produced documents outlining his defense strategy.

The government moved to compel these AI-generated documents on February 6, 2026, arguing they were not privileged. The court heard arguments on February 10, and Judge Jed Rakoff ruled for the government on February 17, 2026.

The court's reasoning on privilege

Judge Rakoff ruled that Claude is not an attorney. The court stated this fact alone disposes of any attorney-client privilege claim over the exchange.

The court then examined Claude's consumer terms of service and privacy policy. Because these terms reserved Anthropic's right to log prompts, use them for model training, and disclose them to third parties, the court found they were incompatible with any reasonable expectation of privacy. You can read a detailed legal analysis of these terms in the Proskauer client alert.

The court's reasoning on work product

The defendant argued the documents were prepared in anticipation of litigation. However, because they were not prepared by or at the direction of counsel, and did not reflect the lawyers' own strategy, the court ruled the work-product doctrine did not apply. This reasoning is examined in the Harvard Law Review Blog.

The narrow exception

The court suggested the outcome might have differed if counsel had directed the client to use the AI tool as part of the representation. This could potentially bring the tool under privilege using the Kovel doctrine, which covers a lawyer's use of third-party experts or translators. The court did not rule that attorney-directed AI use is definitely privileged, but it explicitly distinguished that scenario from Heppner's unsupervised use.

This case converts an abstract ethics warning into a binding judicial holding. If you or your clients use consumer tools with standard training and logging terms, a court can compel you to hand those conversations over to your adversaries.

The Nippon Life case: what happens when a client, not a lawyer, uses consumer AI

The risk to client confidentiality is not limited to what happens inside your law firm. Your clients can compromise their own privileges by using consumer AI tools without your knowledge.

On March 4, 2026, Nippon Life Insurance Company of America filed a lawsuit against OpenAI in the Northern District of Illinois (Case No. 1:26-cv-02448). The lawsuit arose after a claimant in a settled dispute uploaded her own attorney's correspondence to ChatGPT. She asked the chatbot whether she was being misled by her counsel and Nippon Life. ChatGPT answered that she was, and drafted motions for her. The claimant then fired her lawyer and moved to reopen the settled case.

Nippon Life's lawsuit is directed at OpenAI, not the claimant or her former attorney. The lawsuit alleges tortious interference with contract, abuse of process, and the unauthorized practice of law. It seeks roughly $300,000 in compensatory damages and $10 million in punitive damages. This case is analyzed in detail by Stanford Law School's CodeX.

This is not a case about a legal AI vendor suffering a data breach. Instead, it is a clear example of the dangers warned about in Texas Opinion 705 and D.C. Opinion 388. When a client feeds confidential attorney correspondence into a consumer AI tool, they do so with no legal protection. The results can disrupt settlements and lead to severe litigation fallout.

A checklist before you input a single client document

Before you or your staff upload any client files into an AI system, use this checklist to verify the tool's security posture.

  • Verify the product tier: Ensure you are using a paid, enterprise-grade deployment with a signed agreement. Do not assume that because your firm has an enterprise license, using the free or personal version of the same tool carries those protections.
  • Get a written commitment against model training: Obtain a written, contractual guarantee that the vendor will not train its models on your inputs or outputs. Do not rely on marketing copy or verbal assurances. Purpose-built platforms like CoCounsel, Everlaw, RelativityOne, Clio Manage AI, and Kira provide this in their terms.
  • Request a DPA and SOC 2 Type II report: Ask for a signed Data Processing Agreement and a copy of the vendor's latest SOC 2 Type II audit. For a detailed guide on how to evaluate tools during a trial run, see our resources on Why Almost No Legal AI Tool Offers a Free Trial and Why Almost No Legal AI Tool Has Reviews.
  • Confirm data residency and retention windows: Ask where your data is processed and stored. Inquire whether a zero-retention or same-day deletion configuration is available for sensitive matters.
  • Audit the subprocessors: Confirm that any underlying model providers, such as OpenAI, Google, or Anthropic, are contractually bound by the same no-training and data deletion terms as your primary vendor.
  • Obtain informed client consent: If you are using a tool that does not meet strict enterprise standards, or if your local bar requires it, obtain explicit client consent before uploading their files. General boilerplate language in an engagement letter does not satisfy this requirement under ABA Formal Opinion 512.
  • Establish a firm-wide AI policy: Ban the use of personal, free, or consumer-grade AI accounts for any work that touches client information. This is necessary because both OpenAI and Anthropic train on consumer conversations by default.

FAQ

Do legal AI tools train on client data by default?

It depends entirely on the product tier you use. Purpose-built legal AI tools like CoCounsel, Harvey, Lexis+ with Protégé, Everlaw, RelativityOne, Clio Manage AI, and Kira publish clear policies stating they do not train models on customer data by default. Consumer AI platforms are the opposite. Personal accounts on ChatGPT and Claude have model training turned on by default, and users must manually opt out in their settings to protect their data.

Is information I put into an AI tool still covered by attorney-client privilege?

No, not automatically. In United States v. Heppner, a federal court ruled that documents generated by a defendant using a free consumer version of Claude were not protected by attorney-client privilege or the work-product doctrine. The court found that Claude's consumer terms of service, which allowed data logging and model training, were incompatible with a reasonable expectation of privacy. The court left open the possibility that AI use directed by counsel as part of a legal representation might be treated differently.

What is the difference between a consumer AI account and an enterprise one?

Consumer accounts are built for the general public and typically have model training turned on by default. They also carry terms of service that allow the provider to log, review, and share data. Enterprise, business, and API tiers contractually commit to not training on your data, offer custom retention settings, and will execute a formal Data Processing Agreement (DPA) to ensure regulatory compliance.

What should I ask a legal AI vendor before uploading client documents?

You should ask for a written contractual commitment that the vendor will not train its models on your data. You should also request their SOC 2 Type II security report, a signed Data Processing Agreement, and details about their data residency and retention policies. Finally, ask whether any underlying model providers are bound by the same no-training and security terms.

Do state bars require lawyers to vet an AI tool's data practices before using it?

Yes, state bars are increasingly requiring this step. Texas Ethics Opinion 705 directs lawyers not to input sensitive client data into public or unvetted AI platforms. D.C. Bar Ethics Opinion 388 requires lawyers to verify whether client data will be visible to third parties. Pennsylvania and Philadelphia Joint Formal Opinion 2024-200 states that lawyers must not input confidential client data into any AI that lacks adequate security. Additionally, the State Bar of California has proposed amending Rule 1.6 to make exposing client data to tools with a material risk of misuse an ethical violation.

The bottom line

The confidentiality risk in legal AI is not evenly distributed. The software vendors built specifically for legal work generally publish clear, contractually binding commitments not to train on customer data. The risk concentrates almost entirely in consumer-grade AI products used casually by lawyers, support staff, or clients themselves outside of an enterprise agreement. This is the exact pattern that triggered the litigation in both Heppner and Nippon Life.

The year 2026 is when AI confidentiality stopped being an abstract ethics warning and became active case law. The SDNY's decision in Heppner is the first ruling to strip privilege from AI-assisted work product, turning directly on the consumer terms of service the user accepted. Meanwhile, California's proposed Rule 1.6 amendment shows that regulators are prepared to make the mere use of unvetted AI tools a disciplinary offense.

The practical response for law firms is straightforward. Know exactly which product tier your firm is using, get no-training commitments in writing, and establish clear policies so that a client's unsupervised AI use does not undo your own hard work.