Where legal AI research tools still have jurisdiction gaps

Legal AI research tools claim national coverage, but hallucination data and vendor reviews show real gaps by state and court level. Here is what to check.

By Caleb Mercer10 min read

Legal AI vendors have built their marketing around a highly compelling promise. They claim their systems search the full body of case law across all fifty states. For litigators and partners at multi-state firms, this sounds like a complete solution. You can type a natural language prompt and receive a clean, cited summary of the law in any US jurisdiction.

However, there is a major gap between having a document indexed in a database and being able to research it accurately with AI. Independent evaluations show that the depth and reliability of these tools are highly uneven. The underlying technology works well when retrieving high-profile federal appellate decisions. It struggles when navigating the more fragmented worlds of state trial courts and administrative agencies.

To evaluate these systems, buyers must look beyond vendor marketing. This guide examines where general-purpose legal AI platforms encounter jurisdiction gaps. We will look at what independent research says about error rates, why specific practice areas are more vulnerable, and how to protect your firm before signing an expensive contract. To understand the broader performance of these tools, you can also read our analysis of How accurate is legal AI, really? What the benchmarks show in 2026.

The coverage claim versus the accuracy record

The most rigorous independent study of legal AI research tools to date highlights the contrast between vendor claims and actual performance. Researchers at Stanford University's RegLab published a preprint study in May 2024, which was later peer-reviewed and published in the Journal of Empirical Legal Studies in 2025. You can view the original publication on the Stanford RegLab portal and access the peer-reviewed version through the Journal of Empirical Legal Studies.

The researchers tested the leading legal AI tools and found meaningful hallucination rates. LexisNexis's Lexis+ with Protege, which was evaluated under its original name Lexis+ AI, hallucinated in roughly 17 percent of queries. Thomson Reuters's Westlaw AI-Assisted Research, a predecessor product to the AI features now bundled into CoCounsel, hallucinated in roughly 33 percent of queries. For comparison, a bare GPT-4 baseline with no legal retrieval layer hallucinated in 43 percent of queries.

These findings directly contradicted the vendors' marketing. LexisNexis had claimed 100 percent hallucination-free linked legal citations. Thomson Reuters had claimed its tools avoid hallucinations by relying on trusted content. The Stanford researchers concluded that "providers' claims are overstated."

Both vendors have updated and rebranded their products since this study was conducted. Lexis+ AI became Lexis+ with Protege in February 2026. CoCounsel is now deeply integrated into the Westlaw ecosystem. However, as of mid-2026, neither vendor has published a comparable, peer-reviewed accuracy study to disprove these baselines. Buyers should treat these historical error rates as the most reliable independent data points available. Unchecked hallucinations can lead directly to court-ordered penalties, as detailed in our guide on how Courts are sanctioning lawyers for AI-hallucinated citations.

Why the gaps cluster in state trial courts and administrative law

The errors identified in legal AI benchmarking are not distributed evenly across all types of law. Commentary from legal research librarians, including an LLRX summary of what the science says about hallucinations, points to a clear structural pattern. You can read the analysis on LLRX and follow the discussion from the AI Law Librarians.

The research shows that large language models perform best on United States Supreme Court decisions and major federal appellate cases. They perform worst on lower-court opinions, state trial court rulings, and administrative-law decisions.

The reason for this gap comes down to data density. Large language models require massive amounts of training data to learn how concepts relate to one another. The published opinions of federal appellate courts are highly structured, widely distributed, and heavily cited. They form a thick web of data that AI models can easily map.

In contrast, state trial court opinions and administrative-law rulings are thinly represented in training datasets. Many local jurisdictions do not publish comprehensive, digital collections of trial-level orders. Administrative bodies often maintain separate, poorly indexed databases. Because the AI model has less high-quality data to learn from, its retrieval mechanisms are much less reliable in these areas.

This structural gap creates a silent risk for specific practices. A firm handling routine state-level personal injury cases, family law disputes, or local zoning matters operates in the exact areas where AI tools are most likely to hallucinate. The vendor's marketing materials will state that they cover your state, but they will not warn you that their retrieval accuracy drops off when you leave the federal system.

The multi-state comparison problem

Another significant limitation involves multi-state legal synthesis. Many general-practice and mid-sized firms need to research how a single legal issue is treated across different jurisdictions.

A hands-on review of Westlaw Precision, which contains the AI research layer for the CoCounsel and Westlaw bundle, evaluated this capability. You can read the detailed review on VaQuill. The reviewer found that the AI system performed noticeably worse when asked to draft a memo comparing an issue across three different states side-by-side than when handling a single-jurisdiction question. The model struggled to keep the distinct state standards separate and occasionally blended different statutory elements together.

This limitation is well known to the developers of these platforms. Both LexisNexis and Thomson Reuters scheduled major model updates in mid-2026 specifically designed to address these multi-jurisdiction synthesis failures.

For buyers, this is a critical detail. If your practice requires you to compare the laws of multiple states, you cannot assume a generalist AI can do this reliably. The fact that major vendors are actively rebuilding their tools to fix this issue proves that cross-jurisdiction comparison remains a live technical challenge.

When narrow beats broad: the CompFox model

To understand the alternative to a generalist tool, look at CompFox. This platform does not attempt to cover the laws of all fifty states. It does not index federal appellate tax decisions or Delaware corporate disputes.

Instead, CompFox is built exclusively for California workers' compensation law. It is trained on a highly specific, proprietary database of more than 34,000 California Workers' Compensation Appeals Board and panel decisions. The developers spent two years building and fine-tuning this single-state, single-practice dataset. To see how these tools fit into a practice, you can read our Legal AI for Workers' Compensation Firms: A Buyer's Guide.

This narrow approach offers a major advantage in accuracy. Because the tool's dataset is restricted to one jurisdiction and one practice area, the retrieval layer is highly optimized. The system understands specialized local terms, local procedural rules, and specific state board habits better than a generalist model ever could.

For firms whose practices are highly concentrated, the lesson is clear. A broad "nationwide" marketing claim is often a signal of shallow local coverage. A tool that is honest about its narrow limits and deep within its single lane is often the safer, more reliable choice.

The hidden cost of uneven coverage

When a legal AI tool fails to return accurate citations in a specific jurisdiction, the cost is measured in attorney hours. AI research software is sold on the promise of efficiency. If a tool can summarize a legal issue in ten minutes instead of two hours, it saves the firm money.

However, if you are practicing in a lower-court or state-level jurisdiction where the tool's accuracy drops, you cannot trust the output. You must verify every single citation, locate the original case PDF, and confirm that the holding supports your argument.

This verification process takes time. If a tool has a 33 percent error rate in your practice area, the time spent double-checking hallucinations can completely erase the efficiency gains. You may end up spending more total time verifying untrustworthy AI output than you would have spent conducting traditional, manual research using a standard database.

Furthermore, relying on inaccurate outputs creates severe malpractice risks. General-practice partners who rely on AI to quickly draft briefs in unfamiliar state courts are especially vulnerable. The lack of transparent, state-by-state accuracy metrics from major vendors means that buyers are effectively acting as unpaid testers for these tools.

Questions to ask before you buy

Because vendors do not publish detailed, state-by-state or court-level accuracy breakdowns, you must conduct your own due diligence. Use this checklist during your next software trial or sales demonstration.

  1. Ask for localized accuracy metrics. Ask the sales representative if they have published or internal reports showing retrieval accuracy and hallucination rates specifically for your state's trial courts or administrative bodies. Note that as of 2026, neither CoCounsel nor Lexis+ with Protege publishes this data. If they cannot provide it, you must assume their aggregate "high accuracy" claim is based mostly on federal appellate data. To prepare for this conversation, you can read our advice on Why So Many Legal AI Vendors Hide Their Pricing (And How to Get a Real Number).

  2. Conduct a "cold case" trial. Do not let the vendor run the demo using their own pre-selected questions. Bring three real, previously answered research questions from your firm's active cases. Ensure at least one question involves a lower state court or local administrative agency. Run these queries during the trial and manually verify every single citation the tool returns. If you want to compare multiple platforms during this trial, refer to our directory of the Best AI Legal Research Tools for Law Firms (2026).

  3. Test a multi-state synthesis query. If your firm handles work across multiple borders, type a query that asks the tool to compare a specific standard across three of your active states. Carefully check the output to see if the tool blends the standards together or incorrectly attributes one state's rule to another.

  4. Identify the exact data sources. Ask the vendor what percentage of their database consists of administrative-law rulings, local municipal codes, and unpublished state trial court orders for your specific region. If your practice relies heavily on these sources, and they are thin in the vendor's database, the tool's utility will be severely limited.

  5. Compare generalists against specialized tools. If your firm's practice is highly concentrated in a single area, evaluate a specialized platform. Weigh the depth of a tool like CompFox against the broad claims of a national vendor. The narrow tool will lack nationwide scope, but its local accuracy may protect your firm from costly errors.

FAQ

Do Lexis+ AI and Westlaw's AI research tools hallucinate?

Yes. The peer-reviewed Stanford RegLab study published in the Journal of Empirical Legal Studies in 2025 found that LexisNexis's Lexis+ AI hallucinated in roughly 17 percent of queries, and Westlaw's AI-Assisted Research hallucinated in roughly 33 percent of queries. Both tools performed better than a bare GPT-4 baseline at 43 percent, but neither tool was entirely free of errors.

Do these tools have less coverage in some states than others?

At the database level, both major vendors index all state and federal case law. The primary gap is not missing documents, but rather retrieval accuracy. Research shows that AI models are significantly less accurate when retrieving and synthesizing state trial court opinions and administrative-law rulings than when handling federal appellate case law.

Can I get a jurisdiction-by-jurisdiction accuracy report from a vendor before buying?

No. Neither CoCounsel nor Lexis+ with Protege currently publishes detailed accuracy or hallucination metrics broken down by individual state or court level. Buyers must rely on independent academic studies or conduct their own internal testing during product trials.

Is a narrow, single-state tool a worse choice than a nationwide one?

Not necessarily. For a firm with a highly specialized practice, a narrow tool can be a much safer and more reliable choice. CompFox, which focuses exclusively on California workers' compensation law, is built on a proprietary, fine-tuned database that offers deeper, more reliable coverage in its niche than a generalist platform.

Has anything improved since the 2024 Stanford study?

Both major vendors have updated and rebranded their AI tools since the original study was conducted. Lexis+ AI was renamed Lexis+ with Protege in February 2026, and CoCounsel is now integrated into Westlaw. Both vendors scheduled mid-2026 updates to address known errors in multi-jurisdiction comparison. However, no independent, peer-reviewed studies have yet replicated the Stanford methodology on these current versions.

The bottom line

When evaluating legal AI, "nationwide coverage" is a marketing term, not a technical guarantee. The underlying technology remains highly dependent on data density. It performs exceptionally well on the heavily documented rulings of federal appellate courts, but it is prone to significant errors when navigating the thin, fragmented records of state trial courts and administrative agencies.

For a law firm, relying on an inaccurate AI tool is an operational and ethical risk. The hours you save on drafting can easily be lost to manual citation checking. Before your firm commits to an expensive subscription, you must bypass the standard sales presentation. Run your own practice-specific test queries, verify every citation against primary sources, and remember that a specialized, narrower tool is often the more reliable choice.