Responsible Use >
Summarization
Core Concepts
Generative AI can be helpful for summarization tasks, but it is imperfect
GenAI can be effective at distilling lengthy documents into concise summaries, creating chronologies from disparate information, and organizing information into structured formats. However, LLMs focus on word patterns and frequency. They may miss language that prioritizes or qualifies concepts, such as negation (e.g., the words ‘no,’ ‘not’ and ‘without’). This is an inherent part of LLM architecture, not a flaw per se. But this architecture can undermine the accuracy of a GenAI summary. As with all GenAI output, the summary must be reviewed for accuracy, neutrality, and completeness.
Accuracy of Generative Ai may degrade as input volume increases ('context rot')
An LLM’s context window is the maximum amount of information a GenAI system can process or analyze during a single interaction. It represents all the information the AI system can see in a given conversation; the system cannot process any information outside of its context window. Information entered into a GenAI system is broken down into tokens, sequences of textual characters that make up part or all of a word. The context window is measured in tokens and its size varies system to system, depending on the AI system’s architecture. Earlier AI models have context windows of about two to three thousand tokens (about three to six pages or fifteen hundred to three thousand words), while newer models range up to one to two million tokens (fifteen hundred to five thousand pages or seven hundred fifty thousand to one million words).
The larger the context window, the more information an LLM can process. As engagement with a GenAI system proceeds (through a series of prompts or conversations), its context window fills. GenAI systems tend to favor tokens at the beginning and those at the end of the input. Notably, the size of a system’s context window is not always obvious and the system may not provide this information to the user.
When using GenAI to summarize documents, if the quantity of information the system is asked to summarize exceeds its context window, the system will not retrieve information outside the context window and may struggle to retrieve information near the middle of the context window, leading to diminished accuracy. This phenomenon is called 'context rot.' Even LLMs with large context windows often prioritize information at the beginning and at the end of the input, potentially omitting important information in the middle. This is referred to as 'lost in the middle.'
Appropriate use cases for Generative AI summarization and organization
GenAI can summarize lengthy briefs, exhibits, or transcripts, create case timelines, organize facts, and structure complex regulatory materials. It also has been used to decipher handwriting in filings by self-represented litigants (with moderate but not 100% accuracy). Although GenAI tends to hallucinate less when working with a circumscribed set of materials, it still can make errors. In a recent case, counsel used GenAI to summarize documents and deposition transcripts for a declaration provided to the court; the declaration included hallucinated facts. GenAI summaries require human review and verification.
The neutral prompting principle
How a prompt is phrased has an impact on GenAI output. Prompts given to GenAI tools should be objective and neutral, without presupposing or signaling a particular outcome. There is a difference between asking GenAI to ‘objectively summarize the parties’ arguments’ and asking it to ‘explain which party has the better argument and why.’
Summarization and organization versus analysis
While it is not improper to use GenAI for some types of analysis, it is vital to be clear about what you are asking it to do. There is an important distinction between asking GenAI to summarize or organize the parties’ arguments (appropriate) and asking GenAI whether these arguments are correct (not appropriate). This issue can manifest in subtle ways. For example, there is a difference between asking GenAI to identify ‘all allegations’ versus ‘key allegations.’ The former is more clearly summarization while the latter involves analysis and requires the GenAI to make choices about what is legally significant. One cannot rely on a GenAI tool to clearly differentiate among the tasks of summarization, organization, and analysis. GenAI-generated summaries and chronologies should always be verified and reviewed for accuracy.
Short Videos
-
Preparing for Hearings with GenAI Tools
In this video, U.S. District Judge Jonathan Hawley describes how he uses Westlaw CoCounsel to prepare for hearings and conferences. He explains how the tool assists with developing summaries and creating timelines using publicly available documents.
-
GenAI Systems and Data Protection: How to Configure Privacy Settings
This video explains the implications of the type of GenAI system used for data protection and includes a demonstration of how to configure privacy settings.
Practical Guides
-
Talking to your Law Clerks about GenAI
Checklist of topics to address when discussing GenAI use with law clerks.
Frequently Asked Questions
GenAI use falls along a spectrum. Every summarization task inherently involves some judgment about what to include and what to exclude; every organizational task inherently involves some judgment as to hierarchy. It is not necessarily wrong to allow GenAI to engage in analysis. It is important to be aware that it is being done. To limit the degree of judgment and analysis exercised by a GenAI system, judges should be intentional about crafting their prompts and which GenAI functions they select (for example Westlaw CoCounsel’s Summary versus Review features).
While many legal (and some general-purpose) AI systems provide links that enable you to verify the accuracy of what the system has included, there is no way to know or check what was omitted from the summary without reading the original source. When confronted with a voluminous set of case-related documents, GenAI summaries can still be helpful. However, judges should be aware that they may not be 100% complete and may contain errors. The original source should always be reviewed before using an AI-generated summary.
It depends. Factors that can impact the reliability of GenAI case law summaries include the complexity of the fact pattern, procedural history, and legal issues; whether there are concurrences or dissents; and whether the case is well known and has been subject to significant media coverage; and the wording of the prompt. There has been little independent empirical research studying the accuracy of GenAI case law summaries.
Curated Resources
-
Silent but Deadly — Context Rot Problems in Legal
Syntheia (2026)
-
Judging AI: How Judges can Harness Generative AI Without Compromising Justice
Judicature (2025)