(Title Image: Summit Art Creations | shutterstock.com

Generative AI and its Potential Use/Misuse in War Crimes Cases

By Terry Beitner

Introduction

If you have been following the news about AI and the law, you have likely seen many articles about how AI chatbots (in this article, I will refer to AI chatbots as “chatbots,” “machines,” and “models” interchangeably) make errors that end up in written pleadings around the world (see here, here, and here).  Lawyers have been sanctioned by judges through the imposition of fines and, in one Ontario case, potential criminal contempt of court proceedings involving the use of ChatGPT (see here). Generally, counsel are sanctioned for not checking the accuracy of their cited references (e.g., case citations or other cited authorities) before filing their pleadings with the court. The impact of counsel’s failure to check their citations is that, upon verification, the cited references are occasionally incorrect. Chatbots are known to have simply made up case references to support an argument (otherwise known as “hallucinations”).

Despite the widely publicized pitfalls of early AI chatbots, this article argues that the integration of AI into the legal profession provides practitioners with unprecedented capabilities, transforming the management of massive datasets in both everyday practice and complex international criminal investigations and trials. AI offers the ability for an almost instantaneous analysis and retrieval of information from large datasets without having to manually tag relevant pieces of data. The machine will operate based on requests from users asking for a specific type of analysis and result. The machine will scan the entire dataset and decide what information is relevant to respond to a user’s query. In doing so, the machine will conduct its own “analysis” of the sources and provide an answer in various user-defined formats.

While technology has impacted knowledge workers before (notably through the personal computer and the internet), the introduction of these models represents a qualitative shift. The promise offered by this technology is that it moves beyond information retrieval and calculation into generative reasoning and linguistic synthesis. Today’s models aggregate disparate pieces of information and produce the information as a whole, applying the type of analysis requested by the user. AI puts enormous computing power in the hands of an individual or a small team that can work at previously unimagined speed, combined with the ability to produce outputs in a variety of formats. The work that goes into producing the final product will still, however, require high-level human analysis to verify and refine the output to make sure that it is fit for purpose.

Although the end product produced by some of these machines may convincingly resemble high-level cognitive reasoning, it is not necessarily so. These machines simulate intelligent behaviour; today’s AI tools do not understand their own output and meaning in the same way that humans do (see here for a brief discussion on the “reasoning capabilities” of these machines).

The Introduction of Frontier Models Is a Game Changer for Our Profession 

When OpenAI released ChatGPT to the public in November 2022, we saw the beginning of a new era of technology that put a powerful tool into the hands of the public. However, we, as early adopters of this technology, were not quite sure what it could do other than carry on conversations on a wide array of topics or write real estate listings very quickly with simple prompts. Since 2022, OpenAI (ChatGPT), Google (Gemini), and Anthropic (Claude) have regularly released new, more powerful state-of-the-art (“frontier”) models every few months. We have come a long way since 2022.

The Morgan Stanley Institute indicates that “nearly $3 trillion of AI-related infrastructure investment will flow through the global economy by 2028, with more than 80% of that spending still ahead” (see here). The investments in this technology are phenomenal, and Morgan Stanley also posits that “AI has become a structural force in economic expansion.” The legal profession has also bought into AI. For example, Kirkland & Ellis LLP, the world’s highest-grossing law firm, has recently announced that it will spend $500 million building its own AI platform (see here). Furthermore, there are recent media reports of law firms offering “multi-tier” pricing models where clients choose between an “AI-heavy” output or “slower and costlier human-led legal advice” (see here). This technology is here to stay and will likely form a major part of our lives and the economy.

The View From The Bench 

As mentioned above, courts have already stepped in to deal with counsel carelessly submitting shoddy work into the record. However, judges in the United Kingdom and the United States have offered more guidance to the profession. Maritza Braswell, Magistrate Judge of the U.S. District Court for the District of Colorado suggests that the frequency of hallucinations may be a diminishing concern in light of ongoing improvements in AI technology. In her view, the legal community should focus on bias caused by these models’ training data rather than the diminishing concern of hallucinations (see here). Judge Braswell adds that the issue of bias is not sufficiently discussed in the current discourse on AI. She explains that hallucinations are easy to verify by checking the citations given for sources. They are visible. Bias is not as easily discernible. Chatbots are trained on the internet and reproduce its biases. For example, she notes that in certain legal areas, such as employment discrimination, a chatbot may provide a framework that is “slightly weighted” toward the defence perspective. In her view, this is because of the higher number of blog posts or articles written by large employment defence firms with substantial resources and published on the internet.

Judge Braswell also mentions that her colleagues on the bench use AI to summarize transcripts of testimony, identify where witnesses may have commented on a particular issue, and synthesize testimony. She also uses a chatbot for plain-language translation when presiding over cases with unrepresented litigants. She says that she employs the chatbot to help her explain rulings to these litigants who may not understand the “industry speak” and uses the model to provide a plain-language, accessible explanation.

Additionally, Judge Braswell discusses how she uses Claude (by Anthropic), a publicly available model comparable to ChatGPT. Using a model called Cowork, Judge Braswell explained that she creates local folders and adds specific sources to the folders, such as the Federal Rules of Civil Procedure, the local rules of her district and Standing Orders. She then instructs the chatbot to answer her questions relying only on the provided sources and to not provide an answer if it is not based on a rule. If no rule is available, she then allows the chatbot to formulate an answer based on other rules or sources, provided the machine identifies the sources relied upon.

The judiciary in the United Kingdom also has also incorporated AI into their practice. The Right Hon. Sir Colin Birss, Chancellor of the High Court, spoke at The City of London Law Society on April 22, 2026, where he discussed how the judiciary in the UK is using AI (see here). He began by pointing out that judges can use a secure form of Microsoft Copilot that is available to all judges in England and Wales.

In addition to highlighting the impact of AI as an in-house transcription tool, Lord Justice Birss added that judges have found AI useful for producing anonymized judgements. He indicated that after a judge completes the draft ruling, they can ask AI for suggestions to anonymize the text. He noted that “…some judges have commented that the AI has identified pieces of information as candidates for anonymisation, which are not the obvious things to redact (like the names and so on). The AI identified information combinations which might risk a kind of jigsaw identification of the individuals concerned.”

Lord Justice Birss also mentioned that he uses AI to identify “internal inconsistencies in my own work.” He explains that after writing a judgement, he submits it to a chatbot to identify any internal inconsistencies. The Chancellor says the machine is “remarkably effective” and adds that, although he doesn’t always agree with the model’s output, “it has been helpful and I have clarified wording in draft judgments as a result.” Lord Justice Birss noted that AI has also “transformed” his ability to “find things in emails and files” since he no longer must conduct word searches in old emails.

The Chancellor also discussed the issue of access to justice for unrepresented litigants. He observed that globally, unrepresented litigants are increasingly using AI to draft materials submitted to court. He noted that although the volume of material can pose challenges, this use is, in his view, “pro-access to justice.” He explains that material may be “very long and not right” but adds that not all of it is wrong or of poor quality. Lord Justice Birss stated that in his experience and the experience of other judges, an unrepresented litigant’s case is often “…presented more clearly and coherently than I would have expected in similar circumstances in the past.” For further brief commentary on AI and access to justice see here.

The View From the Bar 

Based on their 2024 study of the issue, Thomson Reuters indicates that the top five use cases for law firms using these machines are legal research, document review, briefing or memo drafting, document summarization, and correspondence drafting (see here).

Harmonic, a company that offers tools to help businesses secure their data by monitoring and controlling AI use in the workplace, adds further insight. Harmonic identifies the following uses of AI by in-house corporate counsel: contract review, litigation strategy, regulatory analysis, policy drafting, compliance analysis, and intellectual property review (see here). Their study indicates that the legal and governance department is the largest user of AI across company departments and that, as a result, in-house legal teams have reduced the amount of work outsourced to outside firms.

In her article “Most Legal Work Isn’t Worth Paying For Anymore,” Mary Shen O'Carroll argues that the expanded use of AI by in-house teams should move the enterprise/law firm financial model away from hourly billing to value-based purchasing (see here).

Generally speaking, the use of AI for legal research is taking a firm hold. Commercial legal service providers in Canada and the United States, such as Lexis+ and Westlaw, supplying traditional boolean-based research tools are now incorporating chatbots into their offerings. You simply ask the chatbot a legal question in a conversational manner, and the machine will provide a narrative answer with citations to cases and other legal reference material. Even CanLII, the research tool offered by the Federation of Law Societies of Canada, has incorporated AI search abilities in addition to the traditional boolean search methods.

Additionally, the Access to Algorithmic Justice (A2AJ), a joint initiative of Osgoode Hall Law School and the Lincoln Alexander School of Law in Toronto, offers a powerful, free AI assistant for legal research. One commentator posits that, in their view, the quality and utility of the system that relies on A2AJ far exceed that of CanLII, Lexis+ Protégé, or Westlaw Edge searches (see here). Of course, the commentator adds that the usual cautions apply. Counsel must verify the cases by reading the relevant passages and decide whether client information can be shared with the machine.

The introduction of artificial intelligence (AI) as a tool for legal research is another example of how technology has impacted the profession over time. In Canada, over the last forty years, legal research has progressed from purely paper-based methods to computer-assisted research designed to run boolean searches. As with any new technology, AI creators will need time to deal with issues regarding accuracy; however, rapid improvements to these models demonstrate that it will not be long before AI is regarded as another standard, reliable research tool.

Document Review, Analysis, Summarization, and Plain-Language Translation 

In my view, document review is one of the most exciting use cases for the profession. As I mentioned above, one can interact with a collection of information with no need to prepare materials for analysis other than to make sure that they are in a format that is machine readable for the chatbot in question. The input can be almost any type of file used by today's knowledge workers (Word, PowerPoint, email, PDFs, YouTube videos, etc.). The analysis of the datasets and the output are done in minutes.  The form of output is chosen by the user and can include a wide array of formats, such as narrative text, bulleted summaries, data tables, videos, audio files, and visual data maps that resemble the products criminal analysts have been using for years, such as those produced by i2 (see here). One does not need any specialized knowledge to engage with AI, although learning how to get the best results is an iterative process through prompt experimentation.

The art of working with a chatbot lies in human-to-machine communication. Interactions with chatbots are similar to the way we have been interacting with Google Search for almost 20 years. We ask questions, feed keywords and get results in the form of web links. In the context of working with AI chatbots, our communications with chatbots are called prompts. Working with chatbots is a learning process; although today's models are good at understanding human prompts, in my experience, it usually takes a few tries to get the desired result. This means rewording the prompts until there is clear communication between the user and the machine.

Chatbots are designed to follow the instructions provided in prompts, however their responses are also guided by parameters that include guardrails developed to prevent the machines from causing harm. In this context, prompt design is key; effective communication with a machine is an iterative process of trial and error that yields almost instantaneous results.

For example, one can provide a machine with a corporation’s financial report for the first quarter of the fiscal year and ask for a summary. The user evaluates the response and determines that the output is too vague and lacks quantitative data. Additionally, the narrative structure is not suitable for use by corporate executives. The user then proceeds to adjust the prompt with more specific instructions. The next prompt will tell the model to act as a senior product manager presenting its findings to executive leadership and summarize the report in bullet points. The user then instructs the machine to focus heavily on metrics and keep the response to under 150 words. If further nuances are required, the user will proceed with the same evaluation and gap analysis until satisfied with the response.

Similarly, in the legal context, counsel can work with a secure chatbot to draft correspondence to clients relating to a recent decision or a description of the legal steps in an action. Counsel can adjust the output of the machine to fit the circumstances and needs of the client. We should remember that this technology is new; however, the competition among the key corporate players continues to drive improvements and innovation. As these machines become “smarter,” they are better able to discern the user’s intent, and prompt design becomes less important.

In the context of litigation and investigations, secure models may assist in the analysis of witness statements, forensic reports, or other relevant documents. These models can generate customized tables and data maps, charting connections (and gaps) between pieces of evidence and the elements of a case that must be proven. Ms. Joyce King, Deputy State’s Attorney, Frederick County State’s Attorney’s Office, MD, mentioned her own uses of AI during a recent webinar (see here). She discussed using AI for trial preparation (e.g., to organize discovery, build timelines for juries, and act as investigative aids) but warned against its use to craft legal arguments because of its error rate.

Google’s Gemini Notebook (see here) is a publicly available free tool that demonstrates the power of AI to analyze large amounts of information using prompts as described above. This machine is similar to Anthropic’s Cowork used by Judge Braswell mentioned above. Document collections are contained in “notebooks” that hold up to fifty “sources” in the free version and up to six hundred in the premium tier. Once sources are uploaded, one can query an individual source, a selection of specific sources, or the entire collection. The uploading process is simple. Once one gathers the sources in a local folder, Gemini Notebook can be instructed to collect the selected sources and copy them. The process is similar to a simple copy and paste. Cowork can operate within a local folder on one's computer without having to move files. There are limits to Gemini Notebook’s “memory” that vary depending on whether you are using a free or a paid account.

When interacting with Gemini Notebook, the machine limits its answers to the information contained in the uploaded material. In other words, the model does not access the internet to search for information to supplement its answers, thereby operating in a closed environment. The model also footnotes its assertions with references to the source(s) in question, highlighting the specific text relied upon to formulate its responses. This reduces the likelihood of hallucinations because the chatbot uses only the information identified as sources. One should, however, remember that the machine can still make mistakes, but it is less likely to do so and verification of assertions is easy. For example, you can upload the trial judgement of a case and the decision of an appellate court and have the chatbot tell you what holdings were reversed or confirmed.

There are other interesting features from Gemini Notebook, such as having the machine create a short podcast based on a source by using its “audio” feature. There are two AI-generated voices, one male and one female, that will discuss the information in the source, and you could even interrupt the conversation and ask questions. Gemini Notebook will also create other products (based on the sources) that it refers to as mind maps, infographics, flash cards, etc., and these operations are coded into the machine and therefore do not require a special prompt. One just clicks on the various buttons on the site. In my own use of Gemini Notebook, I have found the tool particularly helpful for recalling the content and findings of a case after completing my own analysis and when returning to a ruling to recall specific points. As illustrated above, AI offers simple-to-use tools with powerful and fast analytical capabilities that should improve any knowledge worker’s workflow.

War Crimes Cases and AI 

The investigation and prosecution of core crimes cases (war crimes, crimes against humanity, and genocide) present some unique challenges to investigators, their legal advisors, and, once these matters end up in court, defence lawyers and judges who are tasked with establishing the truth. While the domestic trial preparation use cases highlighted by Ms. King such as automated timeline creation and discovery organization may improve workflows in national litigation, their utility is significantly amplified when applied to mass atrocity situations involving large and complex fact patterns, such as those found in core crimes cases. Core crimes cases usually require the analysis of large, varied datasets in digital or analog form, such as witness statements, electronic communications, paper records, and as physical evidence.

I include all types of litigation when I refer to the “prosecution of core crimes cases” and not solely criminal prosecutions. There is a vast array of legal remedies available in the core crimes “accountability space” that is not limited to criminal prosecution. These include immigration remedies that bar alleged war criminals from entering or remaining in refugee/immigrant receiving states, revocation of citizenship, and state-imposed sanctions. Therefore, the following discussion of the use of AI applies to any litigation involving complex facts and law, including complex civil cases, organized crime, or terrorism-related matters.

AI’s analytical capability is particularly relevant for complicated core crimes cases because of the need to analyze an often-complex factual matrix. For example, in prosecutions of commanders for the criminal acts of their subordinates, it is necessary to establish links among various actors to recreate the operational chain of command. Such analyses may involve a great volume and wide variety of information (e.g., witness statements, government regulations, military regulations, government documents, written orders issued by commanders, intelligence intercepts of coded communications, etc.). To illustrate the sheer volume of such material, in the Prlić case before the ICTY, the trial chamber heard 207 viva voce witnesses and admitted 9,756 documents into evidence (see here at para 268). AI can sift through the data and help draw the lines between the commander, their subordinates, and the crimes in question. Similarly, in the Ongwen matter before the International Criminal Court (ICC), the court heard 186 witnesses (see here at para 31) and 4,200 items were submitted into evidence (see here at para 245). The trial ruling was 1,077 pages long, and the appeal judgement was 611 pages long. These cases not only involve significant amounts of evidence, but the court rulings in international criminal law cases are often quite long, especially at the ICC.

One of the challenges in the investigation and prosecution of core crimes cases that distinguish them from ordinary criminal offences concerns the unique legal elements of proof required to establish criminal liability. For example, to prove a criminal offence in Canada, one must establish both the material element of the offence (i.e., the conduct or the actus reus) and the mental element (i.e., the intent or knowledge, mens rea). Under international criminal law there are additional elements that must be proven.

For example, the war crime of “wilful killing”, that is, causing the death of a person contrary to the law governing armed conflict, requires proof of additional elements than those required to prove the similar national offense of murder. One variation of murder in Canadian law occurs when one causes the death of a human being by means of an unlawful act with the intention to to cause death or bodily harm that the person knows is likely to cause death. The prosecution must prove the material element of the offence (actus reus) with evidence establishing the unlawful physical act(s) of the accused that caused the death. Additionally, the prosecutor must establish the mental element (mens rea) that is, the intent to kill.  

Similarly, under international law, the prosecutor must prove that the person meant to engage in the unlawful conduct (acts with intent) and the material physical act causing the death to prove the crime of wilful killing.  However, the following additional elements must be established before the court. The prosecution must establish that the victim was not taking active part in hostilities and prove that the perpetrator’s conduct took place in the context of, and was associated with, an armed conflict. In other words, the existence of the armed conflict played a substantial role in the perpetrator’s ability to commit the crime, his decision to commit it, the manner in which it was committed, or the purpose for which it was committed (see here and here at para 58).

Furthermore, the prosecutor must also establish that the perpetrator was aware of the factual circumstances that established that the victim was not taking active part in the hostilities and that he the perpetrator was aware of the armed conflict. Additionally, all of these elements are to be proven according to the customary standard of proof in criminal cases: beyond a reasonable doubt. The context in which the alleged crime occurred (the existence of an armed conflict, the status of the victim at the time of the event, and the impact of the armed conflict upon the accused’s ability and decision-making related to the criminal act) must also be established beyond a reasonable doubt. Therefore, a great amount of evidence relating to the context of the event is required to prove a single war crime compared to a national offence of a similar character. This evidence can involve a substantial number of witnesses, documents, intelligence intercepts, and other evidence, including expert testimony. As a result, trials of war crimes cases are extremely complex and require a substantial amount of evidence to establish facts not required in the prosecution of national crimes of a similar character.

Core crimes often involve mass atrocities that generally, but not necessarily, occur during armed conflict. Consider the armed conflicts that occurred in Rwanda in 1994 and in the Balkans following the breakup of the former Yugoslavia in 1991; these conflicts involved mass atrocities, including estimates of more than 800,000 killed in Rwanda, and a legal finding of genocide by the International Criminal Tribunal for Rwanda (ICTR). The conflict in the former Yugoslavia also led to the establishment of the International Criminal Tribunal for the Former Yugoslavia (ICTY), which also found that genocide occurred in Srebrenica, Bosnia and Herzegovina. Today, it is up to the ICC, as the permanent international criminal tribunal, to hear these types of cases and deal with the challenges of large evidentiary data sets.

One of the unique features of the trials at the ICTR is that evidence was predominantly witness testimony. Lawyers and the Canadian public are well aware of the frailties of eyewitness testimony due to Canada’s history of wrongful convictions revealed by various Royal Commissions of Inquiry (for a summary of findings, see the Government of Canada  “Report on the Prevention of Miscarriages of Justice” here).

Witnesses to the atrocities in Rwanda were frequently called to testify in multiple proceedings relating to these events. For example, it was not uncommon for witnesses to give evidence before an international tribunal, a national court in their home state, a court in a third-party state (e.g., Canada) and make additional statements to other state or human rights investigators. The issue of multiple previous and sometimes contradictory statements loomed large in several trials.

These evidentiary discrepancies, which often emerge between early investigative statements and later courtroom testimony, or even between different previous out-of-court statements, must be viewed within the context of investigations where teams are looking at crime scenes with hundreds, if not thousands, of victims and a comparable number of perpetrators. Additionally, witnesses may give numerous statements and provide formal testimony over a period of several decades. Furthermore, witnesses may be called to testify repeatedly regarding the same events, but during trials against different accused.

Working in parallel with the ICTR and ICTY, countries in Africa, Europe, and Canada began prosecutions to seek accountability for perpetrators of core crimes from these and other conflicts. For example, Canada prosecuted two Rwandan nationals in separate proceedings. Désiré Munyaneza was convicted of war crimes, crimes against humanity, and genocide by the Superior Court of Quebec in 2009 (see here, and upheld on appeal, see here). He was the first person charged under Canada's Crimes Against Humanity and War Crimes Act, SC 2000, c 24; in 2013, Jacques Mungwarere was acquitted by the Superior Court of Justice (see here). Additionally, proceedings are currently underway in Newmarket, Ontario, against Ahmed Eldidi for alleged ISIS-related war crimes committed in Iraq. The trial is scheduled to begin later this year (see here).

In Mungwarere, the court noted that one witness provided ten previous statements between 1999 and 2011 regarding the events that occurred in Rwanda in 1994 (see here). These included statements to ICTR investigators in 1999, testimony in ICTR trials in 2001 and 2002, and statements to American investigators in 2002; the witness also made declarations in a 2003 Canadian asylum application, along with statements to the RCMP in 2004 and on four other occasions between 2010 and 2011, showing that in the Rwandan context there are situations where witnesses have given dozens of previous statements relevant to specific proceedings.  To further complicate matters, statements are typically given in Kinyarwanda and then translated into French or English. The issue of reliable translations adds an additional challenge to the analysis of this eyewitness evidence, either during an investigation or trial.

Consequently, eyewitnesses in core crimes cases are often expected to have provided prior statements to a wide range of actors. These include state and non-state criminal investigators, such as those with the Commission for International Justice and Accountability (CIJA; see here), international tribunals, and national courts. Witnesses may also give statements to UN or regional state-based international commissions of inquiry (such as those established by the Organization of American States; see here), or to national truth and reconciliation commissions established by states to examine alleged human rights abuses on their territory (e.g., Chile 1990, (see here); South Africa 1995, see here; Canada 2008, see here; etc.).  Furthermore, human rights workers and dedicated investigative mechanisms, such as the International, Impartial and Independent Mechanism for Syria (see here), the Independent Investigative Mechanism for Myanmar (see here), and the Independent International Commission of Inquiry on Ukraine (see here), also collect evidence and witness statements to support investigations and future prosecutions.

Although the issue of prior inconsistent statements is not uncommon during ordinary litigation, core crimes cases are unique due to the extraordinary number of previous statements that witnesses may have given in the ordinary course of events and the time span over which they are produced. Therefore, the sheer number of previous statements increases the likelihood of previous inconsistent statements, making these witnesses more vulnerable during cross-examination than in the course of ordinary litigation. In this context, the use of AI can assist in the analysis of eyewitness evidence, readily identify inconsistencies across statements, and cross reference this evidence with other information that may mitigate or highlight evidentiary discrepancies.

Additionally, modern approaches to the investigation of core crimes involve variations of the classic core crimes investigative method, where one often starts with an allegation against a specific individual and then builds the case around the allegation. Modern-day conflicts have pushed investigative bodies to adopt different investigative methods. For example, Canada has adopted a methodology previously used by its European partners as well as the ICC. These operations do not necessarily involve an investigation of a specific individual or event. The investigative agency in question puts out a call for people to come forward and identify themselves as potential witnesses to atrocities. Practices vary across investigations, and therefore investigative agencies may take formal statements or decide to limit their interactions with potential witnesses to the gathering of general information.

The objective is to gather information that may be useful for future investigations. The Royal Canadian Mounted Police describe the operation as a “broad, intelligence led intake process designed to collect, preserve, and assess information” potentially relevant to investigations under Canada's Crimes Against Humanity and War Crimes Act (see here). Additionally, this information may be shared with other states’ investigators who may be launching an investigation relevant to a witness's information. The RCMP has two structural investigations underway: one relates to the war in Ukraine, and another concerns the crimes carried out by ISIS against the Yazidi population between 2014 and 2017. The RCMP also indicate that they are in the developmental stage of a structural investigation into the Israel-Hamas armed conflict.

The reliance on novel information-gathering processes demonstrates the large amount of information that is accumulated in the context of core crimes cases, and as noted above, often over the course of many years before a case comes to court. Government investigative agencies and international bodies, such as the CIJA or the United Nations Investigative Team to Promote Accountability for Crimes Committed by Da’esh/Islamic State in Iraq and the Levant (UNITAD - see here), operate with sophisticated information management systems (see here). Their systems may include computer vision, a type of artificial intelligence that enables machines to analyze and interpret images and videos and describe their content (see here and here), and their operations are designed to support prosecutions and collect vast amounts of material to enable the prosecution of core crimes. For example CIJA has secured almost 1,200,000 pages of documents and interviewed 5, 500 witnesses (see here).

Historically, even basic AI chatbots were not widely available, and therefore defence counsel would not have had these easily accessible and powerful resources for the analysis of large datasets.

Today, AI may help address fair trial concerns where defence teams may not have the full resources of governments and international organizations behind them. Indeed, in the Ongwen case mentioned above, the defence argued, unsuccessfully, that it was prejudiced by the amount of and the manner that 4,200 items of evidence were admitted into evidence (see here at para 245). The trial court admitted the material without making admissibility rulings on each item but as part of a “holistic” assessment of the disparate pieces of evidence submitted to the court (as is permissible at the ICC, see Appeal Chamber ruling here at para 505). The court completed the case without issuing a ruling on every piece of evidence in either interlocutory decisions or its final judgement. In Sainovic et al., a matter before the ICTY, the prosecution ultimately disclosed 1,755,372 pages to the defence (see here at para 36). Today, defence counsel may look to AI to assist them in the review of these vast amounts of material.

Finally, experts such as historians, linguists, criminal forensic specialists, military analysts, and others play an important role in core crimes cases. Counsel may also use AI to prepare for cross-examination or preparation of their own experts. Dealing with expert witnesses is labour-intensive, requiring a comprehensive review of their prior publications and testimony, again an area where AI can assist.

Given the sheer volume and complexity of evidence in mass atrocity cases, AI provides new tools to quickly cross-reference vast quantities of information and uncover critical connections that might otherwise be overlooked during an investigation or trial. Consequently, any tool that helps analysts, investigators, or counsel deal with this vast quantity of information/evidence is, in my view, an aid to the search for the truth, especially in the context of core crimes cases.

Conclusion 

Core crimes cases involve an enormous amount of evidence, which is largely due to the fact that the elements of the offences that must be proven go beyond the direct perpetrator and the victims. The context of the events in question must be established to the criminal standard of beyond a reasonable doubt, in addition to the conduct of the perpetrator and their state of mind. As illustrated above, there are many opportunities for lawyers, investigators, analysts, and judges to incorporate AI into their workflow to achieve their objectives.

We are at the beginning of a journey where the incorporation of AI into the workplace for knowledge workers may have profound impacts on our profession and society at large. Although lawyers involved in cases with large datasets may benefit from the use of these new and powerful tools, we should proceed with caution and awareness of our ethical obligations when using AI (see here). In a nutshell, check your work, understand these new tools and their limitations, and respect client confidentiality.

In the words of Ronald Reagan, “Trust, but verify.”

________________________________________

Citation: Terry Beitner, “Generative AI and its Potential Use/Misuse in War Crimes Cases" (2026) 10 PKI Global Justice Journal 2.

About the Author:

Patrick Diotte

Terry Beitner is the former Director and General Counsel of the Crimes Against Humanity and War Crimes Section of the Department of Justice Canada. He is now a part-time professor of international criminal law in the Faculty of Law at the University of Ottawa.