“AI in the crime lab” has become a headline phrase, and like most headline phrases it hides more than it reveals. When people say it, they are usually talking about one of two very different things and conflating them is the single most common mistake I see.
The first is established statistical and algorithmic software that has been producing courtroom evidence for more than a decade. Probabilistic genotyping (the software that interprets complex DNA mixtures) is the marquee example. It is partly a black box, and in a growing number of cases its output is not merely an input to the evidence. It is the evidence.
The second is the newer machine-learning and deep-learning wave everyone pictures when they hear “AI.” That wave is real, but for now it lives mostly upstream (in investigative triage, screening, and comparison-assist) and has not yet become the load-bearing basis for many convictions.
The practical story, then, is not that crime labs have been automated. It is that a small number of algorithmic tools have quietly become dispositive in individual cases, while the transparency, standards, and disclosure infrastructure around them lags well behind. That gap is where the legal questions live and where a well-prepared defense does its work.
How Forensic Labs Actually Use AI
Ordered roughly by courtroom maturity, the picture looks like this.
DNA mixture interpretation — the most mature and most litigated
Probabilistic genotyping software such as STRmix (used by many state and federal labs, including the FBI) and TrueAllele from Cybergenetics takes a complex DNA mixture and produces a likelihood ratio. The algorithm’s output is itself the evidence a jury hears. That is a categorically different thing from an examiner using a tool to help form an opinion. New York’s older Forensic Statistical Tool was later found to contain undisclosed coding problems, and STRmix experienced a validation-and-coding issue in Australia that affected roughly 60 cases. When the software is the witness, its defects become the case’s defects.
Two Texas cases, cutting both ways
Texas has watched this play out from both directions. In a 2018 Travis County homicide case, the judge excluded the STRmix results from two items of evidence after the DPS analyst acknowledged she had not followed the laboratory’s own interpretation procedures for the software. This is a textbook example of the “properly applied” problem, decided in a Texas courtroom under Texas’s own reliability rules.
Now cut the other way: Houston’s Lydell Grant was convicted of a 2010 murder and sent to prison. Years later, a TrueAllele reanalysis of the DNA mixture excluded him. When the reworked profile was run through the national CODIS database, it identified another man, who confessed to the killing. Grant walked out of prison in 2019. Same category of technology, opposite results. The lesson is not that probabilistic genotyping is inherently good or bad. It is that the output is only as trustworthy as the inputs, the version, and the analyst’s judgment sitting behind it, which is exactly why you have to be able to see all three.
Latent prints
AFIS and ABIS systems use algorithms to return candidate lists, and newer deep-learning models score similarity. A human examiner still makes the final call but the algorithm decides what gets compared in the first place, which quietly shapes the outcome.
Firearms and toolmarks
NIBIN and IBIS correlate cartridge-case and bullet images. Newer research layers 3D imaging and machine learning on top to generate objective similarity scores, with the explicit goal of replacing subjective “match” testimony with something measurable.
Facial recognition
This one is used heavily as an investigative lead rather than as trial evidence which is precisely why it is so easy to miss. Facial recognition can drive an arrest without ever surfacing in discovery, and it has been linked to several documented wrongful arrests.
Digital and multimedia forensics
Machine learning now handles image and video classification, CSAM detection, deepfake detection, and phone-extraction triage. Extraction tools increasingly add AI categorization on top of the raw data pull.
Emerging and research-stage
Bloodstain pattern analysis, forensic document and handwriting examination, forensic anthropology (trauma and age estimation), toxicology, and wildlife forensics all have active AI research but little courtroom footprint yet.
The honest summary: today AI is mostly a triage, screening, and comparison-assist layer, with DNA probabilistic genotyping as the standout exception where the algorithm’s output is the evidence.
The Standards Are Catching Up Slowly
There is real movement, but there is no unified “AI in forensics” standard yet. NIST and OSAC (the principal U.S. body) currently treat AI as a set of research-and-development needs scattered across disciplines, rather than as one overarching governance standard. The NIST AI Risk Management Framework is not forensics-specific, but it is increasingly cited as the backbone labs should map their AI tools onto. ISO/IEC 17025 accreditation and SWGDE guidance cover software validation broadly, and that is the existing hook labs are expected to use for any AI tool.
The most-cited critique remains the 2016 PCAST report, which set the “foundational validity” bar against which these tools are measured and specifically flagged that probabilistic genotyping had been validated only within narrow limits with a limited number of contributors and a limited mixture ratio.
Here is the gap to watch: standards for validating a tool exist, but standards for the transparency, disclosure, auditability, and bias-testing of AI specifically are still embryonic. That gap is exactly where the litigation lives.
How Would You Even Know It Was Used?
This is the sharpest question, and the answer is uncomfortable: often you would not know unless you go looking. A few practical levers:
- Read the bench notes and the case file, not just the report. A lab will happily report “a likelihood ratio of 540 million to one” without foregrounding that software produced it. The software name, version, parameters, and the assumed number of contributors live in the underlying documentation and not the summary.
- Ask directly in discovery. Demand the software name and version; internal and developmental validation studies; parameter settings and any analyst “judgment calls”; run logs; and the underlying STR data and electropherograms so a defense expert can re-run the analysis independently.
- Watch for investigative-lead laundering. Face recognition and similar AI leads can generate a suspect and then never appear in the file. Ask specifically whether any algorithmic or AI tool was used at any stage of the investigation (not just in the bench work).
- Check for version and validation mismatches. Was the specific version used validated by this laboratory, for this type of sample?
- Treat unreplicated numbers with caution. Different programs (STRmix, EuroForMix, TrueAllele) can differ by a thousand-fold to millions-fold on identical data. A single, unreplicated likelihood ratio deserves scrutiny, not deference.
The Source-Code Fight
Increasingly, AI is embedded in the instrument-and-software pipeline rather than bolted on which makes it easy to overlook, and which is exactly why the fight over source-code access matters so much.
State v. Pickett (N.J. App. Div. 2021) was the watershed: the first appellate court to order TrueAllele source-code disclosure to the defense, under a protective order, rejecting the trade-secret shield and holding that limiting review to “a device at the prosecutor’s office” was an undue burden. Vendors resist hard. Cybergenetics’ Mark Perlin has argued that TrueAllele’s roughly 170,000 lines of code would take “eight and a half years” to review. That’s an argument that doubles as an admission of how much unreviewed logic sits inside a single expert’s testimony.
United States v. Ortiz, 736 F. Supp. 3d 895 (S.D. Cal. 2024), exposed STRmix’s validation limits in stark terms. The software had never been validated for samples with six or more contributors, yet the analyst made a “judgment call” to set the count at five thereby forcing the sample into the software’s validated range. After the defense expert showed the mixture likely contained six contributors, the court excluded the STRmix evidence, invalidating the roughly 540-million-to-one likelihood ratio the software had produced.
That issue did not go away. In United States v. Lopez (D. Conn. 2025), a defendant again pressed the number-of-contributors question under Daubert, arguing the state lab’s own SOP forbade using STRmix above four contributors. That’s a sign that the Ortiz line of attack is now a live, recurring theme rather than a one-off.
The recurring pattern is easy to spot once you know it: the vendor invokes trade secret plus “too complex to review,” and the defense invokes the Sixth Amendment’s Confrontation Clause and Due Process. Courts split but the trend is slowly toward more access.
Reliability Gatekeeping — Daubert, Frye, and (in Texas) Kelly
Nationally, the Daubert and Frye standards are the usual standards: the proponent must show the method is reliable and (in Frye jurisdictions) generally accepted. AI strains both. A model’s output cannot always be explained, training-data bias may be undisclosed, and “general acceptance” is murky for a tool only the vendor fully understands. There is an active push in the bar and the academy (including proposals to amend the Federal Rules of Evidence for AI-generated evidence) and this is likely to be the primary battleground for the next several years.
For Texas cases, it is worth being precise: Texas does not apply Daubert or Frye by name. Reliability gatekeeping under Rule 702 runs through Kelly v. State, 824 S.W.2d 568 (Tex. Crim. App. 1992), for hard science, and Nenno v. State for softer disciplines. The Kelly framework asks whether the underlying scientific theory is valid, whether the technique applying it is valid, and whether the technique was properly applied on the occasion in question. That last prong maps almost perfectly onto the AI problem: it does not matter that STRmix is generally validated if it was never validated for this sample, run at this contributor count, in this version.
This Isn’t New: Breath, Blood, and Now Toxicology
If the “opaque software deciding guilt” problem sounds familiar, that is because DWI defense has been fighting it for years (long before probabilistic genotyping existed). The newest AI is now extending the same problem into forensic toxicology.
Breath alcohol: the original source-code fight
Breathalyzers are embedded computers. Firmware converts a raw infrared or fuel-cell signal into a BAC number, and the algorithms in between make consequential judgment calls: slope and breath-flow detectors deciding when a sample is “valid,” minimum breath-volume thresholds, radio-frequency-interference detection, and rules for averaging and rounding. None of that appears on the printed ticket the jury sees.
State v. Chun (N.J. 2008) had the New Jersey Supreme Court, through a special master, order extensive review of the Draeger Alcotest 7110 and impose detailed conditions on its use. This was an early landmark for looking inside the machine. In Massachusetts, Commonwealth v. Camblin established a defendant’s right to a reliability hearing on breath-test source code, leading to the consolidated Commonwealth v. Ananias litigation over the Alcotest 9510. In 2016 the court authorized defense review of the source code using both static testing (reading the code for defects) and dynamic testing (running it live on real instruments). That same litigation exposed serious problems at the Massachusetts Office of Alcohol Testing (methodology and calibration failures plus withheld exculpatory documents) and resulted in tens of thousands of breath tests being excluded. Independent code reviews of the Draeger 9510 later flagged defects including no correction for breath temperature (which can inflate a reading by roughly 6–8% per degree Celsius above normal), weak error handling, and rounding behavior that could mask problems.
Blood alcohol: peak-integration
The gold standard for blood alcohol is headspace gas chromatography (GC-FID, often with GC-MS confirmation). The reliability question here is subtle. The instrument software (Agilent ChemStation/OpenLab and similar) uses an automatic peak-integration algorithm to decide where each chromatographic peak starts and stops and to compute its area. That area is what becomes the reported BAC. Shift the baseline slightly and the number moves.
Analysts can accept the algorithm’s integration or manually redraw it (injecting subjectivity that can push a result over a per se limit) and the audit trail records whether re-integration occurred. Usually only the final number is produced. Obtaining the electronic raw chromatographic data and the audit trail would allow your expert to re-integrate independently. It is the direct analogue of demanding the STR electropherograms in a DNA case.
Drug analysis and toxicology: where genuine AI is arriving
Modern confirmatory identification in a tox lab relies on mass spectrometry (LC-MS/MS and GC-MS) matching a sample’s spectrum against a spectral library (the NIST mass-spectral library and vendor libraries) using a match-factor algorithm. That score is algorithmic, and the threshold and search parameters are rarely stated in the report. The genuinely new frontier is deep learning for novel psychoactive substances. When a compound is not in any library, deep-learning systems now predict MS/MS spectra and infer identity with tools such as PS2MS and related spectrum-prediction models published in 2023–2024. The risk is obvious: a black-box model “identifying” a drug that was never matched to a reference standard, then offered as though it were a confirmed identification. (Keep presumptive color/reagent field tests separate. Those are a well-known reliability problem, but not an AI one.)
What to Ask For
The discovery levers are consistent across every one of these tools: name the software, demand the raw data, and check whether the version and validation actually fit the sample in front of you.
- DNA: software name and version; internal and developmental validation studies; parameter settings and analyst judgment calls; run logs; the underlying STR data and electropherograms.
- Breath: instrument make, model, and firmware version; source code and any prior source-code review findings; calibration and certification records; the specific error, slope, and RFI settings.
- Blood: the electronic raw chromatographic data (not just the printout); the audit trail showing automatic versus manual integration; the lab’s integration SOP; calibration and control data.
- Drugs: the instrument and library used; the match-score threshold and search parameters; whether the compound was confirmed against a reference standard; and whether any predictive or machine-learning identification tool was used at any stage.
The Bottom Line
“AI in the forensic lab” is not one thing and treating it as one thing is how good defenses get lost. Sort the load-bearing tools (probabilistic genotyping, above all) from the upstream assists. Understand which standards supposedly govern them and where those standards run out. Then use discovery and the reliability rules (Kelly in Texas, Daubert and Frye elsewhere) to expose the gap between what the software decided and what anyone was ever allowed to check. That gap has been winning breath and blood cases for years. It is now available in DNA and toxicology cases too, for the lawyers willing to look.
Deandra Grant earned the ACS-CHAL Forensic Lawyer-Scientist designation. She holds an M.S. in Pharmaceutical Science and a Graduate Certificate in Forensic Toxicology, and she has spent three decades taking apart the science behind the State’s evidence. This post is part of the Deandra Grant Law forensic science series.
Further Reading
- Harvard Cyberlaw Clinic — Victory for Transparency in Probabilistic Genotyping (Pickett)
- The Markup — Powerful DNA Software Faces New Scrutiny
- ProPublica — Where Traditional DNA Testing Fails, Algorithms Take Over
- Texas DPS — Notification letter documenting the 2018 Travis County STRmix exclusion
- NBC News — A Texas jury convicted him; a DNA algorithm led to the real suspect (Lydell Grant)
- NIST / OSAC — Research and Development Needs
- NIST — OSAC Registry
- Quinn Emanuel — Adapting the Rules of Evidence for the Age of AI
- Maryland State Bar — Applying Daubert and Frye to AI Evidence
- Commonwealth v. Ananias — Alcotest 9510 source-code appeal (DelSignore)
- CBS News — Breathalyzer source-code flaws cast doubt on convictions
- Barone Defense — Gas chromatography blood-alcohol raw data and defenses
- Machine Learning in Forensic Toxicology — review (ScienceDirect, 2025)
- PS2MS — deep-learning NPS identification (ACS Analytical Chemistry)
This post is an informational synthesis for educational purposes and is not legal advice. Case citations should be independently verified against the official record before use in any filing.