Skip to main content
< All Topics
Print

Sleuth Report: Confession, I Accidentally Published Victim Names

Overview

A major investigative report by Rye Howard Stone documents how the creator of the public research database Epstein Data discovered that thousands of pages from the DOJ Epstein Files release contained unredacted victim identifying information. The report explains how the database unintentionally mirrored those records and how the issue was eventually corrected using targeted searches, AI assisted review, and large scale redaction methods. (rhowardstone.substack.com)

The article has become one of the most discussed investigative posts surrounding the DOJ Epstein Files release because it directly challenges government claims that removing tens of thousands of documents was necessary for victim protection. (rhowardstone.substack.com)

What Happened

According to the report, the database hosted OCR text from approximately 3.5 million DOJ Epstein Files pages that had been released publicly in January 2026. The creator later discovered that some of those files included:

• Real victim names
• Phone numbers
• Dates of birth
• Partial Social Security numbers
• Home addresses

The report states that these details had already been published by the DOJ itself and mirrored across multiple research archives online. (rhowardstone.substack.com)

The author explains that the discovery happened while reviewing removed DOJ documents connected to broader analysis of the missing files controversy.

The Redaction Failure Problem

The article describes widespread failures in the DOJ redaction process. According to the report:

• Victims’ attorneys identified numerous exposed records
• Some documents allegedly exposed multiple minors
• Sensitive information remained searchable online
• Public mirrors inherited the same exposure problems

The report claims that the issue was not isolated but spread across hundreds of files and thousands of searchable entries. (rhowardstone.substack.com)

How the Cleanup Was Performed

The report provides a detailed explanation of how the cleanup process worked.

Initial Failures

The first approaches reportedly failed because OCR text created massive detection errors. Early attempts included:

• Regex pattern matching
• Automated entity recognition
• Broad metadata scanning

These approaches created high false positive rates and damaged unrelated records. (rhowardstone.substack.com)

The Breakthrough

The author explains that the process only became manageable after identifying actual victim names from contaminated documents. Once those names were known, the database could be searched systematically.

According to the report:

• A corpus wide scan completed in roughly 95 seconds
• More than 4,000 matches were identified
• Over 400 affected documents were reviewed
• Multiple passes of AI assisted redaction were performed

The final result reportedly removed identifiable victim information without deleting entire documents from public access. (rhowardstone.substack.com)

Criticism of the DOJ Response

One of the most significant claims in the article involves the DOJ decision to remove approximately 67,784 documents from public access after the redaction failures were discovered.

The report argues that:

• Most removed files may not have contained victim information
• Targeted redaction would have been more appropriate
• Entire files were removed instead of repaired
• The government had resources far beyond what independent researchers possessed

The author claims that one individual with AI tools and search indexing was able to complete targeted redactions in a single day, while the DOJ removed entire collections instead. (rhowardstone.substack.com)

Why This Matters

The report highlights a growing issue facing large public evidence archives:

Transparency vs Privacy

Investigative archives involving trafficking, abuse, and criminal networks require balancing public accountability with survivor protection. This article demonstrates how difficult that balance becomes when millions of scanned pages are involved.

OCR and Searchable Archives Create New Risks

The report also illustrates how OCR searchable databases can unintentionally expose sensitive data at scale. Even when information appears buried inside PDFs, searchable indexing can instantly surface personal information across massive archives.

Independent Investigators Are Becoming Major Archivists

The article reflects how independent researchers and open source investigators are increasingly managing document systems that rival institutional archives in scale and accessibility.

Key Takeaways

Large Scale Evidence Releases Require Better Redaction Systems

The report suggests that traditional review systems may not be sufficient when millions of pages are released publicly.

AI Tools Can Both Create and Solve Problems

The article demonstrates how AI assisted tools can help locate sensitive material quickly but also introduce new challenges when dealing with noisy OCR data.

Public Trust Depends on Accuracy

The report repeatedly stresses that transparency, careful sourcing, and responsible handling of evidence are essential for maintaining credibility in large public investigations.

Sources

Previous Sleuth Report: Celina Dubin’s Colorado Wedding Renews Questions About the Dubin Family’s Long Relationship With Jeffrey Epstein
Next Sleuth Report: Daniel Siad’s Epstein Recruitment Emails and Pro Trump Account
Table of Contents