Skip to main content
< All Topics
Print

Sleuth Report: The Missing 641,502 Pages: A Forensic Audit Challenges the DOJ Epstein Files Count

The Department of Justice announced on January 30, 2026, that it had published nearly 3.5 million pages under the Epstein Files Transparency Act. A new forensic page count audit finds that the publicly numbered production cannot support that figure.

The audit follows the Department’s own Bates numbering system to its final page. That sequence ends at EFTA02858498, creating a maximum possible total of 2,858,498 uniquely numbered EFTA pages. The difference between that ceiling and the government’s public claim is approximately 641,502 pages.

This finding does not prove that 641,502 pages were unlawfully removed. It does establish that the Justice Department has not published a reproducible accounting that explains what it counted as part of its reported 3.5 million pages.

Key Takeaways

  • The Justice Department claimed that it published nearly 3.5 million responsive pages.
  • The final EFTA Bates stamp is EFTA02858498, limiting the numbered production to no more than 2,858,498 pages.
  • The Epstein Data archive contains 2,858,311 EFTA pages, which represents nearly the entire possible Bates sequence.
  • Adding House Oversight records, FBI Vault files, estate productions, and other public collections raises the documented total to approximately 2,923,763 pages.
  • The combined public collection remains approximately 576,237 pages below the Justice Department’s announced total.
  • The audit found 23 Data Set 9 entries that are empty, broken, or limited to placeholder pages.
  • The 187 page difference between the Bates ceiling and the archived EFTA count is explained by blank entries, placeholder pages, and unused numbers.
  • Independent archives and news organizations have reported totals between approximately 2.5 million and 2.9 million pages.
  • No independent public recount cited by the audit reaches 3.5 million accessible pages.
  • The Justice Department has not released a complete accounting of duplicates, withheld pages, nonresponsive material, sealed records, privileged records, or files removed after publication.

What the Justice Department Claimed

On January 30, 2026, the Justice Department issued a press release announcing that it had published nearly 3.5 million responsive pages in compliance with the Epstein Files Transparency Act.

The Department said that the new release included more than 3 million additional pages, more than 2,000 videos, and approximately 180,000 images. Attorney General Pam Bondi and Deputy Attorney General Todd Blanche repeated the cumulative 3.5 million page figure in a letter to Congress.

However, the Department did not provide a public manifest explaining how it calculated that total. It did not specify whether it counted duplicate pages, withheld material, videos, photographs, native computer files, or pages that were reviewed but never released.

That omission makes the announced figure impossible for the public to reproduce.

The Bates Numbers Create a Verifiable Ceiling

Bates numbers are sequential identifiers applied to pages in a legal production. Each numbered page receives its own unique stamp. This makes it possible to determine the highest possible number of pages within a production.

The final document identified in the EFTA release is EFTA02858497. That document contains two pages. Therefore, the final page is numbered EFTA02858498.

This creates a maximum ceiling of 2,858,498 numbered pages.

The Epstein Data forensic audit reports that its archive contains 2,858,311 of those possible pages. Only 187 numbers separate the archived total from the theoretical ceiling.

The audit accounts for that small difference through blank entries, placeholder pages, and numbers that were apparently never assigned within portions of the court records production.

The central conclusion is simple. A Bates sequence ending at 2,858,498 cannot contain 3.5 million uniquely numbered pages.

What the Public Archive Actually Contains

The audit provides several separate totals because documents, pages, extracted records, and native files are not interchangeable units.

The Epstein Data database contained 1,425,565 document records from all included sources as of July 31, 2026. Those records represented 2,923,763 document metadata pages.

Within the numbered EFTA production, the database contained 1,393,164 document records and 2,858,311 pages.

The database also contained 2,924,310 extracted page records. The audit explains that this slightly larger figure results from the way spreadsheets, videos, and other native files are converted into searchable records. Those additional search records are not additional public PDF pages.

The distinction matters because several news headlines inaccurately described the Justice Department release as millions of documents or records. The Department’s official claim referred to pages, not documents.

The Public Totals Still Do Not Reach 3.5 Million

The audit expanded its count beyond the EFTA production. It included records from the House Oversight Committee, the Justice Department Office of Government Relations, the FBI Vault, estate productions, and related public sources.

That expanded collection contained approximately 2,923,763 pages.

Even after those materials were added, the public collection remained approximately 576,237 pages below 3.5 million.

The adjacent productions add approximately 65,000 pages. They do not supply the more than 641,000 pages needed to reconcile the EFTA Bates ceiling with the Justice Department’s announcement.

The Twenty Three Blank Data Set 9 Entries

The audit compared its archive with the Justice Department’s own document manifest and identified 23 Data Set 9 entries that did not contain ordinary documents.

Twenty entries reportedly returned empty files. One returned a small broken HTML fragment. Two contained pages stating that no images had been produced for underlying native material.

The affected entries include the following records:

The audit states that the entries were checked against Justice Department downloads, an independent mirror, and archived copies. The supporting Data Set 9 blank slot inventory records the affected Bates numbers and the file sizes served by the government website.

These entries raise legitimate questions about the underlying native files. However, they cannot explain a difference of more than 641,000 pages.

The Entire 187 Page Difference Is Accounted For

The Bates ceiling contains 187 more numbers than the 2,858,311 EFTA pages stored in the archive.

According to the audit, 21 numbers correspond to empty or broken entries. Two correspond to placeholder pages. Approximately 164 numbers were never assigned within the less densely numbered court records portion of the production.

This accounting is important because it prevents the 187 number gap from being misrepresented as proof that 187 substantive documents disappeared.

A gap in Bates numbering can identify something that requires investigation. It does not establish that a document was created, withheld, or removed.

Independent Counts Support a Lower Public Total

The audit compared its findings with several independent archives and reporting projects.

CBS News reportedly developed software to crawl the Justice Department library and counted approximately 2.7 million publicly accessible pages. CBS also found that more than 47,000 files, representing approximately 65,500 pages, had been removed or taken offline after publication.

The Justice Department disputed the CBS analysis. However, it also acknowledged that more than 47,000 files remained offline for further review.

Jmail reported approximately 2,474,242 pages across more than 1.4 million files within the scope of its archive. Other projects reported roughly 1.4 million file records, although their methods and source collections differed.

These totals cannot be treated as perfectly interchangeable. Archives count source files, extracted records, emails, attachments, native files, and pages differently. Some also share processed source material, meaning they are not always fully independent confirmations.

Even with those limitations, every public count cited in the audit falls below 3 million pages. None reaches the Department’s announced figure.

The Printed Reading Room Also Falls Short

The Institute for Primary Facts created a physical reading room containing printed copies of the public Epstein files.

According to the audit author’s onsite examination, the collection was organized into 3,437 volumes with a maximum capacity of 800 pages per volume.

Multiplying those figures produces a maximum possible capacity of 2,749,600 pages. Since individual volumes may contain fewer than 800 pages, the actual printed total could only be lower.

The physical collection therefore supports a production closer to 2.8 million pages than 3.5 million pages.

The Public Collection Continues to Change

The audit also documents a serious preservation problem. The Justice Department’s online collection is not static.

Some documents have moved between data sets without public notice. Files that once appeared at one address may now appear elsewhere. A broken link can therefore represent a relocation rather than a deletion.

The Justice Department has acknowledged that more than 47,000 files remained offline for additional review. Independent monitoring projects have also recorded files changing or disappearing.

The audit separately identifies 503 House Oversight pages that were reportedly produced but never posted publicly. Those records run from HOUSE_OVERSIGHT_009974 through HOUSE_OVERSIGHT_010476.

The supporting House Oversight unposted pages inventory provides a machine readable list of those records.

The changing collection makes independent preservation essential. It also means that page totals must include a capture date and a precise definition of what was counted.

Six Million Pages Identified Does Not Mean Six Million Pages Released

The Justice Department reportedly identified more than 6 million pages as potentially responsive during its review. Reuters separately reported that approximately 5.2 million pages were scheduled for review.

Those figures describe a review pool, not a public release.

A potentially responsive collection may include duplicates, unrelated records, privileged communications, sealed materials, victim identifying information, and records withheld under statutory exceptions.

The problem is not simply that the released collection contains fewer pages than the review pool. The problem is that the Department has not published a complete reconciliation showing how the review pool became the final public production.

A transparent accounting would identify the number of pages reviewed, removed as duplicates, classified as nonresponsive, withheld under each legal exception, placed under seal, removed for victim protection, and ultimately published.

Without that accounting, the public cannot determine what the 3.5 million figure actually represents.

The Phang Litigation Adds Legal Significance

The page count dispute exists alongside continuing litigation over compliance with the Epstein Files Transparency Act.

In Phang v. Blanche, United States District Judge Emmet Sullivan made a preliminary finding concerning a limited group of claims that the Justice Department had not substantively answered. The court ordered the Department either to release specified material without the disputed redactions or explain its legal basis for withholding it.

The court later directed the Department to submit unredacted versions for private judicial review. The litigation remains unresolved, and preliminary findings do not establish that the entire numerical gap represents unlawful withholding.

The case still matters because it demonstrates that the disclosure dispute is not limited to competing database totals. Courts are examining whether specific records and redactions comply with the law.

What the Audit Proves

The audit documents four central facts.

First, the EFTA Bates sequence ends at 2,858,498.

Second, the Epstein Data archive contains 2,858,311 EFTA pages.

Third, all public sources included in the database total approximately 2,923,763 pages.

Fourth, the Justice Department has not released a detailed calculation that allows the public to reproduce its 3.5 million page claim.

Those findings establish a substantial and unresolved counting discrepancy.

What the Audit Does Not Prove

The audit does not prove that 641,502 pages were secretly deleted.

It does not prove that every page within the potentially responsive review pool had to be released.

It does not establish that every missing or offline file contains evidence of criminal conduct.

It also does not determine whether the government counted duplicates, native files, images, videos, withdrawn files, or withheld pages when calculating its public total.

Those possibilities remain hypotheses until the Department releases its counting methodology.

Questions Congress Should Ask Under Oath

Congress should ask Attorney General Pam Bondi and Deputy Attorney General Todd Blanche to define precisely what the Department counted as a page.

They should be required to identify the source database used to generate the 3.5 million figure and provide the calculation behind it.

Congress should also ask whether the figure included duplicates, nonresponsive records, sealed records, privileged material, images, videos, native files, or pages that were reviewed but never made public.

The Department should disclose how many files were removed after publication, how many remain offline, and whether those files were included in the announced total.

Finally, the Justice Department should provide a complete reconciliation connecting the potentially responsive review pool, the final Bates sequence, the files currently available online, and every category of withheld material.

The public should not have to reverse engineer a federal disclosure from page stamps while the Department continues defending a number it has never shown its work for.

Why This Audit Matters

The Epstein files concern institutional failures, trafficking allegations, public corruption questions, survivor testimony, and the conduct of politically powerful people. Reliable accounting is therefore part of the evidence itself.

A government agency cannot claim transparency merely by publishing an enormous collection. It must also explain what was reviewed, what was released, what was withheld, what was later removed, and how its public totals were calculated.

The Epstein Data page count audit does not answer every question about the missing total. It does something equally important. It identifies the exact question the Justice Department has avoided answering.

What, precisely, did the Department count to reach 3.5 million pages?

Evidence and Research Tools

The Epstein Data full text corpus allows researchers to search the public collection and reproduce the database counts.

The page level archive contains the searchable page records used in the audit.

The native files catalog identifies videos, audio files, spreadsheets, and related placeholder pages.

The missing EFTA methodology explains how the audit examined apparent gaps in the Bates sequence.

Readers can explore additional background records through the EpsteinWiki Knowledge Base and search the primary evidence collection through Epstein Data.

Sources

Previous Sleuth Report: The Millions Jeffrey Epstein Directed to Darren Indyke and Richard Kahn
Next Sleuth Report: The Missing Jane Doe 4 FBI Notes: What Was Found and What Remains Unverified
Table of Contents