Skip to main content
< All Topics
Print

Epstein Data: An Open Source Research Database For The Epstein Files

Epstein Data is a free, independently operated research database built from the public Epstein files. It combines full text document search, artificial intelligence research tools, image analysis, transcripts, case indexes, relationship maps, evidence tracking, and direct EFTA source pages.

The platform says its full text corpus contains approximately 1.43 million documents and 2.9 million pages from twelve Department of Justice datasets. It also indexes photographs, audio and video transcripts, spreadsheets, handwritten material, redaction data, case records, and other structured research collections.

Unlike a closed commercial search service, Epstein Data publishes source databases and research materials through public repositories. It also provides instructions for downloading the research databases and running the investigative tools locally.

Researchers can access the platform at Epstein Data.


Snapshot

Name: Epstein Data

Website: Epstein Data

Operator: Independent research project associated with the public R. Howard Stone repositories and podcast

Access: Free

Source model: Open source databases with direct EFTA evidence pages

Core corpus: Approximately 1.43 million documents and 2.9 million pages

Main production: Twelve Department of Justice datasets released under the Epstein Files Transparency Act

Additional collections: House Oversight records, FBI Vault material, United States Virgin Islands records, Department of Justice Office of Professional Responsibility material, and other public sources

Minimum age: The site requires users to confirm that they are at least 18 years old

Important warning: Analytical reports and summaries may contain errors and must be verified against linked source documents


What Is Epstein Data?

Epstein Data is a research infrastructure project rather than a single search box. It organizes the public Epstein files into databases that can be searched, compared, mapped, and downloaded.

The site is designed around EFTA identifiers. These identifiers allow researchers to move from a search result or analytical claim to a specific document page.

The project states that it is independent and is not affiliated with the Department of Justice, FBI, any government agency, or Anthropic. It also explains that its analytical reports use Claude and are repeatedly checked against source documents, while warning that errors may remain.

This distinction matters. Epstein Data contains both original public records and analytical layers created from those records. Researchers must identify which one they are reading.


The Full Text Corpus

The Epstein Data full text corpus is the foundation of the platform. It makes scanned and digital documents searchable by name, phrase, EFTA number, and other terms.

The homepage reports approximately 1.43 million documents and 2.9 million pages. The corpus includes material from twelve Department of Justice datasets.

Researchers can search exact phrases by placing words inside quotation marks. They can also search directly for an EFTA identifier, such as EFTA00074206.

Full text search speeds up discovery, but optical character recognition can misread scanned text. Names, dates, telephone numbers, account numbers, and handwritten notes should always be checked against the document image.


Direct EFTA Evidence Pages

One of the platform’s most important features is the use of permanent evidence pages for individual EFTA records. These pages allow researchers and readers to review the source connected to a published claim.

Examples include EFTA00128780, EFTA01433023, EFTA01578684, EFTA01139414, and EFTA00151067.

An EFTA page is more useful than a general archive link because it brings the reader directly to the relevant record. This supports transparent review and makes corrections easier when an interpretation is disputed.

The EFTA identifier still does not explain the document by itself. Researchers must determine the file type, author, date, context, and evidentiary value.


Artificial Intelligence Research Assistant

The Epstein Data research assistant accepts plain language questions and searches across several databases. The interface says it can synthesize findings from approximately 1.4 million documents while citing specific records.

The assistant searches the full text corpus, redaction analysis, image descriptions, audio transcripts, and other structured collections. This can help users locate relevant EFTA records without knowing the exact language used in a document.

The platform states that chat history is stored in the user’s browser and is not saved on its servers unless the user chooses to share it. It also warns that queries are processed by outside artificial intelligence services and tells users not to enter personal information.

Generated summaries can contain errors. Researchers should use the assistant to locate evidence, then read and cite the evidence itself.


FBI Evidence Map

The FBI Evidence Map reconstructs serial numbers used in FBI investigations connected to Jeffrey Epstein. It compares evidence indexes with documents found in the released corpus.

The tool identifies records as found, missing, or illegible. It draws from several source indexes, including Non Testifying Witness Material, Testifying Witness Material, interview records, Sentinel serial reports, and physical evidence logs.

According to the methodology displayed on the page, the project searches the full corpus for FBI serial references through text matching and optical character recognition of secondary stamps.

The map reports a resolution rate of approximately 29 percent and states that more than two thirds of indexed material was not located in the public release. Missing means the project did not find the indexed record within the reviewed corpus. It does not prove why the material was absent or whether another copy exists under a different identifier.


Reverse Image And Face Search

The reverse image search allows users to upload a photograph and search for visually similar images or faces across indexed collections.

The stated scope includes the twelve Department of Justice datasets, House Oversight records, United States Virgin Islands material, FBI Vault records, and Department of Justice Office of Professional Responsibility material.

The tool uses CLIP for visual similarity and InsightFace for facial comparison. Similarity scores can help locate the same image, the same scene, or visually related material.

A similarity result is not an identification. Facial comparison can produce false matches, especially with poor resolution, unusual angles, sunglasses, partial faces, or similar looking people.

The site says uploaded images are held in memory only during the request and are not written to disk. It says embeddings are not retained. Timestamp, IP address, query type, and result count are logged for abuse prevention.

The terms prohibit stalking, harassment, and exposing private individuals. The site also states that the tool cannot reverse Department of Justice redactions protecting victims.


Document Coverage Map

The Document Coverage Map visualizes Bates stamp ranges across the EFTA corpus. Each cell represents a block of sequential identifiers, while color intensity shows how many expected stamps are present.

This can help researchers locate gaps and understand how the release is distributed across the EFTA number range.

A grey cell does not automatically represent a missing document. EFTA numbers can mark pages inside a multiple page file, dataset boundaries can create apparent gaps, and native files may have PDF placeholder pages.

Any suspected gap should be checked against production metadata, concordance files, page counts, Department of Justice links, and archived copies.

The Missing EFTA Document Analysis explains the difference between raw number gaps and page based document gaps.


Images, Transcripts, And Native Files

Epstein Data provides separate collections for photographs, image descriptions, audio and video transcripts, spreadsheets, native files, email metadata, handwritten material, and page classifications.

The homepage reports 92,000 analyzed photographs and 435 transcribed recordings. Its reverse image tool searches a larger image index that includes extracted pages and material from additional sources.

These counts refer to different analytical collections and should not be treated as interchangeable. One database may count photographs selected for detailed analysis, while another counts every indexed visual item or page image.

Automated image descriptions and transcripts require verification. Researchers should inspect the original image or media file and cite timestamps when using audio or video evidence.


Network And Entity Research

The Epstein Data Network combines several evidence layers into a relationship graph. The site also maintains an entity directory and structured knowledge graph.

These tools can help researchers identify repeated people, organizations, locations, communications, and documentary associations.

A graph connection does not prove a social relationship or criminal participation. It may reflect an email reference, shared document, address book entry, court allegation, photograph, or another documentary link.

Researchers should open the evidence supporting each connection. EpsteinWiki’s article on mapping the Epstein network provides additional guidance for interpreting network data responsibly.


Case And Deposition Collections

The platform organizes records from 53 federal and state cases and provides individual case briefings. It also offers searchable deposition transcripts synchronized with publicly available audio or video when possible.

This structure helps users separate records by proceeding rather than mixing every allegation into one general archive.

Court records require precise labeling. A complaint contains allegations. A deposition records sworn testimony. An exhibit may be admitted for a limited purpose. A judicial ruling establishes what a court decided, not necessarily the truth of every statement quoted within it.

Researchers should identify the case, court, filing date, document type, and procedural status before summarizing legal material.


Survivor Centered Resources

The Hear From The Survivors collection centers public interviews and statements from survivors. The homepage identifies 19 survivors represented in the collection.

This resource is important because large document archives can reduce people to names, identifiers, or references inside investigative files. Survivor testimony provides context that database records alone cannot supply.

The platform displays an adult content warning because the archive includes official records involving sexual abuse, trafficking, and exploitation of minors.

Researchers should use survivor testimony with care, avoid unnecessary repetition of graphic details, and never attempt to identify protected victims.


Investigation Reports

The Epstein Data report library contains audits, technical studies, entity research, network analysis, and methodological reports.

The platform states that analytical text is generated with Claude and iteratively checked against source documents. It also clearly warns that reports may contain errors.

This disclosure is a strength because it distinguishes analysis from primary evidence. However, the warning must be carried into downstream reporting. A report can identify useful patterns, but its conclusions should be reproduced only after reviewing the cited records.


Open Source Data And Local Research

The project’s public data repository provides structured exports from its research work. This includes entity data, image catalogs, EFTA mapping, and other analytical resources.

The Run Your Own AI Investigator guide explains how users can download the databases and research tools to a local computer. The graphical walkthrough estimates approximately 17 gigabytes of free storage and uses Claude Code with a paid Claude subscription.

Local access gives researchers more control over their copies and workflows. It also makes the methodology easier to inspect and reproduce.

The research materials are published under a CC BY NC SA 4.0 license according to the site. Users should review the license before republishing databases or using them commercially.


Tools For Developers

Epstein Data offers a public REST interface, an OpenAPI specification, an MCP server, and an instructions file for artificial intelligence agents.

These tools allow developers to connect research applications directly to the public databases. They also make it possible to build independent interfaces, verification tools, and automated research workflows.

Automated access does not remove the need for evidentiary context. A technically correct database response can still be misunderstood if the underlying record contains an allegation, duplicate, transcription error, or incomplete page.


Privacy And Ethical Limits

The Epstein Data privacy and terms page describes how the platform handles logs, uploads, and other technical data.

The research assistant says chat history remains in the browser unless a user shares it. However, prompts are processed through outside artificial intelligence services. Users should not enter personal or confidential information.

The reverse image tool says uploaded images remain in server memory only during processing and that embeddings are not retained. It still logs limited request information for abuse prevention.

Researchers should never use the tools to identify protected victims, stalk individuals, expose private people, or spread unsupported allegations. Public access to a record does not eliminate ethical responsibilities.


How To Use Epstein Data Responsibly

A reliable research process should preserve the distinction between discovery, analysis, and evidence.

  1. Begin with a name, phrase, EFTA number, image, case, or research question.
  2. Use search tools to locate relevant records.
  3. Open the direct EFTA evidence page.
  4. Read the complete document and surrounding pages.
  5. Identify the source type and procedural context.
  6. Confirm names, dates, quotations, and numbers manually.
  7. Review duplicates and related documents.
  8. Separate verified facts from allegations and analytical inference.
  9. Link the final publication directly to the evidence page.
  10. Correct the public record when new evidence changes the conclusion.

How Epstein Data Supports EpsteinWiki

Epstein Data provides the evidence layer needed for transparent EpsteinWiki reporting. Its direct EFTA pages allow readers to move from a Knowledge Base claim to the underlying source.

The structured databases also support EpsteinWiki research on how Jeffrey Epstein’s business and trafficking system worked, Jeffrey Epstein’s companies and financial infrastructure, mapping the Epstein network, and institutional failures in the Epstein case.

EpsteinWiki provides narrative context. Epstein Data provides searchable records, analytical tools, and evidence citations. Used together, they allow readers to examine both the explanation and the receipt.


Key Takeaways

  1. Epstein Data is a free and independently operated research database for the public Epstein files.
  2. Its full text corpus contains approximately 1.43 million documents and 2.9 million pages.
  3. The platform links research findings directly to individual EFTA evidence pages.
  4. Its tools include full text search, artificial intelligence assistance, reverse image search, relationship graphs, transcripts, case collections, and evidence tracking.
  5. The FBI Evidence Map compares indexed investigative material with records found in the public release.
  6. The project publishes structured data and research materials through public repositories.
  7. Researchers can download the databases and run investigative tools locally.
  8. Artificial intelligence reports, transcripts, image descriptions, and relationship maps can contain errors.
  9. A search result, graph connection, face match, or document mention does not prove wrongdoing.
  10. Survivor privacy and protection remain essential throughout the research process.
  11. The strongest Epstein Data citation is a direct link to the relevant EFTA evidence page.
  12. Every important conclusion should be independently verified against the original record.

Why Epstein Data Matters

The Epstein files are not truly accessible if the public must manually open millions of pages without reliable indexes, searchable text, or direct citations.

Epstein Data converts the release into a research system. It helps users search documents, examine images, follow EFTA identifiers, compare case records, inspect network claims, and audit what the government did or did not release.

Its greatest value is not artificial intelligence. Its greatest value is the path from a research question to a specific piece of evidence that another person can open and evaluate.


Sources

  1. Epstein Data
  2. Epstein Data Research Assistant
  3. Epstein Data Full Text Corpus
  4. Epstein Data FBI Evidence Map
  5. Epstein Data Reverse Image Search
  6. Epstein Data Coverage Map
  7. Epstein Data Network
  8. Epstein Data Investigation Reports
  9. Epstein Data Missing EFTA Document Analysis
  10. Run Your Own AI Investigator
  11. Epstein Research Data Repository
  12. Epstein Data Privacy And Terms
  13. United States Department of Justice Epstein Library
  14. Epstein Files Transparency Act
  15. Epstein Data evidence record EFTA00074206
  16. Epstein Data evidence record EFTA00128780
  17. Epstein Data evidence record EFTA01433023
  18. Epstein Data evidence record EFTA01578684
  19. Epstein Data evidence record EFTA01139414
  20. Epstein Data evidence record EFTA00151067
Previous DocETL Epstein Email Archive Explorer: AI Analysis of 2,322 Released Epstein Emails
Next Epstein Exposed Flight Log Database
Table of Contents