Epstein-Data
Epstein Data is a free, independently operated, open source research platform that makes millions of pages from the publicly released Epstein files searchable and easier to examine. As an EpsteinWiki research and evidence partner, Epstein Data provides the document infrastructure that allows contributors and readers to move from an EpsteinWiki claim directly to the underlying EFTA record.
Snapshot
| Detail | Information |
|---|---|
| Name | Epstein Data |
| Website | epstein-data.com |
| Project type | Independent public interest research database |
| EpsteinWiki relationship | Research and evidence partner |
| Cost | Free public access |
| Accounts required | No account or login required for ordinary research |
| Core collection | Approximately 1.4 million documents and 2.9 million pages |
| Main source | Publicly released United States Department of Justice Epstein files |
| Document groups | 12 Department of Justice datasets, plus additional indexed collections |
| Search capabilities | Full text, exact phrase, EFTA number, images, entities, transcripts, and structured data |
| Research tools | Artificial intelligence research assistant, reverse image search, flight index, network graph, geographic map, document viewer, and evidence map |
| Open source access | Structured data and research repositories are publicly available through GitHub |
| Government affiliation | None |
| Required caution | Automated text, entity matches, image descriptions, and analytical reports must be verified against original records |
What Is Epstein Data?
Epstein Data is a searchable research platform built to organize the enormous volume of material released in connection with Jeffrey Epstein, Ghislaine Maxwell, related investigations, court proceedings, and government document productions.
The Epstein Data homepage describes the collection as containing approximately 1.4 million documents and 2.9 million pages across 12 Department of Justice datasets. The collection includes scanned documents, photographs, spreadsheets, emails, audio, video, transcripts, investigative records, court materials, and other public records.
The platform converts a difficult collection of separate files into searchable databases. Researchers can search for a name, organization, location, phrase, EFTA identifier, aircraft, financial term, or other research lead.
Epstein Data is independently operated. It is not affiliated with the United States Department of Justice, Federal Bureau of Investigation, any other government agency, or Anthropic.
Collection totals and tool counts may change as records are added, corrected, reprocessed, or classified.
The EpsteinWiki and Epstein Data Partnership
EpsteinWiki recognizes Epstein Data as a research and evidence partner because the two projects serve complementary purposes.
Epstein Data provides large scale document search, direct document access, structured datasets, analytical tools, and technical infrastructure. EpsteinWiki provides researched profiles, timelines, evidence summaries, legal context, editorial review, source explanations, and connections between people, organizations, places, cases, and documents.
The partnership creates a practical research path:
- A researcher discovers a document or pattern through Epstein Data.
- The original document is opened and reviewed.
- The EFTA identifier, page, context, and limitations are recorded.
- EpsteinWiki contributors compare the record with court filings, testimony, public records, and other evidence.
- The verified information is incorporated into an EpsteinWiki article.
- Readers can follow the citation back to the underlying document on Epstein Data.
The relationship is also visible through the Epstein Data Community Resources directory, which identifies EpsteinWiki as a collaborative knowledge base covering people, documents, court cases, and evidence.
Partnership does not mean that every automated result, report, or interpretation is automatically adopted by EpsteinWiki. EpsteinWiki applies its own investigative, privacy, sourcing, and editorial standards before publication.
Why Epstein Data Matters
The public Epstein document collection is too large for conventional page by page review alone. Important references may appear in duplicate records, attachments, handwritten notes, forwarded emails, transcripts, spreadsheets, or poorly scanned pages.
Epstein Data helps researchers:
- Locate records across millions of pages
- Search exact phrases and name variations
- Resolve EFTA identifiers
- Compare duplicate documents
- Find references across different datasets
- Review images and extracted image descriptions
- Search audio and video transcripts
- Examine flight records
- Explore geographic references
- Identify possible relationships for further verification
- Trace records cited in investigative reports
- Download structured data for independent analysis
The platform does not eliminate the need for human research. It makes human review possible at a scale that would otherwise be extremely difficult.
The Document Collection
The Full Text Corpus contains approximately 1.43 million indexed documents and 2.9 million pages.
The source collection includes material from the Department of Justice production released under the Epstein Files Transparency Act. Epstein Data also indexes other relevant public collections, including House Oversight materials and FBI Vault records where available.
The platform’s open source data repository describes its structured exports as covering all 12 Department of Justice datasets, plus House Oversight estate materials and FBI Vault records.
The corpus includes several types of information:
- Original document images
- Optical character recognition text
- Native digital files
- Email metadata
- Scanned text extracts
- Handwriting transcriptions
- Audio and video transcripts
- Spreadsheet data
- Photograph analysis
- Document classifications
- Redaction analysis
- Entity extractions
- Related document matches
- EFTA statements
- Report citations
These information layers serve different purposes. An original document image is evidence. Optical character recognition, an entity extraction, an image description, or an artificial intelligence summary is an aid for locating and interpreting that evidence.
EFTA Document Access
EFTA identifiers are production or Bates numbers assigned to records in the public document release.
A typical identifier appears as:
EFTA00074206
Epstein Data allows direct document access using the identifier:
The EFTA00074206 document page is an example of the direct link format.
A direct EFTA page may provide:
- The source document
- Page images
- Extracted text
- Document metadata
- Download access
- Dataset information
- Related records
- Searchable content
An EFTA number does not certify that every statement in the document is accurate. It identifies the produced record. Researchers must still determine who created the document, what type of record it is, what it says, whether it is complete, and how much evidentiary weight it deserves.
Required EpsteinWiki Citation Method
Every EpsteinWiki claim based on an EFTA record should link directly to the relevant Epstein Data document.
Use:
Do not use only:
- The Epstein Data homepage
- A general search page
- An entity page
- An artificial intelligence response
- An extracted database row
- An investigation report
- A screenshot without the original document
- A social media post discussing the record
The citation should include:
- Complete EFTA number
- Brief document description
- Relevant page
- PDF page when different
- Printed page when available
- Document date
- Source or author when known
- Necessary evidentiary limitation
Place the direct receipt near the claim it supports. Include important direct EFTA links again in the Sources section.
Search Tools
The Epstein Data search system allows researchers to search by name, keyword, exact phrase, or document identifier.
Keyword Search
A keyword search can locate pages containing a term, surname, company, property, address, or subject.
Search common variations, including:
- Full name
- Surname
- Initials
- Alternate spelling
- Former name
- Nickname
- Company name
- Email address
- Telephone number
- Street address
- Aircraft registration
- Trust or corporate entity
A missing search result does not prove that the information is absent. Optical character recognition may fail when documents contain handwriting, poor scans, unusual fonts, damaged pages, or heavy redaction.
Exact Phrase Search
Quotation marks can be used to search for an exact phrase:
“wire transfer”
Exact phrase searching can help locate repeated wording, duplicate correspondence, legal language, addresses, or distinctive quotations.
EFTA Search
Researchers can enter a complete EFTA identifier to locate a specific document.
Always compare the identifier digit by digit. A single incorrect number may lead to a different record or no result.
Database Search
The Explore Databases section allows researchers to search individual collections, including full text, redaction analysis, extracted entities, photographs, audio transcripts, and optical character recognition text.
Results from one database should be compared with results from other relevant collections.
Ask AI Research Assistant
The Epstein Data Research Assistant allows researchers to ask questions in plain language. The assistant searches across the indexed databases and returns a synthesized response with document citations.
The tool can help:
- Develop search terms
- Locate possible EFTA records
- Identify names or entities for review
- Compare references across databases
- Find relevant document clusters
- Generate preliminary research paths
- Summarize large result sets
- Suggest related records
The assistant is a discovery tool, not an evidentiary source.
Epstein Data states that artificial intelligence summaries may contain errors, hallucinations, or misinterpretations. Every claim must be verified against the linked source document before it is cited or published.
Do not enter personal, private, confidential, or survivor identifying information into the research assistant. Queries are processed through outside artificial intelligence services.
Reverse Image Search
The Reverse Image Search tool allows a researcher to upload an image and search for visually similar material across the indexed image collection.
It can help locate:
- Another copy of a photograph
- The source document containing an image
- A wider version of a cropped photograph
- Similar locations or scenes
- Repeated evidence photographs
- Possible appearances of the same person
- Related image sequences
A visual match is a lead. It is not automatic proof that two people, places, or objects are identical.
Researchers should compare:
- Facial features
- Clothing
- Background objects
- Architecture
- Lighting
- Image dimensions
- Cropping
- Source documents
- Dates
- Captions
- Metadata
- Surrounding pages
Do not identify a person solely from an automated face match. Do not use image tools to identify a protected survivor, witness, or minor.
Flight Network
The Epstein Data Flight Index maps thousands of flight records and allows research by passenger, route, aircraft, date, and destination.
The index can help researchers:
- Find passenger name variations
- Reconstruct flight sequences
- Compare routes
- Identify aircraft
- Review departure and arrival locations
- Examine repeated travel patterns
- Locate the original source records
A database entry remains a transcription or structured interpretation of a source record. Researchers must inspect the underlying log or manifest before publication.
A flight entry may document that a person was listed on a particular leg. It does not establish why the person traveled, what happened at the destination, or whether the passenger knew about another person’s conduct.
Network and Relationship Tools
The Epstein Data network system organizes people, organizations, locations, and documented relationships into searchable visual structures.
The tools can help identify:
- Repeated correspondence
- Shared corporate roles
- Travel overlap
- Financial references
- Legal relationships
- Employment
- Scheduled meetings
- Document co appearances
- Property connections
- Possible research clusters
A network graph is a research map. It is not a guilt chart.
Two people may appear near each other because their names occur in the same document, because they shared an institution, or because an automated system proposed a relationship. Every connection requires review of the underlying evidence.
EpsteinWiki should describe the exact relationship instead of relying on vague language such as “connected to Epstein.”
Geographic Research
The Global Heatmap organizes geographic references across the corpus. The platform reports references involving 145 countries.
Geographic tools can assist with:
- Location searches
- Property research
- Travel patterns
- International corporate records
- Government contacts
- Institutional networks
- Event reconstruction
- Regional document clusters
A country appearing in a document does not establish government involvement. A location count may include travel references, addresses, news articles, legal citations, business records, or incidental mentions.
Researchers must open the cited documents and determine what each location reference actually represents.
FBI Evidence Map
The FBI Evidence Map reconstructs serial numbers, evidence receipts, witness material, and other records associated with FBI investigations.
The map can help researchers compare:
- Indexed records
- Located documents
- Missing documents
- Illegible references
- Evidence receipt numbers
- FBI serial numbers
- Testifying witness material
- Non testifying witness material
- Related case numbers
A record marked as missing means the project did not locate the expected item within the public release under the stated method. It does not necessarily prove that the government destroyed, concealed, or never possessed the record.
Missing record analysis should clearly identify the index, expected serial, search method, and limits of the public production.
Investigation Reports
The Epstein Data Investigation Reports library contains forensic reports organized across multiple research categories.
Topics include:
- Financial forensics
- Institutional failures
- Congressional briefings
- Evidence and device analysis
- Intelligence records
- Social network analysis
- Government officials
- Dataset analysis
- Internet theories
- Prosecutorial questions
- Document removal
- Missing records
- Properties and travel
- Legal records
- Redaction issues
Each report is intended to connect analytical findings with specific EFTA citations.
Epstein Data states that its analytical text and investigation reports are generated with artificial intelligence and iteratively checked against source documents, but may still contain errors. EpsteinWiki contributors must independently verify any report before using it.
The correct research sequence is:
- Read the report.
- Identify the relevant claim.
- Open every cited EFTA record.
- Verify the quotation and page.
- Review surrounding pages.
- Search for contradictory evidence.
- Check related court or government records.
- Write an independent EpsteinWiki assessment.
Do not copy an automated report into EpsteinWiki and present it as independently verified research.
Survivor Voices
The Hear From the Survivors section collects publicly available video interviews and statements from survivors who spoke publicly and on the record.
The page centers survivor accounts as the reason the records matter. It links to recordings hosted by their original publishers rather than reproducing complete testimony as extracted text.
EpsteinWiki contributors using this section must:
- Respect the survivor’s public framing
- Preserve the context of the interview
- Verify quotations against the recording
- Include timestamps
- Avoid unnecessary graphic detail
- Avoid inferring consent
- Avoid identifying protected people
- Link to the original publisher when available
- Follow current survivor privacy requests
A survivor’s public interview is an important primary account. It should be treated with dignity and evaluated according to its context, subject matter, and relationship to other evidence.
Open Source Data and Independent Review
Epstein Data publishes structured exports through the Epstein Research Data GitHub repository.
Available resources include structured information derived from:
- Full text documents
- EFTA mappings
- Entity extractions
- Knowledge graphs
- Images
- Redaction analysis
- Communications
- Transcripts
- Spreadsheets
- Document classifications
- Metadata
- Investigation report citations
Open source access allows researchers to inspect the data structure, reproduce searches, perform independent analysis, and build additional public interest research tools.
Structured data can contain processing errors. Researchers should treat database rows as pointers to evidence and verify them against the source file.
Tools for Developers and Advanced Researchers
Epstein Data provides technical access for researchers who need to work with large collections.
Available options include:
- Public structured data
- Downloadable databases
- A REST API
- OpenAPI documentation
- Model Context Protocol access
- Local database installation
- Command line research tools
- Artificial intelligence agent instructions
The Run Your Own Investigator guide explains how researchers can download the databases and search them locally.
Local analysis may provide greater control over search methods and research notes. It does not remove the obligation to protect private information or verify original records.
Privacy Practices
The Epstein Data Privacy and Terms page describes the site as a public interest research tool.
According to the published policy:
- Ordinary use does not require an account
- The application does not use Google Analytics
- Search queries are recorded without an application level IP address or browser fingerprint
- Recent artificial intelligence chat history is stored in the user’s browser unless the user chooses to share it
- Artificial intelligence questions are processed through an outside service
- The platform provides a privacy and removal request process
- Users are instructed not to republish information that may identify a trafficking or abuse victim
Infrastructure providers may maintain standard security logs. Researchers should read the current privacy policy before submitting sensitive material or using interactive tools.
The Epstein Data contact page includes options for general feedback, privacy concerns, removal requests, document leads, broken links, and press inquiries.
Privacy and Takedown Requests
Epstein Data provides a specific process for reporting:
- Mistaken redactions
- Exposed personal information
- Survivor identifying information
- Photographs of minors
- Nonconsensual intimate imagery
- Other privacy concerns
A report may identify the EFTA number, page, URL, and nature of the concern.
EpsteinWiki contributors who locate exposed survivor or minor information should not repost it. They should preserve only what is necessary to document the issue and report it through the appropriate privacy channel.
See the EpsteinWiki Privacy Safeguards for Minors before handling sensitive material.
What Epstein Data Can Establish
Epstein Data can help establish that:
- A particular document appears in an indexed public collection
- A term appears in extracted or searchable text
- A document is associated with an EFTA identifier
- A record contains a particular image or statement
- Multiple copies or versions of a document may exist
- A name or entity appears within a database
- A source record may support further investigation
The original document, not the search result, establishes the underlying evidence.
What Epstein Data Does Not Automatically Establish
A search result does not automatically prove:
- A person’s identity
- A personal relationship
- Attendance at a scheduled meeting
- The purpose of travel
- Knowledge of criminal conduct
- Participation in criminal conduct
- Accuracy of a witness statement
- Authenticity of every attachment
- Beneficial ownership of an entity
- The meaning of a redaction
- That a missing record was intentionally withheld
- That an artificial intelligence summary is correct
- That an extracted entity match refers to the intended person
Names in the Epstein files can belong to survivors, witnesses, investigators, attorneys, employees, social contacts, public officials, journalists, medical professionals, unrelated third parties, or people mentioned only in news material.
Document presence must never be converted into guilt by association.
Known Research Limitations
Researchers should account for several limitations.
Optical Character Recognition Errors
Scanned pages may produce incorrect words, names, dates, or numbers. Search the original image and common spelling variations.
Handwriting Errors
Automated handwriting transcription may misread initials, names, telephone numbers, or short notes.
Entity Matching Errors
Two people with similar names may be merged. One person may also appear under several spellings.
Image Description Errors
Automated image descriptions may misidentify people, places, objects, or activities.
Artificial Intelligence Errors
Generated summaries may hallucinate facts, combine separate records, omit qualifications, or misunderstand legal language.
Incomplete Public Releases
A search covers the indexed material, not every record that may ever have existed. Absence from the database is not proof that an event did not occur.
Duplicate Records
The same email, filing, image, or report may appear repeatedly. Duplicate appearances do not create independent corroboration.
Redactions
Redacted information should not be guessed. Optical character recognition near a redaction may produce meaningless text rather than recovered content.
EpsteinWiki Verification Workflow
Step 1: Search Broadly
Search the name, phrase, entity, address, or EFTA identifier across relevant Epstein Data collections.
Step 2: Open the Original Record
Do not rely on a result preview, extracted row, or artificial intelligence summary.
Step 3: Verify the Identifier
Confirm the complete EFTA number and document page.
Step 4: Read the Entire Context
Review the full page, surrounding pages, attachments, email thread, or related document sequence.
Step 5: Classify the Record
Determine whether the material is a court filing, witness statement, investigative note, contact entry, email, calendar entry, photograph, financial record, news clipping, or another record type.
Step 6: Define What It Establishes
Write the narrowest accurate statement supported by the document.
Step 7: Identify Limitations
State what the record does not prove.
Step 8: Seek Corroboration
Compare the record with other EFTA documents, court filings, public records, testimony, and reliable reporting.
Step 9: Search for Contradictions
Look for denials, conflicting dates, alternate identities, later filings, corrections, and reasonable alternative explanations.
Step 10: Protect Survivors and Minors
Check names, images, metadata, indirect identifiers, and redactions before saving or publishing material.
Step 11: Add the Direct Citation
Link the complete EFTA identifier directly to Epstein Data near the relevant claim.
Step 12: Conduct Editorial Review
Verify the article against the original evidence before publication.
Questions
Is Epstein Data an official government database?
No. It is an independent public interest research project. The underlying collection includes government records, but the platform is not operated by the Department of Justice, FBI, or another government agency.
Is Epstein Data free?
Yes. The public research interface is available without a subscription or required account.
Why does EpsteinWiki use Epstein Data?
Epstein Data provides searchable access and stable direct links to millions of pages. This allows EpsteinWiki to connect documented claims with the records supporting them.
Does the partnership mean that every Epstein Data report is approved by EpsteinWiki?
No. EpsteinWiki independently verifies records and applies its own investigative and editorial standards.
Can an Epstein Data artificial intelligence answer be cited as evidence?
No. Use the answer to locate relevant documents. Cite and analyze the underlying records.
Is an EFTA number proof that a statement is true?
No. It identifies a record in the production. Statements inside the record must still be evaluated according to their source, context, and evidentiary status.
Can a search result prove that two people had a relationship?
No. The underlying documents must establish the nature, timing, and significance of the connection.
What should a researcher do when optical character recognition conflicts with the document image?
Use the document image. Record the transcription error when it could affect searches or interpretation.
Does no search result mean that no evidence exists?
No. A name may be misspelled, handwritten, redacted, poorly scanned, stored in another collection, or absent from the public release.
Can a reverse image result identify someone?
It can generate a lead. Identity must be independently verified before publication.
Are Epstein Data investigation reports human written?
The platform states that its analytical reports and text use artificial intelligence and are iteratively checked against source documents. Errors may remain, so EpsteinWiki requires independent verification.
Can researchers download the data?
Yes. Epstein Data provides open source repositories, structured exports, technical documentation, and options for local research.
What should a contributor do after finding an important EFTA document?
Open the original, verify the identifier and pages, read the surrounding context, preserve the direct link, seek corroboration, assess privacy, and submit the finding through the EpsteinWiki editorial process.
What should a contributor do after finding exposed survivor information?
Do not republish it. Record only what is required to locate the problem and submit a privacy report through Epstein Data and the appropriate EpsteinWiki safety channel.
Can Epstein Data replace court record research?
No. Court dockets should be checked directly for filing history, later orders, amendments, appeals, dismissals, and final outcomes.
Who operates Epstein Data?
The platform describes itself as an independently run research project. Its public source data and research repositories are maintained through the rhowardstone GitHub account, and contact with the operator is available through the site’s contact page.
Related EpsteinWiki Guides
- EpsteinWiki Investigative Standards
- Epstein Files: Evidence Framework
- Fact Checking and AI Detection Tools
- Handling Contradictory Evidence
- Disinformation and Manipulated Media Guide
- Metadata Extracts
- Privacy Safeguards for Minors
- Upload Instructions
- Version Control
- Style Guide
- Formatting Rules
Sources
- Epstein Data
- Epstein Data Search
- Epstein Data Database Explorer
- Epstein Data Research Assistant
- Epstein Data Reverse Image Search
- Epstein Data Flight Index
- Epstein Data FBI Evidence Map
- Epstein Data Investigation Reports
- Hear From the Survivors
- Epstein Data Community Resources
- Epstein Data Privacy and Terms
- Epstein Data Contact and Privacy Requests
- Epstein Research Data GitHub Repository
- Run Your Own Epstein Data Investigator
- United States Department of Justice Epstein Library
Editorial Note
Epstein Data is an EpsteinWiki research and evidence partner. EpsteinWiki independently reviews all records, analytical claims, automated outputs, and investigation reports before incorporating them into published articles. Collection totals and available tools may change as Epstein Data processes additional records and improves its databases.
Last reviewed: September 6, 2026