Any document. Clean, compliance-ready data.

Drop in messy spreadsheets, PDFs, Word docs or emails. CyberInci pulls PII from tables and from free text — names, IDs, dates and contact data — into a clean roster, with a confidence score on every record.

sample_payroll_export.xlsxdemo · sample data
DEMO CO — SAMPLE DATA   
Payroll export (fictional)   
    
EMP_NOEMPL_NMSOC_SECBRTH_DT
1001DOE, JANE B000-12-345601/14/90
1002SMITH, ALEX J000-23-456707/22/85
1003GARCIA, MARIA000-34-567803/09/92
Branding rows detected & skipped · header row 4 · score 0.97
5-tier AI
Canonical outputACCEPT · 1.00

Employee Identification Number

1001

cache

Full Name (Last, First M.I.)

DOE, JANE B.

cache

Social Security Number

000-12-3456

fuzzy

Full Date of Birth (MM/DD/YYYY)

01/14/1990

embedding
messy_data.xlsx
 ACME Corp  
 Q4 Report  
EMP_IDNAMESSNEMAIL
1001John S.123-XXjohn@
Cache
Fuzzy
Embedding
LLM
Infer
normalized_pi.xlsx
employee_idfull_namessnemail
1001John Smith123-45-6789john@acme.com
1002Jane Doe234-56-7890jane@acme.com
1003Bob Johnson345-67-8901bob@acme.com
Confidence: 0.9280/81 rows keptACCEPT

What can it extract?

Real-world files, however messy. These are the shapes CyberInci handles out of the box — every preview below is fictional sample data.

quarterly_roster.xlsxxlsx
GLOBEX INC — INTERNAL USE
EMP_NOEMPL_NMSOC_SEC
1001DOE, JANE B000-12-3456
1002SMITH, ALEX J000-23-4567

Header found on row 3 — branding skipped

Messy Excel exports

Branding rows, blank rows and logos above the real table — the header is found automatically, wherever it hides.

Download this sample
regional_offices.xlsxls
EastWestCentral

3 sheets detected — all extracted

Multi-sheet workbooks

Every sheet is processed, scored and combined — no copy-pasting tabs together first.

contact_list.csvcsv

Employee Number,Name of Employee,Social Sec #,Date Birth

1001,DOE JANE B,000-12-3456,01/14/1990

"Social Sec #" → SSN"Date Birth" → DOB

CSVs with human headers

Wordy or abbreviated headers are matched to the canonical schema by the fuzzy and AI tiers — no renaming needed.

Download this sample
benefits_report.pdfpdf
NAMEIDDOB
DOE, J100101/14/90
SMITH, A100207/22/85

1 table found on page 2

PDFs with tables

Tables inside PDF reports are detected and extracted directly — same pipeline, same confidence scores.

old_records_scan.pdfpdf
OCR

Needs the free Tesseract OCR add-on

Scanned PDFs (OCR)

Image-only scans are read with the optional OCR add-on, then extracted like any other document.

export_no_headers.csvcsv
Employee ID?Name?SSN?
1001DOE, JANE B000-12-3456
1002SMITH, ALEX J000-23-4567

No header row — columns inferred from values

Files with no headers at all

When there is no header row, columns are inferred from the data itself — IDs, names, SSNs and dates are recognized by shape.

breach_notice.docxdocx

Please notify Marcus Bellweather, SSN 000-31-4477, DOB 03/15/1985, at marcus.b@example.com.

Detected in text — review before notifying

Word docs, emails & letters

No table? PII is pulled from prose — email bodies, letters, case notes — and assembled into one draft person per name, flagged for review.

export.jsonjson
{
  "name": "Doe, Jane",
  "ssn": "000-12-3456",
  "contact": {
    "email": "jane@example.com"
  }
}

Nested keys flattened to columns

JSON, XML & data files

Records nested in JSON, XML, YAML or HTML — and legacy dBase (.dbf) — are flattened and mapped to the same schema.

Built for messy data

Spreadsheets, PDFs, Word docs, emails, breach letters — whatever the source, we turn it into a clean roster.

Any file — tables and free text

Excel, CSV, PDF, Word, Outlook email, JSON/XML and more. Extracts PII from tables and from prose — breach letters, email bodies, case notes.

Smart header & column mapping

Finds the real header in branded, multi-row or headerless files, then maps columns by both header text and value shape — with persistent learning.

Confidence, dedup & review

A confidence score on every record, duplicate people merged across files, and anything uncertain routed to review — never silently accepted.

How it works

Three steps from messy spreadsheet to normalized, validated PI data.

01

Upload

Drop spreadsheets, PDFs, Word docs, emails — or a whole folder. Any format, any structure, any mess.

02

Process

Two engines extract PII from tables and from free text, with confidence scoring on every record.

03

Download / Review

Get a normalized roster instantly, and review anything flagged — including people found in text.

Ready to normalize your data?

Upload your first file and watch it turn tables and text into a clean roster.