Drop in messy spreadsheets, PDFs, Word docs or emails. CyberInci pulls PII from tables and from free text — names, IDs, dates and contact data — into a clean roster, with a confidence score on every record.
| DEMO CO — SAMPLE DATA | |||
| Payroll export (fictional) | |||
| EMP_NO | EMPL_NM | SOC_SEC | BRTH_DT |
| 1001 | DOE, JANE B | 000-12-3456 | 01/14/90 |
| 1002 | SMITH, ALEX J | 000-23-4567 | 07/22/85 |
| 1003 | GARCIA, MARIA | 000-34-5678 | 03/09/92 |
Employee Identification Number
1001
Full Name (Last, First M.I.)
DOE, JANE B.
Social Security Number
000-12-3456
Full Date of Birth (MM/DD/YYYY)
01/14/1990
| ACME Corp | |||
| Q4 Report | |||
| EMP_ID | NAME | SSN | |
| 1001 | John S. | 123-XX | john@ |
| employee_id | full_name | ssn | |
| 1001 | John Smith | 123-45-6789 | john@acme.com |
| 1002 | Jane Doe | 234-56-7890 | jane@acme.com |
| 1003 | Bob Johnson | 345-67-8901 | bob@acme.com |
Real-world files, however messy. These are the shapes CyberInci handles out of the box — every preview below is fictional sample data.
| GLOBEX INC — INTERNAL USE | ||
| EMP_NO | EMPL_NM | SOC_SEC |
| 1001 | DOE, JANE B | 000-12-3456 |
| 1002 | SMITH, ALEX J | 000-23-4567 |
Header found on row 3 — branding skipped
Messy Excel exports
Branding rows, blank rows and logos above the real table — the header is found automatically, wherever it hides.
Download this sample3 sheets detected — all extracted
Multi-sheet workbooks
Every sheet is processed, scored and combined — no copy-pasting tabs together first.
Employee Number,Name of Employee,Social Sec #,Date Birth
1001,DOE JANE B,000-12-3456,01/14/1990
CSVs with human headers
Wordy or abbreviated headers are matched to the canonical schema by the fuzzy and AI tiers — no renaming needed.
Download this sample| NAME | ID | DOB |
| DOE, J | 1001 | 01/14/90 |
| SMITH, A | 1002 | 07/22/85 |
1 table found on page 2
PDFs with tables
Tables inside PDF reports are detected and extracted directly — same pipeline, same confidence scores.
Needs the free Tesseract OCR add-on
Scanned PDFs (OCR)
Image-only scans are read with the optional OCR add-on, then extracted like any other document.
| Employee ID? | Name? | SSN? |
| 1001 | DOE, JANE B | 000-12-3456 |
| 1002 | SMITH, ALEX J | 000-23-4567 |
No header row — columns inferred from values
Files with no headers at all
When there is no header row, columns are inferred from the data itself — IDs, names, SSNs and dates are recognized by shape.
Please notify Marcus Bellweather, SSN 000-31-4477, DOB 03/15/1985, at marcus.b@example.com.
Detected in text — review before notifying
Word docs, emails & letters
No table? PII is pulled from prose — email bodies, letters, case notes — and assembled into one draft person per name, flagged for review.
{
"name": "Doe, Jane",
"ssn": "000-12-3456",
"contact": {
"email": "jane@example.com"
}
}Nested keys flattened to columns
JSON, XML & data files
Records nested in JSON, XML, YAML or HTML — and legacy dBase (.dbf) — are flattened and mapped to the same schema.
Spreadsheets, PDFs, Word docs, emails, breach letters — whatever the source, we turn it into a clean roster.
Excel, CSV, PDF, Word, Outlook email, JSON/XML and more. Extracts PII from tables and from prose — breach letters, email bodies, case notes.
Finds the real header in branded, multi-row or headerless files, then maps columns by both header text and value shape — with persistent learning.
A confidence score on every record, duplicate people merged across files, and anything uncertain routed to review — never silently accepted.
Three steps from messy spreadsheet to normalized, validated PI data.
Drop spreadsheets, PDFs, Word docs, emails — or a whole folder. Any format, any structure, any mess.
Two engines extract PII from tables and from free text, with confidence scoring on every record.
Get a normalized roster instantly, and review anything flagged — including people found in text.