Home/Blog/The Dark Data Problem: What is hidden in unstructured files can kill you
Discovery & Classification
August 9, 2026

The Dark Data Problem: What is hidden in unstructured files can kill you

By Jan Brown

The Dark Data Problem: What is hidden in unstructured files can kill you

Most enterprise data governance programs were built for one shape of data: tables, rows, and columns. Structured databases get access controls, masking, retention policies, and audit trails. That work is real, and it matters.

But it also covers only part of the problem.

The largest share of an enterprise’s dark data — content nobody has classified, indexed, or accounted for — doesn’t live in a database. It lives in file shares, scanned archives, and cloud object storage: PDFs, scanned contracts, HR exports, spreadsheets dropped on a shared drive years ago and never touched again.

A Social Security number buried in a scanned intake form is still Personally Identifiable Information under GDPR and CCPA, whether anyone in the organization can find it or not. The regulation doesn’t distinguish between data you lost track of and data you never had. That distinction only matters to the people writing the incident report after the fact.

This is the gap ROAD was built to close with a new capability: Data Discovery for Unstructured Files.

Structured governance was only half the picture

infoCorvus customers running ROAD already have strong control over structured pipelines — application retirement, data migration, data warehouse ingestion, all governed, all auditable. What most organizations have never had is equivalent visibility into the unstructured backlog sitting alongside those systems.

Data Discovery for Unstructured Files extends the same governance discipline ROAD already applies to structured data — govern first, reduce risk, then enable AI — to the file-based content that has historically sat outside of it.

Three ways to run it, depending on the operational problem

The capability connects to on-premise file servers and NAS drives, AWS S3, Azure Blob Storage, Google Drive and Google Cloud Storage, and SFTP sources, including multi-server environments. It runs in three modes:

File System Analysis — in-place analysis. Files are indexed, classified, and scanned for sensitive data in their current location. Nothing moves.

File System Archive — analysis happens during migration. Files are classified, summarized, and organized while they’re copied from source to a target destination.

File System Collection — a unified searchable index built across multiple storage systems at once, so end users can search one place instead of five.

Which mode you use depends on the operational outcome you need — understanding what you already have, moving it somewhere governed, or making it searchable across every system it currently lives in.

Classification and extraction, without a data science team

Once connected, ROAD uses Apache Tika for text extraction — including OCR for scanned PDFs and image-only files — and applies automated classification to sort documents by type: invoice, purchase order, contract, HR document, legal document, marketing material, or any custom category described in plain language.

Entity extraction works the same way. Users describe what they’re looking for — a part number format, a store code prefix, any domain-specific identifier — and ROAD passes that description to the configured large language model at runtime. No model training, and no need to interact with the LLM platform directly.

Every processed document can also receive a plain-language summary, with sentence count configurable per job, so teams can scan results instead of opening every file by hand.

Finding sensitive data is not the same as proving you found it

For compliance and security teams, the meaningful capability isn’t just detection — it’s evidence. ROAD’s Discovery module scans unstructured content using a library of predefined scanners, covering Social Security numbers, driver’s licenses with state-level pattern matching, passport numbers, bank account and credit card numbers, phone numbers, email addresses, physical addresses, and medical terminology with multi-language term lists, including English and French.

Organizations can also build custom scanners using their own regex patterns or word lists, scoped to a specific territory, matched on a whole-field or partial basis. When a scanner flags a match, ROAD highlights the exact location of that match inside the document — not a probability score, a specific location a compliance team can point to.

That distinction is what turns dark data from a theoretical risk into something an organization can actually audit, report on, and remediate.

Search that returns answers, not folders to dig through

Once content is indexed, ROAD supports three ways to query it: simple keyword search across content and metadata, advanced search using structured criteria like file class, extension, size, path, or owner, and natural language search, where a free-text query is interpreted by the configured LLM.

A query like “find documents that reference Oracle EBS” or “show me files that contain Social Security numbers” runs against the full indexed corpus, including AI-generated summaries, in whatever language the configured LLM supports — English, French, Chinese, and others — with automatic typo correction. Search results can be downloaded as a single ZIP file, rather than tracked down one file at a time.

None of that search capability works without the governance layer underneath it. Classification, discovery, and indexing have to run first. Search is the payoff, not the starting point.

Governance before AI, not AI instead of governance

There’s a pattern showing up across enterprise AI initiatives right now: teams want AI to surface insights from document archives, contracts, and historical records, and that’s a reasonable ambition. But if the underlying files haven’t been classified, the sensitive data hasn’t been mapped, and access hasn’t been defined, pointing an AI agent at that content isn’t productivity — it’s exposure.

Once files are classified and sensitive data is identified through Data Discovery for Unstructured Files, that content can be exposed to AI agents and large language models through ROAD’s Governed AI Bridge, under the same role-based, auditable access controls ROAD already applies to structured sources. AI agents are treated as users. They only see the subset of data they’re permitted to access. Every action is logged. ROAD supports multiple LLM providers for this work, including self-hosted deployment options for organizations that require air-gapped or data-sovereign environments.

Govern the data. Reduce the risk it carries. Then, and only then, put AI to work on top of it.

What this means for the systems already in your environment

Data Discovery for Unstructured Files is built directly into ROAD alongside Application Retirement, Data Migration, and Data Warehouse Ingestion. For infrastructure and operations teams, that also means archive mode can reduce storage footprint through directory organization by file type or date, ZIP compression, and configurable retention rules — cutting the cost and risk of holding file content indefinitely with no plan for it.

For compliance and security teams, the outcome is a file environment that moves from unmanaged to auditable: dormant PII exposure becomes something you can see and search, instead of something sitting in storage nobody reviews.

Structured data governance was never the whole job. It was the part that was easiest to measure. Data Discovery for Unstructured Files is designed to close the rest of it — bringing the same governance discipline ROAD already applies to databases to the documents most enterprises have never been able to search or control.


Data Discovery for Unstructured Files is available now to ROAD customers.

Ready to see what’s actually in your file environment? Request a demo to walk through Data Discovery for Unstructured Files against a sample of your own data.