When something has gone wrong inside a company — a whistleblowing report, a regulator's letter, an auditor's question that will not go away — the first task is always the same: establish what actually happened. And what actually happened lives, overwhelmingly, in documents. Emails, chat messages, contracts, approvals, board papers, expense records. The quality of an investigation is set by what was actually read.

For a long time, the binding constraint was reading time. A team of lawyers can review only so many documents per day at a defensible level of attention, and every reviewed document costs money. So investigations were designed around scarcity: pick custodians, pick date ranges, pick keywords, review the slice that survives the filters, and hope the slice contains the story.

Sampling was a budget decision dressed as method

Keyword lists and custodian selection feel rigorous, but they encode guesses. If the people involved used a project nickname, switched from email to chat, or simply never wrote the obvious words, the filter misses the events themselves. The narrower the slice, the more the investigation finds what it expected to find — and nothing else.

Sampling also creates a defensibility problem at the end. When a board, an auditor or an authority asks "did you look at everything relevant?", the honest answer under the traditional model was no — we looked at what we could afford. That answer is increasingly hard to give.

Full sets instead of samples

Structured, machine-scale review changes the starting point. Instead of filtering first and reading second, systems read every document in the collected set and record what each one is: who wrote it, when, to whom, what it concerns, whether it relates to the questions under investigation. Nothing is excluded because a keyword list failed to anticipate it.

The practical effect is that scope becomes a legal decision rather than a budget decision. The question shifts from "which ten thousand documents can we afford to read?" to "which custodians and systems does the allegation actually touch?" — which is the question an investigation should have been asking all along.

Deduplication and the shape of the record

Corporate document sets are mostly redundant. The same attachment is forwarded a dozen times; every reply in an email thread contains the messages before it; the same draft exists in five versions across three drives. Deduplication and thread reconstruction collapse this volume so that each substantive document is assessed once — and, just as importantly, they make the distribution visible: who received which version, when, and who was dropped from the thread. In many investigations, that distribution is the actual question.

Findings you can trace to a source

Reading everything only matters if the output stays anchored to it. A report that says "management was aware of the issue by early spring" is an assertion. A report that says the same thing and points to the specific message, its date, its author and its recipients is a finding. Machine-scale review makes this discipline cheap: because every document is indexed, every statement in the findings can carry its sources, and anyone checking the work — the board, the auditors, an authority, opposing counsel — can follow the link rather than take the author's word.

Source-linking also keeps the investigation honest with itself. Claims that cannot be tied to a document get flagged as inference or interview evidence, not silently promoted to fact.

What still requires a lawyer in the room

None of this removes the lawyer from the investigation; it changes what the lawyer spends time on. Machines do the structured reading. They do not decide what the investigation is for, and they do not answer for it.

  • Strategy: what is in scope, in which order, and what the company does with what it learns — including whether and when to inform an authority.
  • Privilege: how the investigation is commissioned, who receives the findings, and what circulates in which form. These choices are made before and during the review, not after.
  • Interviews: documents say what was written; only people can explain what was meant, and assessing their credibility is judgment, not extraction.
  • Conclusions: whether conduct breached a duty, what the legal exposure is, and what a proportionate response looks like — questions on which the honest answer often begins with "it depends on the facts", and on which a named lawyer must take a position.

The combination is the point: a review that covers the full record, and a lawyer who challenges what comes out of it and answers for the result. If you are facing an investigation and wondering what a complete review would look like on your facts, we are happy to discuss it.