Skip to main content
Case Study: Research & ContentClient Anonymized

Nobody Could Say Which Document Was The Real One

Acquisition Team, Buy SideA deal team of about ten2 months to production
3939
Sources Indexed And Status Marked
Nobody Could Say Which Document Was The Real One

Confidentiality Note: Fully de-identified. No party, sector, figure, date or document title appears. Only the method is described. This work was done in an operating role inside the client's group of companies rather than under a separate vendor agreement.

The Full Story

A deal team had more documents than it could hold in its head, and no way to tell which copy of anything was the real one. A long management presentation as a scanned file. A seller hosted document room where some folders simply were not visible to the buy side. Spreadsheets with tabs hidden from the viewer and formulas resolving to errors. Attachments scattered across a mailbox. Everyone was reading the same deal and no two people were certain they were reading the same version of it.

The cost of that was not hypothetical and it surfaced in the first week. One of the central commercial terms was being carried verbally, from one conversation, and the whole team had internalised it. When the actual document was read, the term was structured differently. Everyone had been reasoning from a shared belief that was wrong, and nobody could have caught it, because there was nothing to catch it against.

So the first deliverable was not analysis. It was an index. Every artifact in the deal got a number, a type, a pointer to where its machine readable extraction lives, and a verification status: verified, in hand but not yet read, or superseded. That last category is the one that does the work. A source that has been replaced stays in the index, marked, with the thing that replaced it, so nobody quietly re-derives a conclusion from a document that has been overtaken.

The index is also honest about what it has not read. A meaningful share of entries sit at in hand but not yet read, and they are visible as such rather than folded into a total that implies full coverage. An index that claims completeness is worse than no index, because it stops people looking.

The extraction underneath is deliberately mechanical. The presentation was transcribed page by page and then spot checked against the original, including both of the dense financial tables, and the spot check matched exactly. Spreadsheets were dumped programmatically to values only files rather than read through a viewer, which is how the hidden tabs came back: twenty of them, in one workbook, carrying content the on screen version simply did not offer. Image only pages were reviewed visually rather than skipped.

Then the boring, load bearing part. The document room was mirrored locally, and each fresh export was compared against the mirror file by file. The most recent comparison came back with sixty five of sixty five files identical, which is a useful answer and an unglamorous one: nothing new had been posted. Knowing that with certainty is worth more than assuming it.

The last piece is a small script that takes a finished brief and greps every quoted line back to the extracted source text, run with a control so a pass means something. Sixty four claims went through it on the flagship brief.

The standing rule on all of it, set at the start, is that the outputs are recommendations or a statement that more information is needed. The people who own the decision own the decision. Gaps get named, not written around.

The Challenge

A deal team was working from a long management presentation, a seller hosted document room whose folders were not all visible to them, spreadsheets carrying hidden tabs and broken formula references, and attachments scattered through a mailbox. Nobody could say which copy of a given document was the current one, or which claims in the analysis rested on which page. The concrete cost of that showed up early: a central term of the deal was being carried verbally, everyone was working from it, and it was wrong.

Our Solution

A numbered index of every artifact in the deal, each carrying its type, where its extracted machine readable version lives, and a verification status: verified, in hand but not yet read, or superseded. A slide by slide transcription of the management presentation, spot checked against the original. Programmatic dumps of every spreadsheet to values only files, including the tabs the viewer did not show. A local mirror of the document room and an integrity comparison against each fresh export. And a small verifier that greps every quoted line in a brief back to the extracted source text.

The Result

An Index That Is Honest About Its Own Gaps

3939
Sources Indexed And Status Marked
2020
Hidden Spreadsheet Tabs Recovered
65 of 6565 of 65
Files Identical On Integrity Check
6464
Claims Grep Verified Against Source