AI labs are paying for real business records.See what yours could be worth →

Privacy

How business records are de-identified before licensing

Key takeaways

  • Good de-identification replaces people, clients and vendors with consistent stand-ins across every source, so the work still reads clearly but no one is named.
  • Identifiers and secrets are removed outright: tax IDs, account numbers, phone numbers, addresses, passwords and one-time codes.
  • High-risk categories are excluded entirely rather than redacted, including legal correspondence, HR and payroll matters, personal tax records and bank statements.
  • Free-text sweeps and several independent review passes catch what pattern rules miss, but residual risk never reaches zero, so exclusions and contract limits matter too.

Before business records are licensed, they should go through three things in order. High-risk material is excluded entirely. People, clients and vendors are replaced with consistent stand-ins across every source. Identifiers and secrets are removed outright. Then free-text sweeps and several independent review passes catch what the rules miss. Done well, the work still reads clearly, but no one in it is named.

This guide explains each step, why consistency matters so much, and the risk that remains even when the work is done carefully.

What does de-identification mean for business records?

De-identification is the process of removing or replacing the details that connect a record to a specific person or organization. It’s different from confidentiality. As lawyers at Frankfurt Kurnit Klein & Selz point out, de-identification addresses whether a record can be linked to someone, not whether its content is sensitive. A disciplinary thread or a pay complaint is still sensitive with the name removed.

That’s why good practice starts with exclusions, not redaction.

Business records are also harder to de-identify than tidy database fields. Names appear in signatures, salutations, quoted replies, file names, spreadsheet tabs and meeting transcripts. As Forbes summarized the view of data industry veterans, there’s no “on-off switch” for personal information tied to a career’s worth of work.

Step 1: Decide what’s excluded entirely

Some material is too sensitive to rely on redaction. It should stay out of scope altogether:

  • Legal correspondence. Messages with outside counsel, which may also be privileged.
  • HR, personnel and payroll matters. Reviews, discipline, compensation, benefits, leave and separation.
  • Personal tax records. Individual tax returns, W-2s and similar documents.
  • Bank statements and full account records.
  • Health information and anything from regulated systems.
  • Credential stores: password vaults, key files and folders of logins.
  • Consumer-level detail that isn’t needed to understand the work, such as customer-by-customer transaction lists.
  • Anything the owner flags: a sensitive client, a confidential project, a whole department.

Exclusions are applied by source and by rule. Legal correspondence can be filtered by law-firm email domains, for example, and HR material by folder, channel and keyword. Withheld records should be counted, so the owner knows exactly what was left out.

Step 2: Replace people and organizations with consistent stand-ins

This is the step that preserves value. Every person and organization gets one stand-in, used everywhere:

  • People become invented names with stable IDs, grouped by role, such as staff, client-side contacts and outside parties. The same person is the same stand-in in email, chat, tickets, meeting transcripts and file history.
  • Email addresses become pseudonymous addresses on reserved example domains.
  • Client and company names are replaced everywhere they appear, including inside file names, ledger descriptions, links and signatures.
  • Vendors and suppliers are replaced with consistent tokens too, except for widely known platforms, banks and government bodies. Otherwise a distinctive list of suppliers can identify a business on its own.
  • Document and system IDs that reveal identity, like shared-drive links, are replaced with stable tokens.

Here’s a fictional example of what that looks like:

Before: “Hi Marcus, Priya at Harbor Lane Foods says the March invoice doesn’t match PO 4471. Can you check with Delmar Packaging?”

After: “Hi Owen, Rosa at Client-7f2a says the March invoice doesn’t match PO 4471. Can you check with Vendor-91c0?”

Owen and Rosa are invented stand-ins. Every mention of the same real person, in any system, uses the same stand-in. The request, the problem and the next step all survive. So does the purchase order number, because it carries no personal information. That’s the point.

Why consistency matters

Buyers pay for connected records: the email that started the work, the chat that debated it, the ticket that tracked it and the document that finished it. Consistent stand-ins keep those threads intact. The Frankfurt Kurnit post notes that buyers want records to stay linked to each other, and that linkage is what makes the data valuable. It’s also what creates risk, which is why the other steps matter.

Step 3: Remove identifiers and secrets outright

Some details are never needed to understand the work. They’re removed, not replaced:

  • Social Security numbers, tax IDs and policy numbers.
  • Bank and card account numbers. Where an account matters to the story, a neutral label like “account A” stands in.
  • Phone numbers and street addresses.
  • Passwords, API keys, verification and one-time codes.
  • Attachment IDs, tracking tokens in links and other machine identifiers that could be traced.

Credentials deserve special attention. Years of email and chat almost always contain a few pasted passwords or keys. A dedicated sweep for secrets is part of every pass.

Step 4: Handle spreadsheets column by column

Spreadsheets and system exports need a different approach from free text:

  • Columns that hold personal details, such as consumer names, emails, phone numbers, addresses and bank details, are replaced with a redaction marker.
  • Sheets that are mostly personal detail, like consumer-level transaction listings, are withheld whole.
  • Columns that carry the work are kept: dates, amounts, statuses, account names, invoice numbers and system keys.

A manifest records which files were included, which were withheld and which columns were redacted.

Step 5: Sweep the free text

Pattern rules catch the obvious cases. Free text needs a second, broader pass that looks for:

  • Names in quoted email headers, signatures and greetings.
  • Spelled-out or informal names, checked against name dictionaries.
  • Credentials and one-time codes.
  • Account, tax and policy numbers in unusual formats.
  • Street addresses, city and ZIP code patterns.
  • Tokens hidden in links and calendar invitations.

Step 6: Review, then review again

No rule set is perfect on the first try. Run several independent review passes, each looking for what the last one missed, and fold every finding back into the rules before the final run. The owner should approve the scope and be able to see a de-identified sample before anything is delivered.

Document the result: the protocol used, what was excluded and why, and the review findings. If the company is ever sold, a prior data license will come up in due diligence, and the Frankfurt Kurnit lawyers note that buyers will expect to see that record.

What residual risk remains?

Careful de-identification lowers risk a great deal. It doesn’t eliminate it. The honest list of what can slip through includes:

  • An unusual first name mentioned once in passing.
  • A surname mentioned without a first name.
  • Coarse geography, like a state or city named in conversation.
  • Product or project names that are distinctive enough to point back to a company.
  • Inferences about small groups, which linked records can make possible even without names.

There’s also the model itself. Research has shown that large language models can memorize and reproduce parts of their training data, as Forbes noted. That’s why de-identification works alongside two other protections: excluding high-risk material entirely, and contract terms that limit how records may be used and forbid re-identification. Our guide to license agreement terms covers those clauses.

Who should control de-identification?

The seller should, or at least approve it. The Frankfurt Kurnit lawyers advise companies to select or jointly select the provider, approve the protocol and require certification that it was followed. Some buyers will pay more to receive raw records and do the work themselves. In the Spirit Airlines auction, one bidder made exactly that offer, Business Insider reported. If you consider it, put the protocol, the access limits and the deletion deadline in the contract. For more on that conversation, see questions to ask an AI data buyer.

How Cascade approaches it

At Cascade, rights and privacy are reviewed before any records are shared. You choose what’s in scope, sensitive categories stay out, people are replaced rather than just blacked out, and nothing is final until you sign. We’re straightforward about residual risk, because you should make the decision with the full picture. Read more on our privacy and security page, or start with the full guide to licensing your business data.

This guide is general information, not legal advice. Privacy obligations depend on your records, your contracts and the laws that apply to you.

Frequently asked questions

What's the difference between de-identified and anonymized data?

People often use the words interchangeably. De-identification means removing or replacing the details that link a record to a person. Legal definitions vary. Under California's CCPA, for example, data counts as de-identified only with technical safeguards, a public commitment not to re-identify it and contractual obligations on recipients. Treat 'anonymized' as a claim to verify, not a guarantee.

Why replace names instead of just blacking them out?

Because the value is in the connections. If the same person becomes the same stand-in in every email, chat message and ticket, a reader can still follow who did what. Blacking names out breaks those threads, lowers the value of the records and can still leave context that points to someone.

Can de-identified business records be re-identified?

The risk can be reduced a lot, but not to zero. Unusual details, small teams and linked records can still point to individuals or groups. That's why high-risk material is excluded entirely, buyers commit in writing not to re-identify anyone, and the owner approves the scope.

Who should do the de-identification, us or the buyer?

Ideally the seller controls it, or approves a protocol performed by a provider both sides trust. Some buyers will pay more to receive raw records and de-identify them themselves, but that moves control of privacy to the buyer. If you agree to it, set the protocol, access limits and deletion terms in the contract.

This guide is part of our series on how to license your business data to AI labs.

Related guides

Start here

Let’s talk about what you’ve built.

A conversation is all it takes to start. No files or login details needed.

Request a free consultation Get your estimate first
  • You pay nothing unless a license is signed
  • You approve every term