Skip to content

HIPAA limited data set: the 16 identifiers to remove, and how it differs from Safe Harbor

A limited data set and a Safe Harbor de-identified data set are produced by two different removal lists, and the practical difference comes down to three fields: dates, geography, and re-identification codes. This is the side-by-side reference, drawn from 45 CFR 164.514(b)(2) and 164.514(e)(2). It is an operational aid, not legal advice, and expert determination under 164.514(b)(1) is a separate route with no fixed list at all.

What the operating record should prove

The two removal lists side by side

Read the two right-hand columns together. Where they differ is where the choice between the two routes actually bites.

The three fields that decide which route you need

Eighteen rows, and only three of them differ. Those three are usually the entire reason a project chooses a limited data set over Safe Harbor.

What the two routes cost you

A limited data set remains protected health information. It may be disclosed only for research, public health, or health care operations, and only under a data use agreement satisfying 45 CFR 164.514(e)(4). It stays inside the Privacy Rule, which means accounting, safeguards, and the covered entity's cure obligation on breach all continue to apply.

The two removal lists side by side

Read the two right-hand columns together. Where they differ is where the choice between the two routes actually bites.

Limited data set exclusions compared with the Safe Harbor identifiers
IdentifierLimited data set, 164.514(e)(2)Safe Harbor, 164.514(b)(2)(i)
NamesRemoveRemove
Postal addressRemove street address. Town or city, state, and full ZIP code may be retainedRemove all geography below state, except the first three ZIP digits where that area holds more than 20,000 people
Dates, and ages over 89May be retained, including birth date, admission, discharge, service, and death datesRemove all date elements except year, and all ages over 89 and dates indicating such age
Telephone numbersRemoveRemove
Fax numbersRemoveRemove
Email addressesRemoveRemove
Social security numbersRemoveRemove
Medical record numbersRemoveRemove
Health plan beneficiary numbersRemoveRemove
Account numbersRemoveRemove
Certificate and license numbersRemoveRemove
Vehicle identifiers and serial numbers, including license platesRemoveRemove
Device identifiers and serial numbersRemoveRemove
Web URLsRemoveRemove
IP addressesRemoveRemove
Biometric identifiers, including finger and voice printsRemoveRemove
Full face photographs and comparable imagesRemoveRemove
Any other unique identifying number, characteristic, or codeMay be retained, which is what permits a study identifierRemove, subject to the re-identification code conditions at 164.514(c)

The three fields that decide which route you need

Eighteen rows, and only three of them differ. Those three are usually the entire reason a project chooses a limited data set over Safe Harbor.

  • Dates. A limited data set may keep full dates, so length of stay, time to event, seasonality, and longitudinal follow-up all survive. Safe Harbor keeps only the year.
  • Geography. A limited data set may keep town or city, state, and the full five digit ZIP code. Safe Harbor keeps state and, conditionally, three ZIP digits.
  • Study identifiers. A limited data set may carry a code that lets the recipient link records across files. Safe Harbor removes any other unique identifying number, characteristic, or code.

What the two routes cost you

A limited data set remains protected health information. It may be disclosed only for research, public health, or health care operations, and only under a data use agreement satisfying 45 CFR 164.514(e)(4). It stays inside the Privacy Rule, which means accounting, safeguards, and the covered entity's cure obligation on breach all continue to apply.

Data de-identified under Safe Harbor is not protected health information, so no data use agreement is required and the Privacy Rule does not restrict its use or disclosure. The cost is analytic: with year-only dates and state-level geography, many clinical and epidemiological questions become unanswerable. That trade is the decision, and it should be made before the extract is built rather than after a reviewer blocks it.

Checking an extract against the list

The list is a field allowlist in disguise, which makes it checkable. Compare the schema of the proposed extract against the removal column, then check the rows: a free-text note column can carry a name or an address even when the structured columns are clean, and a derived column such as a full date of death can reintroduce an element the route excludes.

Audarel treats the confirmed agreement terms as a typed allowlist and inspects the proposed schema and payload against it, identifying the specific columns or rows that fail before the release is approved. See the data use agreement sample for the agreement that a limited data set disclosure requires.

Why this page exists

Keep decisions human and evidence explicit.

Practical guidance that connects policy documents to observable release controls.

Primary references

Confirm requirements against current source material.

Requirements and vendor capabilities change. Confirm the current source and your approved QC plan before changing a production process.

From evidence to conclusion

Put the guidance inside a reproducible release record.

Audarel is in development. If you review data releases against executed agreements today, we want to understand how.

Contact us Read the guides