It’s 2025 and you’re in between jobs. Money is tight so you visit your local food bank’s website to check hours and fill out an intake form. Unbeknownst to you, a tracking pixel embedded in the page is scraping your address and what a later lawsuit would describe as your “intent to receive nutrition assistance.” It’s also tracking your financial hardship, disability status, and how urgent your need is.
All of this information is, allegedly, packaged up and sent to a third party for ad targeting.This incredibly sensitive information wasn’t collected because the food bank’s evaluation plan called for it, but because a tracking tool was installed and nobody asked what information it was gathering in the background.
Now picture yourself a little older, someone who has spent years trusting an aging-services nonprofit with your most sensitive information. That same year, that organization discovers that an unauthorized party has been inside its network, walking away with names and Social Security numbers the organization had been holding, in some cases, for years. Yours included.
Or maybe you never even interacted with the organization directly. Your information just existed somewhere — scraped, purchased, compiled — in the files of a background-check company called National Public Data. At a scale hard to picture, a breach there exposed an estimated 2.9 billion records: full names, addresses, Social Security numbers, and enough detail to identify people’s relatives. All of it was stored well beyond what any single legitimate purpose required.

None of these organizations set out to cause harm, but each one was sitting on more data, more sensitive data, or more identifiable data than its actual purpose demanded. When something went wrong — a hack, a lawsuit, a subpoena — that surplus data was weaponized against the very people it was collected to help.
If you work in applied research or evaluation, you’ve probably sat in a planning meeting where someone asked, “should we just add a few more demographic questions, in case we need them later?” It seems harmless, right? It might even feel more rigorous.
But that instinct is the same instinct behind every story above, just earlier in its life cycle. It has a name: data minimization, or rather, the failure to practice it.
Data minimization is the principle of collecting only the data you actually need to answer your evaluation question or research aim and nothing more. It sounds obvious. In practice, it’s one of the hardest disciplines to maintain, especially when funders, boards, and well-meaning colleagues keep asking “wouldn’t it be nice to know?”
Why Overcollection Happens
Almost nobody sets out to overcollect with bad intentions. In our experience, it usually comes from one of three places:
- “We might want it later,” where questions get added for a future analysis that may never happen
- “It would be nice to know,” where curiosity drives the request more than necessity
- Boilerplate habit, where a demographic block gets copy-pasted from a prior instrument without anyone asking whether each item is needed for this study
None of that is nefarious. But the consequences are real, and they land on participants, not on the organizations doing the collecting.
The Real Costs of Overcollection
Every additional question is a withdrawal from a limited account: participant goodwill. Longer surveys and intake forms increase respondent burden, drive down completion rates, and — particularly in populations who are already asked to disclose a lot to receive services — can feel extractive. Trust is built or eroded well before analysis begins.
Overcollection also introduces what we sometimes call “hot” data: information that is more sensitive, more identifying, or more legally consequential than the evaluation actually requires. Every Social Security number, exact birthdate, or detailed health disclosure you collect but don’t strictly need is pure downside. It adds breach liability and regulatory exposure without adding analytic value.
The examples above aren’t outliers, either. 2024 alone saw nearly 2,000 federal data-privacy lawsuits filed in the U.S., and Meta’s $1.4 billion settlement with Texas over unlawful biometric data collection now stands as the largest privacy settlement in the country’s history. The throughline in all of these stories isn’t hacking sophistication; it’s that the data existed to be stolen or misused at all.
Don’t Just Ask Fewer Questions. Elicit Less Identifiable Data.
Data minimization isn’t only a question of how many items are on your survey. It’s also a question of how identifiable your dataset is, even if every item on it seems justified.
Demographic questions are the clearest example. Age, race, gender, zip code, job title, and years of service are practically boilerplate in evaluation instruments, but each one narrows the pool of people a response could belong to.
Foundational disclosure-risk research from Latanya Sweeney found that ZIP code, birth date, and sex alone could uniquely identify most of the U.S. population, and a widely cited Nature Communications study estimated that just 15 demographic attributes could correctly re-identify 99.98% of Americans in a supposedly “anonymized” dataset.
Now shrink that population down to a single program, department, or small organization. Evaluate an internal staff survey of 40 employees asking for age range, race, tenure, and department, and you may create a dataset where several respondents are the only person who fits that combination. They could be easily re-identified by a supervisor or colleague with a bit of institutional knowledge. That risk climbs even higher in small, close-knit service populations, exactly the setting many community-based evaluations operate in.
Now, we’re not suggesting that you abandon demographic data collection. It’s usually valuable and sometimes required by funders. We’re suggesting that you ask, for every demographic item: Do we need this exact level of granularity, and does our population size support reporting it safely? Broader age bands, collapsed race/ethnicity categories, or suppressing small cells in reporting can preserve analytic value while dramatically reducing re-identification risk.
What Minimization Looks Like After Collection
Data minimization doesn’t retire once collection wraps. The highest-stakes decisions often happen after the data has already done its job.
Think back to that aging services non-profit. That breach applied to files the organization had been holding for an extended period. Long-held data is what makes a breach catastrophic instead of merely bad. The longer sensitive information sits in a system, the more accounts touch it, and the more opportunities exist for something to go wrong.
Minimization after collection means treating “we still have it” as a liability to actively manage.
Narrow access and de-identify as the project moves forward
During active collection, a wider team often needs access — enumerators, data entry staff, coordinators. Once a dataset moves into analysis and reporting, that circle should shrink, and identifiers no longer needed should be stripped out, with a separate, controlled key linking back to identities only where still required. Access lists that never get revisited are a common, preventable source of exposure.
Safeguard what remains with real controls
Encryption at rest and in transit, role-based access, and secure file sharing are the baseline. The FTC’s own guidance is blunt: making conscious choices about what data you collect, how long you keep it, and who can access it is what actually reduces breach risk. Technical safeguards are a backstop for data that shouldn’t still be sitting there in the first place.
Set a retention schedule before you need one, and destroy data deliberately
“We might need it later” drives over-retention just as often as it drives overcollection. A defensible schedule is written down before collection starts, ties retention to a specific funder requirement, IRB protocol, or legal obligation — not a vague sense of future usefulness — and is followed, with timely, documented destruction. Data you no longer have can’t be breached, subpoenaed, or misused, no matter how good your intentions were when you first collected it.
Where Your IRB Fits In
If you’ve worked with an Institutional Review Board, data minimization should sound familiar. It’s baked into the ethical review process, particularly around minimizing risk to participants and ensuring data handling is proportionate to the study’s purpose. It’s part of why IRB reviewers push back on instruments that ask for more than a study needs, and it’s a theme we’ve unpacked in pieces on what happens after you submit an IRB application and why organizations benefit from an independent IRB.
Treating data minimization as a standing design principle is one of the clearest ways evaluators and applied researchers can practice sound research ethics day to day.
Building The Data Minimization Habit
Data minimization isn’t a single decision made during survey design. It’s a discipline that touches instrument design, sampling, demographic categorization, data storage, and destruction schedules. The organizations that do it well build it into their standard operating procedure rather than relitigating it project by project.
If your organization is designing an evaluation plan, revising a long-standing data collection tool, or trying to build out sound data management practices from the ground up, this is exactly the kind of work we help with. We’d love to help your organization build ethical, minimization-minded research and evaluation practices, the kind that protect your participants and your organization in equal measure.

