Open Data, Privacy, and Ethical Data Collection
Being able to collect data is not a reason to collect it
A fictional city bus operator wants to reduce crowding on its busiest morning route. Technically it can obtain a great deal: every tap of every fare card with stop and time, vehicle positions every few seconds, and counts from door sensors. Four questions are hiding inside “can we use this data?”, and they have different answers. Can we technically obtain it? Is it good enough for the question? Are we permitted to use it this way, by law, licence, or contract? And is using it this way appropriate for the people it describes? A dataset can pass the first two and fail the last two. A source being reliable and well documented does not justify a purpose it was never collected for.
Data ethics is the set of standards and working practices that guide how data is collected, stored, analysed, shared, and used, with attention to the people it affects. Personal judgment is not enough on its own, because analysts carry their own assumptions and blind spots; written standards, review by others, and records of decisions are what make the judgment checkable. The UK government’s Data Ethics Framework, written for public-sector projects, organises its guidance around transparency, accountability, and fairness and asks teams to define the public benefit, understand the applicable law, use no more data than the purpose needs, and be open about how data is used. Those questions travel well beyond government, provided the legal specifics are checked for your own jurisdiction. Everything in this article is a fictional exercise; the operator, its riders, and its numbers are invented.
Decide the purpose and the smallest data before collecting
Write down, before anything is collected, what the data is for, which fields that purpose actually needs, whose data it is, how it will be obtained, what those people would reasonably expect, who else could be affected, and what happens if you collect nothing new. For the crowding question, the answer is smaller than the operator’s ambitions.
| Question | Answer for the crowding project |
|---|---|
| Purpose | Find which weekday morning departures on route 12 exceed comfortable load, and by how much |
| Minimum fields | Boardings and alightings per stop per departure; vehicle capacity. Not needed: card identifiers, fare type, home stop |
| People described | Weekday riders of route 12 during the morning peak |
| Collection method | Counts already produced by door sensors and validators, aggregated at the source |
| Reasonable expectations | Riders expect taps to be used for fares and service planning; they do not expect individual travel histories to be studied or shared |
| Others affected | Drivers, whose vehicle positions can reveal their working pattern |
| Alternative without new collection | Two weeks of manual counts by staff on the four busiest departures |
This is data minimization applied before the fact: the purpose sets the fields, not the other way round. It is also purpose limitation: the tap records exist for fares, and using them for a new purpose is a new decision that needs its own justification, not an automatic extension of the old one. Both are part of data privacy, which is about how handling information affects people, not only about keeping it secure. The manual-count alternative matters because it shows the price of not collecting more. If it would answer the question at acceptable cost, the case for building a new collection path is weaker than it looked.
Be precise about rights and roles, because loose words here cause real mistakes. The statement “people own the raw data they provide” is not a description of the law in general. Depending on the jurisdiction, a person may have rights over information about them, such as to be informed, to access it, to object to some uses, or to have it deleted, while the organization that holds the records has duties of care and accountability, and copyright, database rights, or contracts may govern the dataset as a whole. An internal data owner role is an accountability assignment inside the organization; it is not a property right over riders’ information and does not override their rights. Treat each of these as a separate question to check, not as one idea called ownership.
Consent needs the same precision. Showing a privacy notice informs; it does not by itself record agreement, and a checkbox next to lengthy terms is weak evidence that a person agreed to a specific use. Consent is also not the only possible basis for handling personal data: in the European Union, for example, the GDPR lists several legal bases, of which consent is one, and other jurisdictions have their own rules. Where consent is the basis you rely on, the person must be told the purpose, be able to say no without losing the service where that is required, understand the scope, and be able to withdraw, and the withdrawal must actually change what the systems do. Which of these apply, and how, is a question for the applicable law and an adviser who knows it, not for this article.
Commercial use deserves its own line in the plan. If tap data will be used to sell advertising space at stops, or trip patterns will be licensed to a property developer, riders bear the exposure and the operator gains the revenue. Say so plainly: who benefits, who carries the risk, and what choice riders have. Whether a person is entitled to compensation or a right to refuse depends on the jurisdiction and the legal basis; what does not depend on jurisdiction is that hiding a commercial use behind “service improvement” fails the transparency test.
Say what you do, then make the system do it
Transparency is a description that a rider could read and a system that behaves the same way. The description covers what is collected and why, who inside the organization uses it, which outside parties receive it, whether any use is commercial, how long it is kept, how to ask questions or make requests, and how a person can change a choice. The system has to enforce each of those statements. A notice that promises deletion after a year is a false statement if the analytics copy keeps the records for three. Keeping the promise means the retention clock, the deletion path across copies, and the preference flag all exist and are checked.
When the team cannot answer one of these questions, the honest entry is “unknown, owner: service manager, to confirm before collection starts,” not a reassuring sentence. A list of open questions with named owners is a normal part of an ethical collection plan; a plan with no open questions usually means nobody looked.
Open means more than downloadable
Open data is data that anyone is free to access, use, modify, and share, subject at most to conditions that preserve provenance and openness, such as attribution and share-alike. That is the Open Definition maintained by the Open Knowledge Foundation, and the Open Data Handbook based on it describes three properties: availability and access as a whole at no more than a reasonable reproduction cost, preferably as a download; permission to reuse and redistribute, including combining with other datasets; and universal participation, with no discrimination against persons, groups, or fields of endeavour. Being visible on a web page or free to download is not enough, and a licence alone is not enough either. The Open Definition requires four things together: the work is in the public domain or under an open licence; it is accessible as a whole at no more than a reasonable one-time cost, preferably as a download; it is machine-readable; and it uses an open format. Data is open when all four hold.
| State of the data | Can you reuse and republish it? | Is it open data? |
|---|---|---|
| Viewable on a public web page, no stated terms | Unclear; viewing is not a licence to copy or republish | Not shown to be; the rights status is unknown and must be checked, since the data could be public domain or fully reserved |
| Free download, terms allow personal or non-commercial use only | Only within those terms; a non-commercial restriction is a field-of-use limit | No |
| Provided to approved researchers under an agreement | Only by those researchers, for the agreed purpose | No; this is controlled sharing |
| Published under an open licence, attribution required | Yes, by anyone, for any purpose, with credit to the source, if the licence really grants those permissions | Yes, provided the data is also accessible as a whole, machine-readable, and in an open format |
The Open Definition separates two kinds of condition. Attribution and share-alike are acceptable conditions: they preserve the source and keep derived works open. A restriction on the field of use, such as “research only” or “no commercial use,” is not compatible with the definition, however reasonable it may be for other purposes. Both kinds of terms can be legitimate choices; only one of them produces open data. Read the licence text, not the word “open” on the portal, and check that attribution or share-alike is the only condition; a licence that requires attribution and also restricts reuse is not open.
Openness is not the same as ethical release, and the two should not be run together. An open licence answers who may reuse the data; it says nothing about whether the data should have been released at all. The Open Data Handbook itself notes that open data is concerned with non-personal data, data that does not identify individuals. Publishing personal records under an open licence does not make the publication acceptable; it makes it irreversible. Nor is interoperability the same as openness: interoperability is the technical and semantic ability of systems to exchange data and use it, and it is valuable inside closed arrangements too. Hospitals, clinics, and pharmacies exchanging a patient’s prescription under authorization and agreed standards are interoperating; they are not publishing open data, and the example belongs to controlled sharing.
Sharing has benefits that are easy to list and costs that are easy to forget. Open route, stop, and timetable data lets journey-planning apps, researchers, and residents build on it, and it lets the public check the operator’s claims. The costs are the infrastructure to serve it, the work to keep it current and documented, the support requests it generates, and the standards work needed so that other operators’ data can be combined with it. A release that nobody maintains decays into misinformation with a licence attached.
Removing names does not make a release safe
Suppose the operator replaces each card identifier with a stable code and keeps stop, date, and minute for each boarding and alighting, so every trip of the same card still carries the same code. Consider one code that boards at stop 41 at 05:52 every weekday, alights at the stop beside the regional hospital, and returns at 19:40. No name is present, but the code links those trips into one trail, and a neighbour, an employer, or anyone who knows the person may recognize the pattern; a second dataset, a staff roster or a social-media post, may support the guess, although such sources confirm an identity only sometimes and the link remains an inference. Strip the code as well and the trips become unlinked events. Most rows are then much harder to attribute, yet a rare combination, the only boarding at a rural stop at 05:52, can still be tied to a person by someone who saw it happen. This is re-identification risk: the chance that a record can be linked back to a person through precise location, time, rare characteristics, or combination with other data. Direct identifiers are the easy part; the combination of ordinary fields is the hard part.
Keep the techniques apart, because their names get swapped. Pseudonymization replaces identifiers with substitutes; a pseudonymous card number still links every trip of the same person and is still personal data in many frameworks. Masking removes or hides part of a value, so a reader sees less, while the remaining fields and their patterns are untouched. Encryption protects data in storage or transit from people without the key and hides nothing from those who have it. Each reduces a particular exposure and leaves others in place. Anonymization is a claim about an outcome, that people are no longer identifiable in the relevant context, and it has to be assessed for the actual release, recipients, and other data available. Aggregating into counts helps only when the counts are large enough and the categories coarse enough, and suppressing small cells is not safe on its own: a published total of 10 beside a published 9 reveals the suppressed 1, and overlapping tables or an earlier version of the same release can restore hidden values, so the whole set of publications has to be checked together. Synthetic data needs a disclosure-risk assessment that goes beyond confirming that no real record was copied, because a generator can still reveal whether a person was in the source or what a sensitive attribute is likely to be. There is no universal minimum count that makes a cell safe; the threshold, if one is used, is a judgment for this release, and a group of one is a warning to examine rather than an automatic disclosure, since whether it exposes anyone depends on whether that record can be linked to a person and on what the published attributes reveal.
NIST’s guidance on de-identifying government datasets describes the practical shape of this work: study the re-identification risk of the specific dataset, choose among release models such as publishing de-identified data, releasing synthetic data, offering a query interface, or sharing inside a controlled environment, and put a review board over the decision. The lesson for a small operator is the same as for an agency: the release model is part of the privacy design, not a distribution detail.
Choose the release scope, not only yes or no
Between “publish everything” and “share nothing” there are several scopes, and each pairs a benefit with a residual risk. Publish only what the stated purpose needs. Reduce detail, by aggregating, coarsening time or place, or suppressing small groups, until the residual risk is acceptable for public release. Share the detailed data under a written agreement with named recipients, a stated purpose, and deletion at the end. Or hold the data and release nothing while the open questions are resolved. Remember what publication means: a copy that has been downloaded is out of your control, and removing the original from the portal does not recall it. For every release, decide retention, how corrections are issued, how deletion requests are handled where they apply, how users are told about changes, and which licence and version each file carries.
Here is the operator’s decision table for four candidate datasets. It is a worked example of the reasoning, not a template of correct answers; a different city with different laws, riders, and neighbours could reach different conclusions with the same table.
| Dataset | Purpose and expectation | Residual risk after reduction | Decision and conditions |
|---|---|---|---|
| Routes, stops, timetables | Journey planning, public accountability; riders expect this to be public | None found to individuals after checking, for this fictional case, that stops and timetables do not reveal a private residence, an employee’s schedule, or a sensitive facility beyond what is already public; errors mislead travellers | Publish openly under an attribution licence; version and change log; maintainer named |
| Boardings per stop per departure, weekday averages over one month | Crowding analysis, app developers; riders expect service statistics | Low after aggregation; departures with very few riders at a rural stop could still single out a person | Publish openly after suppressing cells below the count the review agreed for this release and checking that totals, overlapping tables, and earlier versions cannot restore them; document the rule and the residual-risk review |
| Per-card trip records with card identifiers replaced by codes | University study of transfer behaviour; riders do not expect individual histories to leave the operator | High; stop-and-time patterns identify regular riders; codes still link trips | Do not publish. Share under a written agreement: named researchers, stated purpose, minimized fields, no re-identification attempts, deletion at project end, legal basis confirmed by the privacy officer |
| Vehicle positions every 10 seconds | Real-time arrival information; riders welcome it | Low for riders; drivers’ shifts and breaks become visible, and the depot address is inferable | Publish live positions with a short delay for the public feed; consult drivers’ representatives; withhold historical traces or coarsen them before any release |
Two features of this table are the point of the exercise. First, the same source system produces datasets that receive different decisions, because the purpose, the expectations, and the residual risk differ, not because one file is “anonymized” and another is not. Second, every row names conditions and a decision owner: the service manager decides the first two rows, the privacy officer must confirm the legal basis for the third, and the fourth waits on a consultation. A row that says only “approved” has skipped the reasoning that later reviewers will need.
Make the shared data maintainable
A file in a machine-readable format is the beginning of useful sharing, not the end. Each field needs a stated meaning and unit: is “boardings” a count of taps or of door-sensor detections, and do the two differ? Times need a time basis, local time with the zone stated, and a rule for the day boundary. Each release needs a version, a record of what changed and why, its provenance, a contact, and a person responsible for keeping it current. A CSV file that lacks these can be opened by anyone and understood by no one, and the format alone confers neither meaning nor the right to reuse; the licence does that. Interoperability, the ability of two organizations’ data to be combined and understood the same way, comes from shared definitions and standards, and it is a separate achievement from openness that the operator will need if its counts are ever to be compared with the neighbouring city’s.
Ethical collection and sharing, then, is a sequence of separable decisions rather than one verdict: define the purpose and the minimum, check rights and roles, be transparent and make the systems match, assess re-identification for the actual release, choose a scope with conditions and an owner, and keep what you publish maintained. The exercises below ask you to apply that sequence to three requests the operator receives.
1. A team has removed card identifiers from the tap records and proposes publishing every boarding and alighting with stop and minute under an open licence, “since the identifiers are gone.” Publish, share under conditions, or hold?
Solution
Hold. Removing identifiers addresses direct identification only. If a stable code remains, it links each person’s trips into a trail whose regular pattern identifies them; even with no code at all, rare stop-and-minute combinations can be tied to a person by someone who observed them, and an open licence would make the release irreversible. The assumptions that decide the answer are that the data describes individual trips at minute precision and that recipients may hold other observations about riders. If the purpose is crowding or planning, the aggregated per-stop, per-departure counts in the decision table serve it; if a researcher needs trip-level data, controlled sharing under agreement is the path. Remaining checks: the re-identification assessment for the aggregated release, including whether totals or overlapping tables restore suppressed cells, and the suppression rule the review agrees.
2. A university asks for the per-card trip records “for research.” The research would benefit the city. Is that sufficient to share the data?
Solution
Not on its own. A beneficial purpose does not replace the checks: the specific research question, the minimum fields it needs, whether a legal basis exists for this new use of fare data in the operator’s jurisdiction, what riders were told and would expect, a written agreement naming the recipients, forbidding re-identification attempts and onward sharing, and setting retention and deletion, plus a decision owner who signs it. The answer is “conditional sharing” only if those conditions are met; until the privacy officer confirms the basis and the agreement is signed, it is “hold.”
3. Riders accepted terms and conditions when buying their fare cards. Can the operator now license trip histories to an advertising company?
Solution
Not automatically. Accepting card terms is evidence of agreement to what those terms described, typically fares and service operation, not to a new commercial use of trip histories. Whether such a use could be permitted, and on what basis and with what choice for riders, depends on the applicable law and must be confirmed by someone who knows it. Independently of the legal answer, the transparency test requires the operator to state plainly who benefits, who bears the exposure, and what choice riders have; the ethical judgment also weighs riders’ reasonable expectations and the re-identification risk of trip histories in the recipient’s hands. Without those answers the decision is “hold.”
Discover more from Insightful Data Lab
Subscribe to get the latest posts sent to your email.
