A few years ago I inherited a payments service that stored full customer date-of-birth, national ID, and unmasked PAN fragments in the same table it used for feature flags. Nobody had done anything malicious. It just accreted, one migration at a time, because storing a field is always easier than deciding whether you should. The day a regulator asked us to produce every place a single customer's data lived, we spent eleven days answering. That is the real cost of doing personal data badly: not a fine, but the slow discovery that you don't actually know what you're holding.
Handling personal data in .NET is not a library you install. It is a set of decisions you make early and keep making. The tooling is good now, better than it was, but the framework will happily let you do the wrong thing. So this is the opinionated version of what I make my teams do, and why.
Know exactly what you hold
You cannot protect data you can't name. Before any encryption or fancy access control, the first job is a data inventory: every field that identifies a person, where it lives, why you have it, and how long you keep it. This sounds like busywork. It is the single most valuable thing you can do, and almost nobody does it until forced.
In practice I make this concrete in the code itself. I do not want the classification living in a Confluence page that goes stale within a quarter. I attribute the model directly, so the classification travels with the field and shows up in code review when someone adds a new one.
[AttributeUsage(AttributeTargets.Property)]
public sealed class PersonalDataAttribute : Attribute
{
public DataSensitivity Sensitivity { get; }
public PersonalDataAttribute(DataSensitivity s) => Sensitivity = s;
}
public enum DataSensitivity { Low, Standard, Sensitive }
public class Customer
{
public Guid Id { get; set; }
[PersonalData(DataSensitivity.Standard)]
public string Email { get; set; } = "";
[PersonalData(DataSensitivity.Sensitive)]
public string NationalId { get; set; } = "";
}
Once the attribute exists, you can reflect over your DbContext at startup and fail the build if a property that looks like personal data (a field named Email, Ssn, Dob) is not annotated. I have a small unit test that does exactly that. It has caught three fields in two years that would otherwise have slipped in unclassified. Cheap insurance.
Collect less than you are allowed to
The most secure personal data is the field you never stored. I am fairly hardline on this because I have watched the alternative play out. Product asks for date of birth "for personalization". Nobody ever builds the personalization. Five years later it is a liability sitting in six backups, and deleting it is now a project.
So the default answer to "should we store this" is no, and the person who wants it has to justify the retention, not the other way round. If you need age verification, store a boolean that says the check passed and the date it was performed. You almost never need the birth date itself. If you need to contact someone, an email is usually enough; you rarely need a phone number too. Minimization is not just a compliance nicety. It directly shrinks the blast radius of any breach.
The cheapest data to secure, audit, delete, and reason about is the data you decided not to collect in the first place. Every field is a standing cost, not a one-time one.
Encrypt in the right place, not everywhere
Transparent Data Encryption at the database level protects you from someone walking off with the disk. It does nothing against a compromised app connection string or an over-broad query. For genuinely sensitive fields — national ID, bank account numbers — I want application-level encryption so the plaintext never touches the database at all. In .NET this is clean with EF Core value converters backed by a key from Azure Key Vault or AWS KMS.
var converter = new ValueConverter<string, byte[]>(
plaintext => _protector.Encrypt(plaintext),
ciphertext => _protector.Decrypt(ciphertext));
modelBuilder.Entity<Customer>()
.Property(c => c.NationalId)
.HasConversion(converter)
.HasColumnType("varbinary(max)");
The tradeoff is real and you should feel it: an encrypted column cannot be indexed or searched in the usual way. If you need equality lookups on an encrypted field, keep a separate blind index — a keyed HMAC of the normalized value — and query on that. Don't reach for column-level encryption reflexively on every field. It has a cost in query flexibility and in key-rotation complexity, and applying it to low-sensitivity data just makes your system harder to operate without buying you much.
Keep personal data out of your logs
Logs are where privacy discipline quietly dies. You do everything right in the database, then someone writes logger.LogInformation("Processing order for {Customer}", customer) and the whole object, email and all, is now sitting in plaintext in your log aggregator with a 400-day retention and access for the entire on-call rotation. I have seen a clean database undermined completely by a Splunk index nobody thought of as a data store.
Enjoying this article?
Get more like it in your inbox — practical engineering leadership, fintech, and AI. No spam, unsubscribe anytime.
The fix is structural, not vigilance. Vigilance fails. I standardize on log types that never serialize raw personal data, and I add a destructuring policy to Serilog that redacts anything marked with the PersonalData attribute. The customer object logs as an id and nothing else unless someone very deliberately overrides it.
- Never log a whole entity, log its id and let people join to the record if they have the access.
- Redact by default in the logging pipeline so a careless log statement fails safe.
- Treat exception messages and stack traces as personal-data hazards too, since parameters leak into them.
- Scan your log aggregator with a regex for email and card patterns on a schedule. You will find things.
Deletion has to actually work
The right to erasure is where a lot of otherwise-tidy systems fall over. Soft-deletes with an IsDeleted flag are fine for undo, but they are not deletion, and pretending otherwise in front of an auditor is a bad afternoon. When someone exercises their right to be forgotten, the data has to be genuinely gone or genuinely anonymized, including in the places you forget about: read replicas, search indexes, event streams, and backups.
Backups are the awkward one. You usually cannot surgically delete one person from a point-in-time backup without rendering it useless. The defensible position most regulators accept is a documented backup retention window, say 35 days, after which the data ages out naturally, combined with a rule that you never restore a deleted subject from an old backup into production. Write that policy down before you need it. My rule of thumb: if you can't explain in one paragraph how a deletion request propagates through every store you own, you can't honor one.
Pseudonymize, but don't pretend it is anonymous
There is a persistent confusion I keep having to correct. Hashing an email is not anonymization. If I can take a candidate email, hash it, and check whether it matches, the data is still linkable, which means it is pseudonymous and still personal data under most regimes. Anonymization means the link is irreversibly broken, and it is much harder than people think, especially with rich behavioral data where a handful of data points re-identify someone.
This matters because teams reach for a SHA-256 of an email and declare the analytics table "anonymous", then feel free to keep it forever and share it widely. It isn't, and you can't. For analytics I prefer proper pseudonymization with a separately-held mapping key, or genuine aggregation where individual rows are gone entirely. Be honest with yourself about which one you actually built.
Access is a privacy control, and audit it
Encryption gets the attention, but most real incidents I have dealt with were not clever attacks. They were an internal account with more access than it needed, used carelessly or curiously. A support agent who can read every customer's full record is a bigger day-to-day risk than a hypothetical attacker breaking your ciphers.
So access to personal data should be least-privilege and, for sensitive fields, logged as an event in its own right. When a support tool decrypts someone's national ID, that read should produce an audit record: who, when, which customer, and ideally why (tie it to a ticket). This is not paranoia. The first question after any suspected internal misuse is "who looked at this record", and if you can't answer it you are guessing. Build the audit trail before you have a reason to want it, because you cannot reconstruct it after the fact.
Test with fake data, always
The fastest way to turn one copy of personal data into twelve is to seed staging and local databases from a production dump. It is convenient and it is a slow-motion incident. Now every developer laptop and every lower environment, none of which have production's controls, is holding real customer data. I ban it outright.
Generating realistic synthetic data is genuinely easy in .NET, so there is no excuse. Bogus and similar libraries produce believable names, emails, and addresses that look right in demos and screenshots without being anyone real. Where you truly must use production-shaped data for a hard-to-reproduce bug, mask it on the way out of production, not on the way in to staging, so the plaintext never leaves the secure boundary.

Conclusion
If I had to compress all of this into one habit, it is this: treat every piece of personal data as something you are borrowing, not something you own. Borrowed things you keep track of, you handle carefully, and you give back when you're asked. That framing quietly answers most of the specific questions, whether to store a field, how long to keep it, who gets to see it, because the answer is always the more conservative one. The .NET tooling will support you once you have decided, but it will never decide for you, and the teams that get burned are the ones waiting for a framework to make a judgment that was always theirs to make.
Get new posts in your inbox
Occasional, practical notes on engineering leadership, fintech, and building with AI. No spam, unsubscribe anytime.
Comments (0)
Leave a Comment
No comments yet. Be the first to comment!
