logoalt Hacker News

Rampart: Browser native on-device PII radaction

58 points • by nateb2022 • yesterday at 5:52 PM • 25 comments • view on HN

Comments

dwa3592 • today at 3:51 PM

I have worked in this field and I am the author of this package - https://github.com/deepanwadhwa/zink

A few things jump out since this is done by the government:

- the lowest hanging fruit for this problem is to clearly tell people (citizens) not to share any personal info with chatbots which can cause financial harm or identity theft. the example on the page shows a person sharing their SNN with a chatbot to help them find an apartment - "My name is Maria Garcia, my Social Security number is 123-45-6789, and I make $1,950 a month. Can you help me find affordable housing?" - why?? this is the opposite of what i would expect a government to advise their citizens.

- it's never too late for a good policy; the government should have extended HIPPA and other data privacy laws to AI companies - the AI company must not store anyone's SSN, no matter how stupid the user is. It should be on the AI company to not store it; so this type of layer should be on the AI company's side.

- technical; there are quasi identifiers of privacy (that's what my package targets) that are asymptotically hard to to deal with - meaning - if you remove everything that can leak your privacy the text would become meaningless. i don't think rampart can solve for that either and it should be clearly said on the website.

bob1029 • today at 3:51 PM

I have presented approaches like this to banking clients and they are still not very interested. The only thing that makes these people happy is zero data retention and deterministic redaction at the source. Regex over arbitrary string literals does not represent determinism in this context.

If your product is handling natural language conversations from end customers, there is not much you can do to prevent the occasional PII leak without ruining the rest of the pie. ZDR is your best mitigation if you actually want the magical AI experience to work the way the investors hope it can.

PII can often become disclosed by way of many correlated factors that are not considered PII on their own. Even a perfect AI system cannot capture all of these relationships. You could probably locate where I live within a 20 mile radius if you spent enough time analyzing my HN comments over the years. Not one of these comments on their own would trigger a PII filter.

throw03172019 • today at 8:27 PM

Plain text in a chat input is only one piece of the problem. What about files like PDFs and documents filled with PII.

iAMkenough • today at 3:25 PM

So that’s why big ballz (co author of the OP) put our PII in an insecure AWS instance via DOGE’s starlink terminal

https://www.csoonline.com/article/4046997/whistleblower-doge...

nhinck2 • today at 3:54 PM

98.4% is nowhere near good enough to call it PII redaction.

handfuloflight • today at 4:53 PM

Why did the National Design Studio see the need to put all the readable text on the right column of the page?

swiftcoder • today at 5:05 PM

What we'd really like is a PII redaction model for video...

Onavo • today at 5:43 PM

Is this good enough for HIPAA?

➕ show 1 reply
Cider9986 • today at 4:36 PM

The source code should be public domain, no?

➕ show 1 reply
sgnelson • today at 3:35 PM

The skeptic in me really can't trust our current government to protect my information.

➕ show 2 replies
theritik12ee1 • today at 9:11 PM

[flagged]