Rampart: Browser native on-device PII radaction

ndstudio.gov

54 points by nateb2022 a day ago


dwa3592 - 5 hours ago

I have worked in this field and I am the author of this package - https://github.com/deepanwadhwa/zink

A few things jump out since this is done by the government:

- the lowest hanging fruit for this problem is to clearly tell people (citizens) not to share any personal info with chatbots which can cause financial harm or identity theft. the example on the page shows a person sharing their SNN with a chatbot to help them find an apartment - "My name is Maria Garcia, my Social Security number is 123-45-6789, and I make $1,950 a month. Can you help me find affordable housing?" - why?? this is the opposite of what i would expect a government to advise their citizens.

- it's never too late for a good policy; the government should have extended HIPPA and other data privacy laws to AI companies - the AI company must not store anyone's SSN, no matter how stupid the user is. It should be on the AI company to not store it; so this type of layer should be on the AI company's side.

- technical; there are quasi identifiers of privacy (that's what my package targets) that are asymptotically hard to to deal with - meaning - if you remove everything that can leak your privacy the text would become meaningless. i don't think rampart can solve for that either and it should be clearly said on the website.

bob1029 - 5 hours ago

I have presented approaches like this to banking clients and they are still not very interested. The only thing that makes these people happy is zero data retention and deterministic redaction at the source. Regex over arbitrary string literals does not represent determinism in this context.

If your product is handling natural language conversations from end customers, there is not much you can do to prevent the occasional PII leak without ruining the rest of the pie. ZDR is your best mitigation if you actually want the magical AI experience to work the way the investors hope it can.

PII can often become disclosed by way of many correlated factors that are not considered PII on their own. Even a perfect AI system cannot capture all of these relationships. You could probably locate where I live within a 20 mile radius if you spent enough time analyzing my HN comments over the years. Not one of these comments on their own would trigger a PII filter.

throw03172019 - an hour ago

Plain text in a chat input is only one piece of the problem. What about files like PDFs and documents filled with PII.

simontwiggyahuk - 2 hours ago

https://discuss.google.dev/t/trusted-automation-with-google-...

iAMkenough - 6 hours ago

So that’s why big ballz (co author of the OP) put our PII in an insecure AWS instance via DOGE’s starlink terminal

https://www.csoonline.com/article/4046997/whistleblower-doge...

nhinck2 - 5 hours ago

98.4% is nowhere near good enough to call it PII redaction.

handfuloflight - 4 hours ago

Why did the National Design Studio see the need to put all the readable text on the right column of the page?

swiftcoder - 4 hours ago

What we'd really like is a PII redaction model for video...

Onavo - 4 hours ago

Is this good enough for HIPAA?

Cider9986 - 5 hours ago

The source code should be public domain, no?

sgnelson - 6 hours ago

The skeptic in me really can't trust our current government to protect my information.

theritik12ee1 - 9 minutes ago

[flagged]