Show Me the Code Path: Why Open Source Wins the PII Redaction Audit
There is a question that ends a lot of redaction vendor demos. The vendor says their AI catches 99% of PII, the slide looks great, and then someone on the security team asks the only question that matters: “Show me how.”
A closed-source vendor cannot answer that question with anything but more words. They can show you a datasheet, a benchmark they ran themselves, a SOC 2 report about their company rather than their detection logic. What they cannot show you is the actual code that decided a given nine-digit number was a Social Security number and not an account ID. For a buyer who will have to defend that decision to an auditor, a regulator, or a court, “trust the datasheet” is not an answer. It is a liability they are being asked to absorb.
Open source answers the question directly. The code path is the answer.
“Show me how” is the whole job for regulated buyers
For a lot of teams, redaction is not the deliverable. The defensible account of how redaction happened is the deliverable. The redaction is just the thing the account is about.
- Government (FedRAMP, CMMC, FISMA). These frameworks are built around documenting control implementations and proving they work as described. When an Inspector General or an assessor asks how a data-handling control is implemented, “here is the source for the detection rule, and here is the policy that governs it” is a control narrative you can stand behind. “The vendor assured us it works” is the kind of sentence that turns a finding into a problem.
- Healthcare (HIPAA). Safe Harbor de-identification under 45 CFR 164.514(b) requires removing 18 categories of identifiers. Demonstrating that your process reliably does that is much easier when the detection logic is readable than when it lives behind an API you can only probe from the outside. Auditable source satisfies a privacy officer in a way API documentation never quite does.
- Legal (FRBP 9037, FRCP 5.2). An attorney filing redacted documents has to certify the redactions are complete, and that certification is personal. An audit trail that traces a redaction back to inspectable detection logic is protection for the person whose name is on the filing. A black box is the opposite: you are certifying something you were structurally prevented from verifying.
In every one of these cases the burden of proof sits with the buyer, not the vendor. So the buyer needs tooling that produces proof, not assurances.
A contractual promise is not an audit trail
The closed-source counterargument is usually some version of “we are certified, we are audited, here are our attestations.” Those things are real and they have value. They are also attestations about the vendor’s organization, not about the line of code that made a specific redaction decision in your pipeline. When the question is “why was this token redacted and that one missed,” a corporate certification does not contain the answer. The detection code does.
This is the same structural point we make in Open Source vs Black Box “Trust Me” Privacy: a promise is something you enforce after a failure, by audit and argument. Inspectable code is something you verify before you deploy, and keep verifying. One is a contract. The other is evidence.
What auditability actually looks like in the Philterd toolkit
“Open source” on its own is necessary but not sufficient. What makes a redaction pipeline defensible is that every link in the chain, from detection to output, is something you can read and reproduce:
- Every detection rule is in source you can read. The regex patterns, the format validators (Luhn checks, SSN structure, date validity), the dictionaries, and the policy conditions are all in the Phileas library under Apache 2.0. There is no hidden classifier making decisions you cannot inspect.
- Every model is purpose-built and measurable. The NLP models that find names and clinical entities are trained on synthetic and public data, and their precision and recall are measurable against your own gold-standard set with Philter Scope. You do not have to take a 99% claim on faith; you can produce the number yourself and put it in the audit file.
- Every redaction decision is logged. The output is not just redacted text. It is redacted text plus a report of what was found, where, with what confidence, and which policy rule applied. That report is the audit artifact.
- The whole chain is reproducible. Because Philter runs inside your environment and the code is public, an auditor can rebuild from source, run the test suite against their own inputs, and confirm the behavior independently.
That is the difference between auditable and audited. An audited system has a report about it. An auditable system lets you generate the proof yourself, on demand, for the specific decisions in question.
The competitive landscape
It is worth being concrete about where the alternatives sit:
- AWS Comprehend, Google DLP, and Private AI are closed-source services. They may detect PII well, but you cannot read the detection logic, you cannot run it inside an air-gapped boundary without their infrastructure, and when an auditor asks “show me how,” you are back to documentation.
- Microsoft Presidio is genuinely open source, which is a real point in its favor. What it does not bring is the domain-specific models, the policy engine, the benchmarking tool, and the broader toolkit (discovery, monitoring, format-preserving encryption) that a regulated production deployment needs. Open source is the floor, not the whole building.
Philterd’s position is to be open source and to ship the rest of what a serious deployment requires, so the auditability is not a hobbyist’s “you could read it if you wanted” but a practical “here is the code path, here is the measured accuracy, here is the log.”
The license is part of the argument
There is also a quieter advantage that compliance teams feel immediately. The Apache 2.0 license removes procurement friction. No commercial-use review, no per-seat negotiation, no lock-in concern to escalate. A legal team can read the license and approve it in a single meeting, which means the security and compliance evaluation can focus on the actual question (is the detection sound) instead of the contract.
The bottom line
For a buyer who has to defend their redaction to someone else, the deciding factor is not which datasheet claims the highest accuracy. It is which tool lets them answer “show me how” with something an auditor will accept. Closed source answers with documentation. Open source answers with the code path, the measured numbers, and the log.
If you are evaluating redaction tooling for a regulated workload, start with the compliance matrix and the government and federal guide, read the code on GitHub, or join the community and ask the engineers who wrote it.