The redaction runs where your data already is
Philter is deployed into your own cloud account and inserted into the pipeline you already have, so sensitive text is removed in transit and nothing is shipped to a third party to be processed. Where the step sits is below.
This page walks through what a PII data pipeline engagement with us actually looks like, from the first call to handoff. It is a representative example to help you picture the work and scope it, not a fixed package. Every engagement is shaped to your data, your systems, and the regulations you answer to.
The premise is that sensitive text is already moving through your systems and you need it redacted somewhere along the way. That might be an EHR export landing in an analytics warehouse, a Kafka topic carrying support transcripts, a nightly batch of documents, or an API that has to strip identifiers before a downstream service sees them. The redaction runs inside your own cloud, so the data never leaves your boundary and never comes to us.
When this engagement fits
This is the right engagement if you can point at a specific flow of data and say what has to be removed before it reaches the other end, and you can run software in your own environment. It fits teams who have the pipeline already and need the redaction step designed, built, and measured, and teams who want the design and plan so their own engineers can build it. It is not a fit if you want a hosted service that ingests your raw data, because keeping the data inside your perimeter is what makes the rest of it defensible.
How the engagement runs
After the intro call, the work follows the same Discovery, Implementation, and Handoff as any Philterd engagement. Phase 1 is Discovery. Phases 2 through 4 are what Implementation and Handoff look like for a production pipeline. The shape holds whether you are redacting one stream or a dozen. What changes is how long each phase takes.
- Discovery. Over a few meetings we map the flows. What data moves, in what formats, at what volume, between which systems, and which regulations apply at each hop. We agree on what has to be removed, what has to survive for the downstream system to still work, and what “good enough” means before anyone writes code. You get a written assessment and an implementation plan, and the plan is yours whether or not we build it.
- Policy design and gold standard. We author the redaction policy against your actual data, deciding per field whether a value is redacted, masked, or consistently pseudonymized so that joins and relationships survive downstream. Alongside it we build a labeled gold-standard sample from your own records, which is what makes the next phase measurable rather than a matter of opinion.
- Build and measure. We deploy Philter into the pipeline inside your cloud, wiring it in where it belongs, whether that is a Kafka Connect transform, a batch step, or a service call. Then we score the policy against the gold standard with Philter Scope and iterate until it meets the targets we agreed in Discovery. Cases the models are unsure about can route to human reviewers through Arbiter.
- Cutover and handoff. We move it onto live data, watch it under real traffic, and hand over the policy, the deployment configuration, the measurement report, and a runbook for tuning it as your data changes. If you want ongoing visibility, Phield can monitor PII counts over time and alert when the shape of the traffic shifts.
On timelines: the phases above are deliberately not given fixed durations. How long each one takes depends on how many flows are in scope, how messy the data is, how many systems have to be touched to insert the redaction step, and how much review the uncertain cases need. We give you a realistic schedule during Discovery, and it can change as the work reveals what the data actually contains.
What you keep
The point of the engagement is that you own the result and can run it without us.
- The Discovery assessment and implementation plan, yours whether or not we do the build.
- A tuned, documented redaction policy written against your data.
- A precision and recall report scored on your gold standard, which is the artifact compliance teams ask for.
- The deployment running inside your own account, on open source software you can read.
- A runbook for re-tuning as the data drifts, so your team can adjust it as things change.
Detection is probabilistic, so the measurement report is the honest picture of where the policy stands rather than a promise that nothing gets through. Part of the handoff is making sure your team knows how to re-measure when the data changes.
How your data is handled
The toolkit runs inside your account, so there is no third-party redaction service to give read access to your pipeline. Most of the work happens against a labeled sample and on policy design rather than against your full production data. Where an engagement genuinely needs us to work with real records, how that data is handled is spelled out in the engagement contract before any work starts. See how we think about data handling in engagements.
Pricing
Engagements are scoped to the flows in scope and run as a fixed-scope project, with an optional retainer for ongoing tuning as the data changes or as new pipelines come along. Because the tooling is the open source stack (Philter, Phileas, and Philter Scope), there is no per-record data-processing fee and no software license to buy. For a number, the best path is the intro call.
This page describes a representative engagement. It is illustrative, not a statement of work, and the actual scope, sequence, and schedule are set with you before work begins.