Span disambiguation is one of the least visible features in Phileas, and one of the easiest to misjudge. It is off by default. Turning it on does nothing at all for many policies, and for others it changes which type a value is reported as. Whether it helps you depends on your filters, how long your process lives, and how many instances you run.
This guide is the decision-oriented companion to the Phileas span disambiguation documentation. The docs explain the mechanism. This page is about whether to turn it on and what to watch for when you do.
The problem it solves
When Phileas finds something sensitive, it records a span: a start and end position in the text, a type such as ssn or identifier, a confidence, and the words around it. Each filter in a policy runs independently, so nothing stops two filters from claiming exactly the same characters.
The classic case is a nine-digit number. The SSN filter will claim 123456789. If your policy also has a custom identifier filter that matches nine-digit employee IDs, that filter claims the same nine characters too. You now have two spans covering identical text with different types, and something has to decide which type is right.
The more filters a policy enables, the more often this happens. Filters whose patterns are structurally different never collide, though. A US phone number is ten digits and an SSN is nine, so the same characters are never both. A policy with a few filters that never overlap gains nothing from disambiguation.
What it actually does
Disambiguation does what a person reading the sentence would do: it looks at the surrounding words. In The employee id 123456789 was assigned, the words “employee” and “id” point away from an SSN.
Mechanically, it works like this:
- Phileas takes the words in a window around the span (five tokens on each side by default, set by
span.window.size), lower-cases them, and drops common stop words. - Each remaining word is hashed into a fixed-size vector (512 positions by default, using murmur3).
- For every span that only one filter claimed, Phileas adds that vector to a running total for that type within the current context. These unambiguous spans are the training data.
- For a span that two or more filters claimed, Phileas compares its vector against the accumulated vector for each candidate type using cosine similarity, and picks the closest.
There is no model and no training step you run. It learns as it filters, from the spans that were never in doubt.
A context, here, is the name you pass with each filter request. What Phileas learns is kept separately per context, so a context for HR documents and a context for clinical notes learn different associations.
What triggers it, precisely
Disambiguation only considers spans that cover the same characters (the same start and end) and differ by type. Spans that merely overlap, such as Washington inside Washington State University, are not its job. Those are resolved by overlap dropping, described below.
In current Phileas releases, the competing spans do not need to have the same confidence. Different filters routinely assign different confidences to the same text, and those are exactly the cases disambiguation is meant to settle. Older Phileas releases (including the one inside Philter 3.x) only disambiguated spans whose confidence also matched, and let the higher-confidence span win otherwise. If you are reading older documentation that says confidence must match, that is the behavior it describes.
How it fits with the other mechanisms
Three different mechanisms decide what ends up redacted, and it helps to keep them apart:
- Confidence rules in the policy decide whether a span is acted on at all, or which strategy applies. A strategy
conditionsuch asconfidence > 0.9belongs here. - Span disambiguation decides which type a value is, when two filters claimed identical text.
- Overlap dropping decides which span wins when spans overlap. It always runs, after disambiguation, and prefers the longest span, then the highest confidence, then the highest priority.
So after disambiguation relabels the competing spans with the winning type, overlap dropping still collapses them into one. Disambiguation chooses the type. It does not decide whether the value gets redacted.
A worked example
Suppose a policy has the SSN filter and an identifier filter for nine-digit employee IDs, disambiguation is enabled, and all requests use a context named hr.
Earlier documents in hr have produced unambiguous spans of both types. A hyphenated value like 482-91-3057 next to “SSN” and “social security” was claimed only by the SSN filter, because the identifier pattern does not match hyphens. Values the SSN filter did not claim, near words like “employee”, “badge”, and “assigned”, were claimed only by the identifier filter. Each of those added its surrounding words to its type’s running vector.
Now this sentence arrives:
The employee id 123456789 was assigned to the night shift.
- Both filters claim
123456789, at the same position. That is the trigger. - The window around the span is
the employee idbefore andwas assigned to the nightafter. Stop words drop out, leaving roughlyemployee,id,assigned, andnight. - Those words are hashed into a vector and compared with the
hrcontext’s accumulated SSN vector and identifier vector. - The identifier vector shares
employee,id, andassigned. The SSN vector, built from words like “social” and “security”, shares little or nothing. The identifier type is closer. - Both spans are relabeled as
identifier, and overlap dropping keeps one of them.
If the hr context were brand new, step 3 would have nothing to compare against. In that case Phileas falls back to a deterministic default among the candidate types. That is consistent, but it is not an informed choice. This is the most important practical point in this guide.
It learns, so where it stores what it learned matters
Disambiguation only gets good after it has seen enough unambiguous examples in a context. Everything operational follows from that.
Phileas keeps what it learns in a vector store, supplied by the application that embeds it. Two implementations ship with the library:
- In-memory (the default). Nothing to set up. What it learns is lost when the process stops, and it is not shared with any other process.
- File-based. Loads saved vectors at startup and writes them back when you call
save()or close it. Learning survives restarts. Saving is explicit because writing on every insert would slow filtering down, so pick a cadence, such as periodically and at shutdown.
The store is an interface (VectorService), so an application that needs something else, such as a store shared across machines, can implement it against its own database.
The short-lived process trap
A fresh in-memory store learns from scratch every time the process starts. If Phileas runs inside a short-lived process (a batch job, a serverless function, a container that is replaced often), it may never accumulate enough examples to do better than the default. In that setup, enabling disambiguation adds work for little benefit. Either persist the vectors or leave the feature off.
The multi-instance trap
Run several instances behind a load balancer, each with its own in-memory store, and each one learns only from the documents that happen to reach it. The same ambiguous sentence can then resolve to different types depending on which instance handled it. That is hard to debug, because each instance is behaving correctly on its own data.
The fix is a single store that every instance reads and writes. With Phileas you provide that through your own VectorService implementation. With Philter 3.x, span disambiguation is enabled instance-wide with the span.disambiguation.enabled setting, and a multi-instance deployment should enable Philter’s cache service so the instances share what they learn. That is the same Redis-compatible cache used for referential integrity.
Settings you should not change later
Phileas exposes a handful of settings, listed in the Phileas settings reference. Two of them deserve a warning: span.disambiguation.vector.size and span.disambiguation.hash.algorithm.
A word’s position in the vector is its hash modulo the vector size. Change either the size or the algorithm and every word maps to a different position, so everything learned so far no longer lines up with new input. The file-based store records both values and discards a saved file that does not match, which turns a configuration change into a cold start rather than silently wrong answers. A shared store needs every instance to use the same values.
Pick these once, before you accumulate vectors you care about, and leave them alone. The defaults are fine for most deployments.
Turning it on, and turning it off per policy
Disambiguation is a deployment-wide switch: set span.disambiguation.enabled=true in the Phileas configuration. It is not something a redaction policy can turn on.
A policy can turn it off, though. If a policy’s filters never collide, opting out saves the per-document cost:
{
"config": {
"analysis": {
"spanDisambiguation": false
}
}
}
A policy that leaves this out gets disambiguation whenever the deployment has it enabled. See Span Disambiguation in the policy documentation for details.
The vector store holds data derived from your documents
The vectors are hashed, not raw text, but they are built from the words that appeared next to sensitive values in your documents, grouped by context name and type. Treat a persisted vector file, or a shared cache holding the vectors, as data derived from sensitive documents: restrict access to it, keep it inside the same boundary as the documents, and include it in your retention and deletion procedures.
Should you enable it?
Enable it when:
- Your policy has filters that can claim identical text, such as a custom identifier pattern that overlaps a built-in filter.
- Your process is long-lived, or you persist the vector store.
- Requests are grouped into contexts that mean something, so each context learns consistent associations.
Leave it off, or opt out per policy, when:
- Your filters never claim the same characters.
- Phileas runs in short-lived processes with an in-memory store.
- You run several instances and have no shared store yet.
Disambiguation improves the odds that an ambiguous value gets the right type. It does not guarantee it, and it does nothing for values no filter detected in the first place. Detection is probabilistic, so validate the output against samples of your own data, especially when you change filters or start a new context.
Phileas and Philter are part of Philterd’s open source PII redaction software toolkit.