Skip to main content
View source

Anonymize

View as Markdown

A RocketRide text filter that detects configured entity types and replaces their spans before text continues downstream. Pick it to redact PII-like content while preserving the rest of the document, rather than dropping the document or relying on a fixed pattern matcher alone.

About GLiNER

GLiNER is the named-entity recognizer used by this node through RocketRide's GLiNER model adapter. The recognizer receives a list of labels, returns entity spans, and is used here to replace those spans in pipeline text. This node can also combine those detections with spans and labels supplied by an upstream classification stage.

What it does

The node accumulates text for an input object, detects entities when that object closes, and writes the redacted text to the text lane. Without classification input it detects the configured entity types; with classification input it also uses its reported character spans and associated rule labels. Choose it over an exact-match redaction step when the data needs entity detection over configurable labels or integration with classifier results.

In mask mode each detected span is replaced with the configured character for the original span length. Token mode replaces it with a label such as [EMAIL]; a classifier-only match becomes [REDACTED] unless an overlapping GLiNER match supplies a more specific label.

Lanes

Lane inLane outDescription
texttextCollect text for an object, redact detected spans at closing, and emit it.

Profiles

Default: GLiNER Small - Lightweight general-purpose model (glinerSmall).

ProfileModel
glinerMultiPIIurchade/gliner_multi_pii-v1
glinerPIILargeknowledgator/gliner-pii-large-v1.0
glinerMergedLargexomad/gliner-model-merge-large-v1.0
glinerSmall (default)urchade/gliner_small-v2.1
glinerMediumurchade/gliner_medium-v2.1
glinerLargeurchade/gliner_large-v2.1
glinerMultiurchade/gliner_multi
gretelSmallgretelai/gretel-gliner-bi-small-v1.0
gretelLargegretelai/gretel-gliner-bi-large-v1.0
glinerKotaeminlee/gliner_ko
glinerItDeepMount00/GLiNER_PII_ITA
glinerArNAMAA-Space/gliner_arabic-v2.1
glinerCommunitySmallgliner-community/gliner_small-v2.5
glinerCommunityMediumgliner-community/gliner_medium-v2.5
glinerCommunityLargegliner-community/gliner_large-v2.5
glinerBiomedSmallIhor/gliner-biomed-small-v1.0
glinerBiomedLargeIhor/gliner-biomed-large-v1.0
custom(your own GLiNER-compatible model)

Configuration

Choose a profile when one of the supplied model choices suits the text you process, or choose custom and supply a model name. Most pipelines then only need to tailor the entity labels and decide whether retaining entity labels in the output is useful. The profile selected when adding the node is glinerSmall; the configuration field itself defaults to glinerMergedLarge.

Model

Profiles supply a model name, while the custom profile exposes Model name for a name at least two characters long. Change models when the supplied profile's intended language, domain, or resource trade-off is a closer match for the data. The recognizer loads the configured model through the RocketRide GLiNER adapter, which uses a model server when one is configured and otherwise runs locally.

Entity types to anonymize

Entity types to anonymize is a list of zero-shot labels. Its default has 15 common PII-style labels, including person, email, phone number, and credit card number. Replace or extend this list for the kinds of entities your pipeline must hide. Blank and non-string list items are ignored; a non-list or an empty resulting list falls back to the default labels, so an attempt to clear the list does not silently disable detection.

Redaction style and character

Use mask (the default) when it is important to preserve the original text length: each matched character is overwritten by Character to use for anonymization, which defaults to . Use token when downstream processing needs to know what was removed; it substitutes label tokens instead and does not use the masking character. Adjacent mask spans merge into one replacement; token spans remain separate unless they actually overlap.

Requirements

The node declares GPU capability and installs GLiNER plus an ONNX Runtime. Outside macOS its requirements select onnxruntime-gpu; on macOS they select the non-GPU onnxruntime package. Model loading is therefore a runtime cost to plan for, especially before selecting a larger profile.

Notes

Long text and prediction failures

Text is processed in 1,024-character chunks with 128-character overlap, label lists are sent in batches of 32, and up to four chunks run concurrently. Duplicate entities from overlapping chunks are removed. A model error for one label batch or chunk is logged and processing continues, so review redaction results when complete coverage is required.

Classifier input

When classification data arrives, the node uses three independent inputs. classificationPolicy idRef values become GLiNER labels only when the local Nucleuz rulePack.dat maps them; if the rule pack is unavailable, those labels are omitted. <Term> values from classificationRules remain GLiNER labels without the rule pack. Classifier-reported textMatches are redacted directly, independent of both label sources. If neither source produces labels, GLiNER adds no classification-derived detections, but the reported text matches are still redacted.

Upstream docs

Schema

FieldTypeDescriptionDefault
anonymize.modelstringModel name
Gliner model to use for anonymization
anonymize.profilestringModel
Anonymize model
"glinerMergedLarge"
anonymizeCharstringCharacter to use for anonymization
Character
entityTypesarrayEntity types to anonymize
PII / entity types to detect and mask. Pre-filled with common types; remove any you don't want, or add your own (the model is zero-shot, so any label works).
["person","name","email","phone number","address","social security number","credit card number","date of birth","organization","company","location","ip address","bank account","passport number","driver license"]
redactionStylestringRedaction style
How detected entities are replaced. 'mask' overwrites each entity with the anonymization character (████). 'token' replaces each entity with a labelled tag like [PERSON] or [EMAIL].
"mask"

Dependencies

  • gliner
  • onnxruntime-gpu ==1.22.0; platform_system != 'Darwin'
  • onnxruntime ==1.22.0; platform_system == 'Darwin'