Skip to main content
View source

Named Entity Recognition

View as Markdown

A RocketRide text-processing node that identifies named entities with a configured Hugging Face model. Choose it when documents need entity metadata for downstream filtering or analysis, rather than LLM-generated structured records.

About Hugging Face

Hugging Face provides tools and models for machine-learning workflows. This node uses its model identifiers with RocketRide's Transformers pipeline to perform named-entity recognition.

What it does

The node runs named-entity recognition on incoming text and documents, filtering the model's results by the configured confidence threshold. Text continues unchanged on the text lane. Documents are copied and, by default, enriched with entity lists and a total count in their metadata. Use it instead of dictionary when you need model-recognized entity spans and categories, not LLM-authored definitions of internal terminology.

Lanes

Lane inLane outDescription
texttextRuns recognition and passes the original text through unchanged.
documentsdocumentsRuns recognition on each document and writes an enriched copy downstream.

Profiles

Default: BERT Large (English) - High accuracy for English text (bertLarge).

ProfileModelContext
bertLarge (default)dbmdz/bert-large-cased-finetuned-conll03-englishEnglish profile; min_confidence defaults to 0.9.
bertBasedslim/bert-base-NEREnglish profile; min_confidence defaults to 0.9.
distilbertDavlan/distilbert-base-multilingual-cased-ner-hrlMultilingual profile; min_confidence defaults to 0.9.
xlmRobertaDavlan/xlm-roberta-base-ner-hrlMultilingual profile; min_confidence defaults to 0.9.
debertadslim/distilbert-NEREnglish profile; min_confidence defaults to 0.9.
biomedicaldmis-lab/biobert-base-cased-v1.1Biomedical profile; min_confidence defaults to 0.85.
custom(your own token-classification model)min_confidence defaults to 0.9.

Configuration

Start with the profile that matches the language or domain of the input, then adjust confidence and output handling only when the default behavior does not fit the pipeline. Select custom when a compatible model identifier is required; the other profiles preconfigure their model value.

Model

The profile selector defaults to bertLarge. Its preset supplies the model identifier, aggregation strategy, and confidence threshold. Use a named profile when it matches the input domain, or select custom to enter a model name yourself; the recognizer otherwise falls back to dbmdz/bert-large-cased-finetuned-conll03-english when no model value reaches it. A custom model must work with the NER pipeline and the selected aggregation strategy.

Entity aggregation strategy

This controls how the underlying pipeline combines word pieces into entities. The default is simple; the available strategies are none, simple, first, average, and max. Change it when your model's subword output produces entity boundaries or scores that are not useful to downstream consumers. Because confidence filtering happens after recognition, a different aggregation strategy can also change which combined entities meet the threshold.

Minimum confidence threshold

The threshold is a number from 0.0 to 1.0 and defaults to 0.9 for most presets (0.85 for biomedical). Raise it when metadata should contain only more confident entities; lower it when the model is missing useful candidates and the pipeline can tolerate more noise. The recognizer discards results below this value before it formats the entity dictionaries or stores them in document metadata.

Store entities in document metadata

This is on by default and affects the documents lane. When enabled, the node copies each document and stores deduplicated, sorted entity words under entities_<type> keys plus entities_count. Turn it off when the lane should preserve document metadata unchanged; recognition still runs, but the extracted entity list is not written to that document's metadata.

These values are written onto the document's DocMetadata as attributes, so in-process consumers read them as doc.metadata.entities_per rather than by subscript. They are extra fields on the model, so they still appear in the serialized metadata returned by toDict().

A document that arrives without metadata is given a DocMetadata built from the object being processed, inheriting its objectId, nodeId, parent, permissionId, and signature rather than a placeholder identity.

Requirements

This node declares GPU capability. It initializes its recognizer once at pipeline start and uses RocketRide's Transformers pipeline, which uses the model server when it is available and otherwise falls back to local execution. The model is not loaded while the node is opened in configuration mode.

Notes

Result handling

Each recognized item is formatted with entity_group, word, score, start, and end. Empty or whitespace-only text yields no entities. If recognition raises an exception, the node reports the error and returns an empty list, while text and documents continue through their normal lane handling.

Text and document behavior

On the text lane, the node collects recognized entities in instance state but writes only the original text downstream. On the documents lane, it processes each document's page_content separately and uses model_copy(deep=True) before changing metadata, so neither the original document nor the metadata it carries is mutated.

Upstream docs

Schema

FieldTypeDescriptionDefault
ner.aggregation_strategystringEntity aggregation strategy
How to combine word pieces into entities
"simple"
ner.min_confidencenumberMinimum confidence threshold
Minimum confidence score (0.0-1.0) for entity detection
0.9
ner.modelstringModel name
HuggingFace model to use for NER
ner.profilestringModel
NER model configuration
"bertLarge"
ner.store_in_metadatabooleanStore entities in document metadata
Add extracted entities to document metadata fields
true