Summarization: LLM
A RocketRide filter node that uses an LLM to distill incoming text into a summary, key points, and named entities.
What it does
Accumulates all text (and table content) of each object flowing through the pipeline, then on close splits the full document into chunks and asks the connected LLM to extract three things from each chunk: a concise summary, a list of key points, and the most significant named entities (people, organizations, products, events, dates, locations). The LLM is instructed to respond as JSON.
Chunking uses LangChain's RecursiveCharacterTextSplitter, sized to the connected LLM's context length and measured with the LLM's own token counter. The token budget of the instruction prompt itself is accounted for, so each chunk plus the prompt always fits the model's context window.
Only the first numberOfSummaries chunks (default 2) are summarized; the rest of the document is ignored. Each of the three extraction sections can be disabled individually by setting its config field to 0. Choose this node when one LLM pass should produce all three outputs; use a simpler text transformation when no LLM-generated summary is needed.
Output is emitted only on lanes that actually have a downstream listener: plain formatted text on the text lane, and/or one structured document per section on the documents lane.
Connections
| Connection | Required | Description |
|---|---|---|
llm | yes | LLM used for extraction and for its context length and token counter. |
Lanes
| Lane in | Lane out | Description |
|---|---|---|
text | text | Summarized output as plain text: a summary block, a Key Points: bullet list, and an Entities: bullet list. Sections that are disabled or empty are omitted. |
text | documents | Summarized output as structured documents: each summary, key-point list, and entity list becomes its own Doc with an incrementing chunkId in its metadata. |
The declared input lane is text; the node accumulates that text for the
current object before it invokes the LLM at close.
Configuration
The default configuration summarizes two chunks, targets 1,500 words per summary, 250 words per key point, and extracts 25 entities. Tune the limits to the length and density of the source material; the splitter uses the connected LLM's token counter and context length, not a fixed character limit.
Number of chunks to summarize
numberOfSummaries limits how many chunks are sent to the LLM after splitting.
Keep the default when the opening chunks are representative. Raise it to cover
longer documents, accepting more LLM calls; set it to 0 when no chunks should
be processed.
Summary, key-point, and entity limits
numberOfSummaryWords, numberOfKeyPointWords, and numberOfEntities control
the instructions passed to the LLM. Set a value to 0 to omit that section.
Use smaller limits for compact indexing output and larger limits when downstream
readers need more detail. These limits do not change how many source chunks are
selected; that is controlled by numberOfSummaries.
Schema
| Field | Type | Description | Default |
|---|---|---|---|
summarization.numberOfEntities | number | Number of entities to extract from the document. Set to 0 to disable entity extraction. | |
summarization.numberOfKeyPointWords | number | Number of words in each key point. Set to 0 to disable key points. | |
summarization.numberOfSummaries | number | Number of chunks to summarize after the document is split | |
summarization.numberOfSummaryWords | number | Number of words in each summary. Set to 0 to disable summaries. | |
summarization.profile | string | "default" |
Dependencies
langchain