Skip to main content
View source

Weaviate

View as Markdown

A RocketRide vector-store node that stores embedded document chunks in Weaviate and retrieves them by semantic or keyword search. Use it when a pipeline or agent needs a Weaviate-backed document store.

About Weaviate

Weaviate is a vector database for storing data with embeddings and structured properties. It supports vector search and filtered retrieval over collections. Use it when a team has selected Weaviate for its retrieval data and wants a RocketRide pipeline or agent to use the same collection.

What it does

The node accepts embedded documents on its documents lane and can retrieve matching documents for incoming questions. It also exposes the store to an agent as tools. A collection is created when documents are first added, and incoming chunks without embeddings are rejected. Pick it over the sibling vector stores when Weaviate is the database the workload already operates, including when the node should use Weaviate's local REST and gRPC connection or its cloud connection.

The driver creates the collection without a vectorizer and gives it an HNSW vector index. It stores document content and metadata as collection properties, allowing semantic or keyword retrieval with the same metadata filters.

Lanes

Lane inLane outDescription
documentsStore embedded document chunks.
questionsdocumentsReturn matching documents.
questionsanswersReturn matching documents as answers.
questionsquestionsEnrich questions with matching documents.

As a tool

The configured tool-server name defaults to weaviate.

FunctionDescription
searchSearches the store for a non-empty query; accepts optional top_k and metadata filter, and returns matching content, metadata, and scores.
upsertAdds or updates a non-empty documents array. Each document requires content and an object ID; it can provide an embedding and embedding model or use the bound embedding provider.
deleteDeletes documents for a non-empty object_ids array and returns the deleted count.

search requires a bound embedding provider for semantic similarity search. The three functions return a failure object when their required input or an embedding cannot be obtained.

search accepts an optional filter object honoring only objectId, nodeId and parent; any other key is rejected. upsert accepts an optional metadata object storing nodeId, parent and chunkId, defaulting to "vectordb_tool", "/" and 0 respectively.

Profiles

Default: Your own Weaviate server (local).

ProfileDefault hostDefault port
Your own Weaviate server (default)localhost8080
Weaviate cloud serverEmpty443

Configuration

Choose the cloud or local profile, then configure the host and collection. The cloud profile connects to the configured cluster URL; the local profile also uses the configured gRPC port, which defaults to 50051. The profile changes the connection method, so decide which service you are connecting to before adjusting search settings.

Host, ports, and API key

The runtime trims whitespace, strips a leading http:// or https://, and removes trailing slashes from the entered host. Cloud mode calls the cloud connection with the API key. Local mode passes the configured REST and gRPC ports, and uses the API key only when one is supplied. Use the local profile for a service that exposes both local endpoints; a startup or validation failure can mean the selected profile does not match the service, or that the local gRPC port is unavailable.

Collection

The collection is the destination for stored chunks and is created with no vectorizer and an HNSW vector index when first needed. The save-time probe requires a name that begins with an uppercase letter and then contains only letters, digits, or underscores. Change it to isolate a corpus or embedding space; it directs the node to a different collection rather than creating a logical partition. A name that begins lowercase or contains punctuation stops configuration before the pipeline starts.

Similarity and retrieval score

The runtime accepts cosine, dot, l2-squared, hamming, or manhattan as the similarity setting, defaulting to cosine; another value stops initialization. It passes the selected value to the HNSW index when it creates a collection, so make the choice match the embedding space before the first ingest. Changing the setting does not recreate an existing collection's index.

The retrieval score defaults to 0.5, while the conversion code also drops results below a hard 0.20 before returning them. Raise the configured threshold if weak documents crowd the context; lower it when expected material is omitted, remembering that nothing below 0.20 can be returned. Cosine scores are converted as (distance + 1) / 2; the other supported metrics use a sigmoid conversion, so threshold values are not directly comparable after a metric change.

Connection timeouts

The client uses separate initialization, query, and insert timeouts of 30, 60, and 120 seconds. They are fixed by this node's implementation, so a slow connection cannot be tuned from this panel. If only large writes time out while queries work, reduce the ingest workload or investigate the service path; if all operations time out, start with the selected profile, host, ports, and API key instead of treating it as a retrieval-score problem.

Tool Server Name

The tool-server name defaults to weaviate and prefixes the three agent functions. Change it when multiple Weaviate stores are connected to the same agent so their functions do not share a namespace.

Authentication

For a cloud profile, provide the API key; the runtime uses it for the cloud connection. For a local profile, the key is optional: an empty key connects without credentials, while a supplied key is used for local authentication.

Notes

Search and document lifecycle

Semantic search requires an embedding and does not accept a non-zero offset. Keyword search matches content with the metadata filters. Re-ingesting an object deletes its existing chunks before replacement chunks are added. The store can mark chunks deleted or active, and its default filters exclude marked-deleted chunks.

Before semantic search, the driver checks that the collection is compatible with the query's embedding model. The collection's batch import reports an error if any objects failed to import, instead of silently treating a partial write as a successful ingest.

Filters cover node, parent, permissions, object IDs, chunks, table fields, and the deletion flag. Normal queries add isDeleted = false, so marked-deleted content is hidden unless a caller requests it. This lets a shared collection be scoped at query time without using collection names as routine access filters.

Rendering

Rendering retrieves an object's chunks in chunkId order and sends joined text to the callback in renderChunkSize groups. The node's document count is a count of stored vectors (chunks).

When an object is re-ingested, the driver deletes its existing objects before adding replacements through Weaviate's dynamic batch API. A batch with failed objects raises an error after the batch ends. Keep the original source available for a retry, since the replacement behavior is not an all-or-nothing transaction across the prior objects and the new batch.

Upstream docs

Schema

FieldTypeDescriptionDefault
vector.cloud.hostEnter the server IP address e.g. .weaviate.cloud
vector.cloud.port443
vector.local.grpc_port50051
vector.local.host"localhost"
vector.local.port8080
weaviate.profilestringType of Weaviate host
Connect to...
"local"
weaviate.providerstringconst: "weaviate"
weaviate.serverNamestringTool Server Name
Namespace for agent-facing tool names, e.g. 'weaviate' exposes tools as weaviate.search / weaviate.upsert / weaviate.delete. Change this when running multiple Weaviate nodes in the same pipeline so their tool names do not collide.
"weaviate"

Dependencies

  • authlib
  • grpcio
  • grpcio-health-checking
  • grpcio-tools
  • httpx
  • pydantic
  • requests
  • validators
  • weaviate-client >=4.20.0
  • numpy