# caption

A RocketRide image node that turns each input image into a natural-language caption. Choose it when downstream text processing needs an image description rather than boxes, masks, or extracted text.

## About Florence-2

Florence-2 is the image captioning model used by this node. The node loads the configured model identifier and optional revision, and sends its configured caption task to the model for each image.

## What it does

The node buffers each incoming image stream, captions the completed image, and writes the resulting string on the `text` lane. It offers short, detailed, and more-detailed caption tasks through its configuration. Choose it for a general natural-language description of image content; reach for Object Detection when you need boxes and OCR when you need the text in the image.

## Lanes

| Lane in | Lane out | Description |
| --- | --- | --- |
| `image` | `text` | Caption string for the completed input image. |

## Configuration

The single supplied profile selects the Florence-2 Base model, so most use cases only need a granularity choice. The implementation uses its default model and task if either value is empty.

### Granularity

Granularity selects one of three configured tasks: `caption` (the default, shown as Short), `detailed_caption`, or `more_detailed_caption`. Keep the default when a concise description is sufficient; use a more detailed task when the next node needs richer text to reason over. The task is passed directly to the captioning facade, so choose one of the values exposed in the configuration panel.

## Requirements

With a model server configured, captioning runs there. Otherwise it runs locally on CPU, Apple Silicon (MPS), or CUDA, and a device lock serializes local caption generation.

## Notes

### Empty captions on failure

The node runs captioning only when the `text` lane has a listener. If image decoding or caption inference raises an exception, it logs a warning and writes an empty string to that listener instead of propagating a result from the failed image.

## Upstream docs

- [Florence-2 Base model page](https://huggingface.co/microsoft/Florence-2-base)

<!-- ROCKETRIDE:GENERATED:PARAMS START -->
<!-- Generated by nodes:docs-generate. Do not edit by hand. -->

## Schema

| Field | Type | Description | Default |
|---|---|---|---|
| `caption.profile` | `string` | **Model** | `"florence-base"` |
| `caption.task` | `string` | **Granularity**<br/>How detailed the caption should be. | `"caption"` |

## Source

[<svg viewBox="0 0 16 16" width="15" height="15" fill="currentColor" aria-hidden="true" style="vertical-align:-0.15em;margin-right:0.35em"><path d="M8 0C3.58 0 0 3.58 0 8c0 3.54 2.29 6.53 5.47 7.59.4.07.55-.17.55-.38 0-.19-.01-.82-.01-1.49-2.01.37-2.53-.49-2.69-.94-.09-.23-.48-.94-.82-1.13-.28-.15-.68-.52-.01-.53.63-.01 1.08.58 1.23.82.72 1.21 1.87.87 2.33.66.07-.52.28-.87.51-1.07-1.78-.2-3.64-.89-3.64-3.95 0-.87.31-1.59.82-2.15-.08-.2-.36-1.02.08-2.12 0 0 .67-.21 2.2.82.64-.18 1.32-.27 2-.27.68 0 1.36.09 2 .27 1.53-1.04 2.2-.82 2.2-.82.44 1.1.16 1.92.08 2.12.51.56.82 1.27.82 2.15 0 3.07-1.87 3.75-3.65 3.95.29.25.54.73.54 1.48 0 1.07-.01 1.93-.01 2.2 0 .21.15.46.55.38A8.013 8.013 0 0016 8c0-4.42-3.58-8-8-8z"/></svg> View source](https://github.com/rocketride-org/rocketride-server/tree/develop/nodes/src/nodes/caption)
<!-- ROCKETRIDE:GENERATED:PARAMS END -->
