# depth_estimate

A RocketRide image node that estimates a dense depth map from one image and emits both a colorized map and summary statistics. Choose it when you need per-pixel depth information instead of object bounding boxes or an image caption.

## About Depth Anything V2

Depth Anything V2 is the monocular depth-estimation model used by this node. The node loads the configured model identifier and optional revision, using the supplied V2 Small profile by default.

## What it does

The node buffers each input image, estimates depth, and restores the dense result to the input's original dimensions. It can send a JPEG colorized depth map on the `image` lane and the depth array's minimum, maximum, and mean as JSON on the `text` lane. Choose it for a frame-wide depth signal; pair it with Object Detection when you need boxes as well as rough distance context for detected objects.

## Lanes

| Lane in | Lane out | Description |
| --- | --- | --- |
| `image` | `image` | JPEG colorized depth map, where red is near and blue is far. |
| `image` | `text` | JSON depth statistics: `min`, `max`, and `mean`. |

## Configuration

The supplied V2 Small profile is the default model, leaving the maximum input edge as the main performance and detail control. The implementation uses its default model if no model is configured.

### Max input edge (px)

This setting limits the input image's long edge before depth inference; its default is `1024`. Lower it when local inference needs less work and memory, accepting less spatial detail in the predicted depth; raise it when sharper depth boundaries matter. The node upsamples the dense result to the original dimensions after inference. At runtime, malformed values use the default and all values are clamped to the range `256`–`4096`.

## Requirements

With a model server configured, inference runs there. Otherwise it runs locally on CPU, Apple Silicon (MPS), or CUDA, and a device lock serializes local model use. Output values are relative depth, not a calibrated physical distance.

## Notes

### Output and failures

The colorized map is encoded as JPEG after the depth result is restored to the original image size. The emitted statistics are calculated from that restored depth array. If image decoding or depth inference fails, the node logs a warning, drops the frame, and does not emit image or text output for it.

## Upstream docs

- [Depth Anything V2 Small model page](https://huggingface.co/depth-anything/Depth-Anything-V2-Small-hf)

<!-- ROCKETRIDE:GENERATED:PARAMS START -->
<!-- Generated by nodes:docs-generate. Do not edit by hand. -->

## Schema

| Field | Type | Description | Default |
|---|---|---|---|
| `depth_estimate.maxEdge` | `number` | **Max input edge (px)**<br/>Downscale input so the long edge <= this value before inference; dense output is upsampled back to original. Lower = faster + less VRAM, higher = sharper depth. | `1024` |
| `depth_estimate.profile` | `string` | **Model** | `"v2-small"` |

## Source

[<svg viewBox="0 0 16 16" width="15" height="15" fill="currentColor" aria-hidden="true" style="vertical-align:-0.15em;margin-right:0.35em"><path d="M8 0C3.58 0 0 3.58 0 8c0 3.54 2.29 6.53 5.47 7.59.4.07.55-.17.55-.38 0-.19-.01-.82-.01-1.49-2.01.37-2.53-.49-2.69-.94-.09-.23-.48-.94-.82-1.13-.28-.15-.68-.52-.01-.53.63-.01 1.08.58 1.23.82.72 1.21 1.87.87 2.33.66.07-.52.28-.87.51-1.07-1.78-.2-3.64-.89-3.64-3.95 0-.87.31-1.59.82-2.15-.08-.2-.36-1.02.08-2.12 0 0 .67-.21 2.2.82.64-.18 1.32-.27 2-.27.68 0 1.36.09 2 .27 1.53-1.04 2.2-.82 2.2-.82.44 1.1.16 1.92.08 2.12.51.56.82 1.27.82 2.15 0 3.07-1.87 3.75-3.65 3.95.29.25.54.73.54 1.48 0 1.07-.01 1.93-.01 2.2 0 .21.15.46.55.38A8.013 8.013 0 0016 8c0-4.42-3.58-8-8-8z"/></svg> View source](https://github.com/rocketride-org/rocketride-server/tree/develop/nodes/src/nodes/depth_estimate)
<!-- ROCKETRIDE:GENERATED:PARAMS END -->
