IROS 2026 · Pittsburgh

IMPRINT: Image-Conditioned Query Enrichment for Long-Tail Object Goal Navigation

Jelin Raphael Akkara1,2, Filippo Ziliotto1,2, Luciano Serafini2, Lamberto Ballan1, Tommaso Campari2

1University of Padova  ·  2Fondazione Bruno Kessler (FBK)

Comparison of text-only grounding with IMPRINT image-conditioned query enrichment

Motivation

Text alone is often not enough.

Semantic-map-based ObjectNav relies on text queries, which become unreliable for rare or semantically specific targets.

IMPRINT enriches a target query with web-retrieved images, providing richer visual cues without retraining or modifying the navigation policy.

Method

A zero-shot, plug-and-play upgrade for queryable semantic maps

IMPRINT enriches text queries with web-retrieved images to improve target grounding in existing navigation systems, without retraining the base policy.

The baseline pipeline comprises three stages: mapping projects visual features onto a 2D grid; map querying compares a text-query embedding with the stored features to produce a similarity map; and target localization selects the highest-scoring grid cell.

IMPRINT modifies only the map-querying stage. It retrieves visual examples of the target, encodes them with the baseline vision–language model, and aggregates their similarity maps to guide target localization. The four steps below describe this process.

IMPRINT mapping and image-conditioned map-querying pipeline
01

Retrieve

Retrieve N images from the web using the target text query.

02

Encode

Extract image embeddings using the baseline vision–language encoder.

03

Match

Compute cosine similarity between each image embedding and the map features to produce N similarity maps.

04

Aggregate

Combine the N similarity maps and select the highest-scoring grid cell as the predicted target location.

Zero-shot and model-agnostic: IMPRINT requires no additional training and leaves the underlying navigation policy unchanged.

Evaluation

Separating grounding from navigation

We evaluate IMPRINT in two complementary settings: static phase measures target localization on a pre-built map, while online phase assesses performance within the complete navigation pipeline. Together, these settings distinguish improvements in spatial grounding from their impact on navigation success.

Static phase

Isolated grounding

We query a pre-built semantic map and compare the highest-scoring grid cell with the ground-truth target location, isolating grounding performance from navigation and detection.

Pre-built map→Query→Grid location

Online phase

End-to-end ObjectNav

We evaluate the complete ObjectNav pipeline, with mapping, querying, detection, and navigation operating jointly throughout each episode to measure how grounding improvements translate into navigation success.

Update map→Query→Navigate

New benchmark

HSSD-rare

A fine-grained ObjectNav benchmark built on HSSD-Hab that tests subcategory-level grounding within shared parent categories.

1,000
Episodes
17
HSSD scenes
20
Parent categories
559
Subcategories
992
Object instances
6.55 m
Mean start–goal distance
Examples of fine-grained chair, lamp, and drinkware categories in HSSD-rare
(a) Fine-grained categories. Examples of subcategories within shared parent categories in HSSD-rare.

Why Viewpoints? HSSD provides object positions and dimensions but no category-specific viewpoints. Valid viewpoints are therefore required to generate ObjectNav episodes where the target is visible and reachable.

Viewpoint Generation: Accessible boundary points are identified around each target. Candidate viewpoints are then sampled radially in nearby traversable space and oriented toward the object.

Viewpoint Validation: The target’s 3D bounds are projected into each candidate view and checked against the depth image. Candidates are rejected if all sampled target points are occluded; otherwise, they are retained as valid viewpoints.

Four map panels illustrating candidate viewpoint generation and filtering around a target object
(b) Viewpoint generation pipeline. Generation and filtering of candidate viewpoints around a target object.
Color and depth views showing visible and occluded target points for viewpoint validation
(c) Viewpoint validation. Color and depth views illustrate target visibility checks at candidate viewpoints.

Results

Better grounding, with detection as the bottleneck

01

Static grounding

Image-conditioned queries improve spatial grounding across both datasets and all evaluated encoders. Gains are especially pronounced for fine-grained HSSD-rare categories.

Static grounding success rate for text, image, and combined queries across BLIP2, SigLIP, and SED

02

Online navigation

Grounding gains transfer to navigation, but target detection limits their full impact. Enhanced detection unlocks substantially stronger performance.

Online ObjectNav success rate comparing Standard, IMPRINT, and IMPRINT with enhanced detection

Reference

Citation

If you find IMPRINT or HSSD-rare useful, please cite our work.

@article{akkara2026imprint,
  title   = {IMPRINT: Image-Conditioned Query Enrichment
             for Long-Tail Object Goal Navigation},
  author  = {Akkara, Jelin Raphael and Ziliotto, Filippo and
             Serafini, Luciano and Ballan, Lamberto and Campari, Tommaso},
  journal = {arXiv preprint arXiv:2607.25106},
  year    = {2026}
}