Description
Embodied navigation — the ability of autonomous agents to move purposefully through complex, unstructured environments — sits at the intersection of perception, cognition, and action. Despite significant progress driven by large-scale pretraining, state-of-the-art navigation systems still struggle with tasks that humans find straightforward: interpreting multi-step spatial instructions, reasoning about occluded regions, generalizing geometric priors across scene types, and composing spatial relations hierarchically. These gaps point to a fundamental challenge: current systems lack robust spatial intelligence — the capacity to form, manipulate, and exploit structured internal representations of the world for goal-directed locomotion.
SISREN brings together researchers from robot learning, computer vision, cognitive science, and natural language processing to examine where spatial intelligence breaks down in embodied agents and how structured reasoning — for instance over maps, scene graphs, affordances, and language-grounded representations — can address these failures. The event is aimed at researchers and graduate students working on embodied navigation, scene understanding, and representation learning, and is especially relevant to the CoRL community's focus on deployable, generalizable robotic systems.
Core Challenges / Research Questions
SISREN will focus on the following concrete, open research challenges:
- C1. Spatial representation learning. What inductive biases, architectures, or training objectives lead to representations that support downstream spatial reasoning, rather than merely fitting perception benchmarks?
- C2. Compositional and long-horizon spatial reasoning. How can agents decompose complex spatial instructions (“go past the chair, turn left at the hallway, stop in front of the blue door”) into executable sub-goals, and how should the knowledge of a sub-task’s failure propagate to subsequent tasks?
- C3. Generalizable cognitive maps. Can learned map representations transfer across environments with different geometry, semantic content, and sensor modalities? What is the right abstraction level (metric, topological, or semantic) for robust generalization?
- C4. Grounding language to spatial reasoning. How should language models interact with spatial representations? When is tight coupling beneficial; when does it introduce brittleness?
- C5. Uncertainty-aware navigation. How should agents represent and communicate uncertainty on spatial states, and how does structured reasoning improve decision-making under partial observability?
- C6. Evaluation methodology. Current benchmarks reward pattern matching rather than genuine spatial understanding. What evaluation protocols better distinguish structured reasoning from shortcut learning?
Invited Speakers
All speakers confirmed.
UCLA
University of Maryland
MIT
Georgia Institute of Technology
Schedule
Half-day workshop · November 9th, 2026 · Austin, TX
| Time | Activity |
|---|---|
| TBD + 0:00 | Opening Remarks (15 min) — Workshop vision and overview of challenges C1–C6 |
| + 0:15 |
Invited Talk Block I (60 min · 25 min talk + 5 min Q&A each) · Dinesh Manocha — C5 (Uncertainty-Aware Navigation), C6 (Evaluation of Spatial Intelligence) · Bolei Zhou — C1 (Spatial Representation Learning), C3 (Cognitive Maps), C4 (Grounding Language) |
| + 1:15 | Contributed Spotlight Session (30 min) — 5-minute spotlights from top-reviewed contributed papers engaging challenges C1–C6 |
| + 1:45 | Breakout Discussion (30 min) — Participant self-select into challenge groups: A: C1 · B: C2–C3 · C: C4 · D: C5–C6. Each group produces a short summary for the panel. |
| + 2:15 |
Invited Talk Block II (60 min · 25 min talk + 5 min Q&A each) · Lu Gan — C1 (Spatial Representation Learning), C3 (Generalizable Cognitive Maps) · Jason Liu — C2 (Compositional Spatial Reasoning), C4 (Grounding Language) |
| + 3:15 | Panel Discussion & Open Q&A (30 min) — Moderated discussion with all invited speakers; synthesis of breakout findings across C1–C6 |
| + 3:45 | Poster Session & Closing (30 min) — Poster presentations, live demos, networking, and identification of key open problems |
Call for Papers
We invite 2–4 page extended abstract submissions (plus references) on topics mapped to the six challenge areas above. Submissions are non-archival to allow dual submission with venue proceedings.
Topics of interest include, but are not limited to:
- Spatial representation learning: inductive biases, architectures, and training objectives for structured spatial understanding
- Compositional and long-horizon reasoning: decomposition of complex instructions, sub-goal planning, failure propagation
- Generalizable cognitive maps: metric, topological, and semantic map representations; cross-environment transfer
- Language grounding: interaction between language models and spatial representations; tight vs. loose coupling tradeoffs
- Uncertainty-aware navigation: uncertainty representation, communication, and decision-making under partial observability
- Evaluation methodology: benchmarks and protocols that distinguish structured reasoning from shortcut learning
Review process. Each submission will receive at least two double-blind reviews from a program committee spanning robotics, computer vision, NLP, and cognitive science. At least three spotlight slots are reserved for junior researchers (graduate students or postdocs as first authors).
Submission link: OpenReview portal coming soon
Organizers
Submit
Paper submissions will be handled through OpenReview. The submission portal will be posted here once available. If you have any questions, do not hesitate to reach out to npbhatt@utexas.edu.
SISREN · Spatial Intelligence and Structured Reasoning for Embodied Navigation · CoRL 2026 · Austin, TX