Embodied AI · Multimodal Reasoning · Robot Navigation
I build embodied AI systems that turn multimodal perception into grounded robot decisions through structured memory, retrieval, world models, and verified execution. I am with Beijing Institute of Technology.
- Knowledge-grounded intelligence — multimodal retrieval, semantic-spatial memory, and open-world ObjectNav
- Reliable autonomy — world-model-assisted control, safety checks, and multimodal terminal verification
- Reproducible robotics — inspectable simulation assets, scripted validation, and clear system boundaries
An embodied multimodal retrieval prototype for open-world object-goal navigation. It connects visual observations, spatial-region memory, semantic evidence, and vision-language model backends, with a path toward simulation-to-real deployment.
A map-independent charging workflow for Go2-W that combines world-model-generated action chunks, cross-embodiment motion adaptation, veto-only safety checks, and multimodal terminal verification.
A reproducible Isaac Sim scene with scripted setup, checksum verification, headless rendering and physics checks, and explicit boundaries for third-party assets.
Embodied AI · Object-goal navigation · Multimodal RAG · Vision-language models · World models · Robot learning · Isaac Sim
I care about clear system boundaries, reproducible setup, and evidence that others can inspect. I try to document both what a prototype demonstrates and what it does not.
- Beijing Institute of Technology
- limingyi@bit.edu.cn

