Chi-Pin Huang
Research Scientist
f11942097[at]ntu.edu.tw
I am a Research Scientist at NVIDIA Research, working on Embodied AI.
My research focuses on reasoning-enhanced vision-language-action models that enable robots to reason about tasks and environments and translate that reasoning into effective actions in the physical world.
Before joining NVIDIA as a Research Scientist, I received my Ph.D. from National Taiwan University (NTU) in January 2026, advised by Prof. Yu-Chiang Frank Wang. I also earned my bachelor’s degree in Computer Science and Information Engineering (CSIE) from NTU in 2022 and worked as an Applied Scientist Intern at Microsoft, developing deep learning models for Bing Maps.
News
| Feb 20, 2026 | Our paper Fast-ThinkAct is accepted by CVPR 2026. |
|---|---|
| Jan 09, 2026 | Received my Ph.D. from National Taiwan University (NTU) and will be joining NVIDIA Research as a Research Scientist. |
| Dec 27, 2025 | Our papers “SANTA” and “TA-Prompting” are accepted by WACV 2026. |
| Sep 18, 2025 | Our paper “ThinkAct” is accepted by NeurIPS 2025. |
| Jun 26, 2025 | Our paper “CNS” is accepted at ICCV 2025, and “MotionMatcher” is accepted at the ICCV 2025 Workshop on P13N: Personalization in Generative AI. |
| Feb 27, 2025 | Our paper “VideoMage” is accepted by CVPR 2025. |
| Feb 03, 2025 | Join NVIDIA Research as a Research Intern. |
| Jul 02, 2024 | Our papers “Receler” and “Select and Distill” are accepted at ECCV 2024. |
| Jan 16, 2024 | Our paper “RAPPER” is accepted by ICLR 2024. |
Selected Publications
- CVPR 2026
Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent PlanningCompresses textual reasoning into compact, verbalizable latent plans for faster vision-language-action inference.
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026 - NeurIPS 2025
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningConnects embodied reasoning to robot actions through reinforced visual latent planning, supporting long-horizon tasks and self-correction.
Advances in Neural Information Processing systems (NeurIPS), 2025 - CVPR 2025
VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion ModelsCombines multiple subject identities and reference motions to generate customized videos with consistent appearances and interactions.
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025 - ECCV 2024
Receler: Reliable concept erasing of text-to-image diffusion models via lightweight erasersUses lightweight erasers to suppress target concepts in diffusion models while preserving unrelated concepts and resisting adversarial prompts.
In European Conference on Computer Vision (ECCV), 2024