Chi-Pin Huang

Research Scientist

profile.jpg

f11942097[at]ntu.edu.tw

I am a Research Scientist at NVIDIA Research, working on Embodied AI.

My research focuses on reasoning-enhanced vision-language-action models that enable robots to reason about tasks and environments and translate that reasoning into effective actions in the physical world.

Before joining NVIDIA as a Research Scientist, I received my Ph.D. from National Taiwan University (NTU) in January 2026, advised by Prof. Yu-Chiang Frank Wang. I also earned my bachelor’s degree in Computer Science and Information Engineering (CSIE) from NTU in 2022 and worked as an Applied Scientist Intern at Microsoft, developing deep learning models for Bing Maps.

News

Feb 20, 2026 Our paper Fast-ThinkAct is accepted by CVPR 2026.
Jan 09, 2026 Received my Ph.D. from National Taiwan University (NTU) and will be joining NVIDIA Research as a Research Scientist.
Dec 27, 2025 Our papers “SANTA” and “TA-Prompting” are accepted by WACV 2026.
Sep 18, 2025 Our paper “ThinkAct” is accepted by NeurIPS 2025.
Jun 26, 2025 Our paper “CNS” is accepted at ICCV 2025, and “MotionMatcher” is accepted at the ICCV 2025 Workshop on P13N: Personalization in Generative AI.
Feb 27, 2025 Our paper “VideoMage” is accepted by CVPR 2025.
Feb 03, 2025 Join NVIDIA Research as a Research Intern.
Jul 02, 2024 Our papers “Receler” and “Select and Distill” are accepted at ECCV 2024.
Jan 16, 2024 Our paper “RAPPER” is accepted by ICLR 2024.

Selected Publications

  1. arXiv 2608
    physcap-framework.png
    PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration

    Guides robots to actively probe hidden physical properties, such as mass and stiffness, for manipulation when vision alone is insufficient.

    Chen-Yu Lin, Jing-Wen Chen , Hsueh-En Chang , Hung-An Chen , Sheng-Hsun Chang, Chi-Pin HuangFu-En YangMin-Hung ChenYi-Ting ChenYu-Chiang Frank Wang, and Shao-Hua Sun
    arXiv preprint arXiv:2608.21031, 2026
  2. CVPR 2026
    fast-thinkact-teaser.png
    Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning

    Compresses textual reasoning into compact, verbalizable latent plans for faster vision-language-action inference.

    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  3. NeurIPS 2025
    thinkact-teaser.png
    ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning

    Connects embodied reasoning to robot actions through reinforced visual latent planning, supporting long-horizon tasks and self-correction.

    Advances in Neural Information Processing systems (NeurIPS), 2025
  4. CVPR 2025
    videomage-teaser.png
    VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion Models

    Combines multiple subject identities and reference motions to generate customized videos with consistent appearances and interactions.

    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
  5. ECCV 2024
    receler-teaser.png
    Receler: Reliable concept erasing of text-to-image diffusion models via lightweight erasers

    Uses lightweight erasers to suppress target concepts in diffusion models while preserving unrelated concepts and resisting adversarial prompts.

    In European Conference on Computer Vision (ECCV), 2024