COMPUTER VISION · WORLD MODELS
Hai Nguyen-Truong.
PhD student · Northeastern
I am a PhD student in Computer Science at Northeastern University, advised by Prof. Lorenzo Torresani at FarSight Lab. I study world models for planning and anticipation, and action understanding in egocentric and exocentric video.
Previously, I earned an MPhil in Computer Science and Engineering at HKUST with Prof. Sai-Kit Yeung, and a BSc in Computer Science with highest distinction at the University of Science, VNU-HCM.
My research experience spans physics-plausible video generation at the Google-funded Fulbright AI Institute, image generation at Huawei Hong Kong Research Center, and spatiotemporal segmentation at VinAI Research (now part of Qualcomm).
TrajPilot predicts possible future camera trajectories from first-person video and uses them to guide action prediction. Motion supplies a fine-grained signal of intent, helping the model plan action sequences and anticipate what someone will do next without observing their future path.
Video generators can make convincing frames but get the physics wrong. OMR uses a frozen world model to steer a frozen video generator during a single sampling run. A small adapter connects the generator’s intermediate latents to the world model, so it can guide the video before it is finished.
VATEX is a novel method for referring image segmentation that leverages vision-aware text features to improve text understanding. By decomposing language cues into object and context understanding, the model can better localize objects and interpret complex sentences, leading to significant performance gains.
SegTransVAE is the first work exploiting the hybrid architecture between CNN, Transformers with the Variational Autoencoder (VAE) branch to the network to reconstruct the input images jointly with segmentation.
Latest news.
TrajPilot accepted at NeurIPS 2026. Our work uses predicted camera trajectories to anticipate actions from first-person video.
Started my PhD at Northeastern University.
OMR accepted at BMVC 2026; concluded my research assistantship at the Google-funded Fulbright AI Institute.
Second prize at VAIC (nearly 1,500 contestants, 340 teams) for Tangent · $5,000 cashOne week after the Mobility win, our AI-for-education project turned lectures into vivid, interactive video lessons.
First prize in Mobility at GenAI Fund’s Agentic AI Build Week (2,000 builders) · $15,000 OpenAI credits + $250,000 Microsoft startup fundThe enterprise track focused on context-aware AI for digital mobility and location intelligence.
TransCues accepted at WACV 2026.
Where I’ve worked.
Research Assistant in Computer Vision
- Researched video diffusion models, focusing on the physical plausibility of generated video.
- Hosted weekly group meetings and research seminars for our computer vision group.
- Mentored junior students on computer vision capstone and thesis projects.
AI Research Engineer
- Contributed to the MindONE project, focusing on multimodal/multitask generation models and optimizing them for Ascend NPUs.
Research Resident
- Researched spatiotemporal tasks: video instance/panoptic segmentation and 4D point cloud panoptic segmentation.
Honors & awards.
Academic
- 2025 Best Poster Presentation Award · AVSTC
- 2024 UGC Research Travel Grant · HKUST
- 2022–2024 Postgraduate Scholarship (PGS) · HKUST
- 2022 Merit Award for Highest Distinction
- 2022 Third Prize · EURÉKA
Competitions & Math
- 2026 🥇 First Prize, Mobility · GenAI Fund Agentic AI Build Week
- 2026 🥈 Second Prize · Vietnam AI Innovation Challenge (Tangent)
- 2023 & 2024 🥈 Runner-up · Maritime CV Workshop (WACV)
- 2021 🥇 First Prize · Ho Chi Minh AI Challenge
- 2020 🥇 First Prize · MediaEval Sports Video Classification
- 2017 & 2018 🥉 Third Prize · Vietnamese Math Olympiad
- 2015 & 2016 🥇 Gold Medal · April 30th Math Olympiad