Project Index

Research Projects

My work spans world models for planning and anticipation, action understanding in egocentric and exocentric video, and earlier research in visual and multimodal segmentation.

5 Publications
NeurIPS Latest, 2026
4 Venues
NeurIPS 2026
TrajPilot uses predicted future camera trajectories to distinguish possible outcomes from an egocentric view
Egocentric Prediction Action Planning

How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction

Sejoon Jun, Hai Nguyen-Truong, Luigi Seminara, Lorenzo Torresani

TrajPilot predicts possible future camera trajectories from first-person video and uses them to guide action prediction. Motion supplies a fine-grained signal of intent, helping the model plan action sequences and anticipate what someone will do next without observing their future path.

Egocentric Video Camera Trajectories Action Prediction
BMVC 2026
OMR method comparison: guidance during generation versus scoring decoded videos
World Models Video Generation

Off-Manifold Refinement: Guiding Video Generators with a Frozen World Model

Hai Nguyen-Truong, Tuan-Anh Vu, Dang Huynh

Video generators can make convincing frames but get the physics wrong. OMR uses a frozen world model to steer a frozen video generator during a single sampling run. A small adapter connects the generator’s intermediate latents to the world model, so it can guide the video before it is finished.

World Models Video Generation Physical Plausibility
WACV 2026
TransCues
Transparent Object Segmentation Transformer

Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues

Tuan-Anh Vu, Hai Nguyen-Truong, Zheng Ziqiang, Binh-Son Hua, Qing Guo, Ivor Tsang, Sai-Kit Yeung

TransCues introduces an efficient transformer-based segmentation architecture capable of handling transparent, reflective, and general objects. By proposing Boundary Feature Enhancement (BFE) and Reflection Feature Enhancement (RFE), we enable the model to better capture subtle details in both glass and non-glass regions, resulting in more accurate and robust segmentation.

Segmentation Transformer Transparent Objects
WACV 2025
VATEX
Referring Expression Segmentation Vision-Language

Vision-Aware Text Features in Referring Expression Segmentation: From Object Understanding to Context Understanding

Hai Nguyen-Truong, E-Ro Nguyen, Tuan-Anh Vu, Minh-Triet Tran, Binh-Son Hua, Sai-Kit Yeung

VATEX is a novel method for referring image segmentation that leverages vision-aware text features to improve text understanding. By decomposing language cues into object and context understanding, the model can better localize objects and interpret complex sentences, leading to significant performance gains.

Segmentation Referring Expression Multimodal
ISBI 2022
SegTransVAE
Medical Image Segmentation CNN + Transformer

SegTransVAE: Hybrid CNN - Transformer with Regularization for medical image segmentation

Hai Nguyen-Truong, Quan-Dung Pham, Nam Nguyen Phuong, Khoa NA Nguyen, Chanh DT Nguyen, Trung Bui, Steven QH Truong

SegTransVAE is the first work exploiting the hybrid architecture between CNN, Transformers with the Variational Autoencoder (VAE) branch to the network to reconstruct the input images jointly with segmentation.

Medical Imaging Segmentation Transformer VAE
No projects match the current search.