BA-T: An Iterative Transformer for Two-View Bundle Adjustment
arXiv / Project Page / Code (coming soon)
An iterative Transformer that applies bundle-adjustment-style updates as a single lightweight, repeatable layer in token space.
Hi! I am a final-year PhD student at TU Munich and University of Oxford through ELLIS, advised by Daniel Cremers and Andrea Vedaldi. I also work closely with Chuanxia Zheng and Nikita Araslanov. I am currently a Research Scientist Intern at Meta Reality Labs Research, and was previously a student researcher at Google XR.
Before my PhD, I received my Master's degree from ETH Zurich, advised by Marc Pollefeys, and worked on 3D vision projects at Microsoft Spatial AI Lab. Earlier, I obtained my Bachelor's degree from CUHK and interned at SenseTime Research.
My research spans 3D/4D reconstruction & generation, world models, and embodied AI: learning how dynamic worlds look, move, and evolve, so that agents can predict and act in them.
arXiv / Project Page / Code (coming soon)
An iterative Transformer that applies bundle-adjustment-style updates as a single lightweight, repeatable layer in token space.
Preprint / Project Page (under construction)
A token-based framework that fuses partial point-cloud sequences over time into complete, temporally coherent 4D geometry.
arXiv / Paper / Project Page / Code
A feed-forward transformer that builds a global latent representation to generate complete, amodal 3D structure from a set of unposed images.
arXiv / Paper / Project Page / Code
A method for consistent dynamic scene reconstruction via motion decoupling, bundle adjustment, and global refinement.
arXiv / Project Page / Code
An unsupervised visual odometry method that improves pose estimation in dynamic scenes.
arXiv / Paper / Project Page / Code
A method for learning camera poses and intrinsics from dynamic casual videos.
Dynamic radiance field reconstruction from only two images, enabled by object-level bundle adjustment.
arXiv / Project Page / Code / Video
A robust visual odometry system leveraging temporal context with long-term point tracking to tackle occlusions and dynamic environments.
A visual localization pipeline using rendered data from NeRF, uncertainty-guided novel view selection, and evidential scene coordinate regression.
arXiv / Project Page / Video
An accurate and reliable pipeline for dense two-view SfM using weighted bundle adjustment with robust outlier filtering and learning-based confidence modeling.
Webly supervised learning for semantic label confusion using visual-semantic graph with metadata-aware anchor selection and GNN-based label propagation.
Webly supervised learning for noisy label classification via sample-wise web label correction with model confidence and pseudo machine label.