Hi, I'm Tianlu Zheng.
I am a Multimodal Algorithm Engineer at NIO’s Autonomous Driving division. I received my M.S. in Robotics Science and Engineering from Northeastern University, China, in June 2026. My work focuses on multimodal representation learning, multimodal large language models, and efficient model training and inference.
Previously, I worked on multimodal algorithms and large-model post-training at Tencent, Baidu, DeepGlint, and China Telecom Beijing Research Institute.
News
- 2026.06 Graduated with an M.S. from Northeastern University and joined NIO's Autonomous Driving division as a Multimodal Algorithm Engineer.
- 2025.09 Our work on robust text-based person retrieval was released on arXiv and accepted to the EMNLP Main Conference.
- 2025.08 Joined Tencent as a Multimodal Algorithm Intern.
- 2025.07 Started a research project on multimodal humor understanding through lateral thinking.
- 2025.06 Joined Baidu ERNIE as a Large Language Model Algorithm Intern.
Education
Recommended admission. Ranked 5/49. Research interests: multimodal representation learning and large language models.
GPA: 3.87/4.00, top 5% of the major. CET-6.
Experience





Publications
Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval
A dual-masking and gradient-attention framework for improving robustness in text-based person retrieval.
arXivDual-Level Open-Vocabulary 3D Scene Representation for Instance-Aware Robot Navigation

Open-vocabulary 3D scene representations designed for instance-aware robotic navigation.
The Solution to the WWW25 Text-based Person Anomaly Search Challenge
A multimodal solution for text-based person anomaly search presented at the WWW '25 workshop.
DOIIndoor Semantic Mapping and Navigation Method and System Based on Multimodal Models
Research
Improving Qwen2.5-VL's ability to understand humor that requires divergent, multi-step reasoning.
- Trained a humor judge on selection, ranking, and rewriting tasks.
- Built a multi-agent generation framework with expander, generator, and supervisor roles.
- Combined LoRA supervised fine-tuning with a second DPO stage and judge-guided preference-data updates.
- Improved performance by about 35% over Qwen2.5-VL-7B, 6% over GPT-4o, and 3% over the referenced state-of-the-art method on the humor dataset.
