Hi, I'm Tianlu Zheng.

I am a Multimodal Algorithm Engineer at NIO’s Autonomous Driving division. I received my M.S. in Robotics Science and Engineering from Northeastern University, China, in June 2026. My work focuses on multimodal representation learning, multimodal large language models, and efficient model training and inference.

Previously, I worked on multimodal algorithms and large-model post-training at Tencent, Baidu, DeepGlint, and China Telecom Beijing Research Institute.

News

  • 2026.06 Graduated with an M.S. from Northeastern University and joined NIO's Autonomous Driving division as a Multimodal Algorithm Engineer.
  • 2025.09 Our work on robust text-based person retrieval was released on arXiv and accepted to the EMNLP Main Conference.
  • 2025.08 Joined Tencent as a Multimodal Algorithm Intern.
  • 2025.07 Started a research project on multimodal humor understanding through lateral thinking.
  • 2025.06 Joined Baidu ERNIE as a Large Language Model Algorithm Intern.

Education

Northeastern UniversityM.S. in Robotics Science and Engineering

Recommended admission. Ranked 5/49. Research interests: multimodal representation learning and large language models.

North China University of TechnologyB.Eng. in Automation

GPA: 3.87/4.00, top 5% of the major. CET-6.

Experience

NIOMultimodal Algorithm Engineer, Autonomous Driving
TencentMultimodal Algorithm Intern
Baidu ERNIELarge Language Model Algorithm Intern
DeepGlintMultimodal Algorithm Intern
China Telecom Beijing Research InstituteGPU Programming and AI Operator Optimization Intern

Publications

EMNLP MainFirst author

Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval

GA-DMS framework overview

A dual-masking and gradient-attention framework for improving robustness in text-based person retrieval.

arXiv
IROS OralFirst author

Dual-Level Open-Vocabulary 3D Scene Representation for Instance-Aware Robot Navigation

DLOV-3D scene representation overview

Open-vocabulary 3D scene representations designed for instance-aware robotic navigation.

WWW '25 Workshop OralCo-first author

The Solution to the WWW25 Text-based Person Anomaly Search Challenge

A multimodal solution for text-based person anomaly search presented at the WWW '25 workshop.

DOI
PatentApplication 2025110619980

Indoor Semantic Mapping and Navigation Method and System Based on Multimodal Models

Research

Multimodal Humor Image Understanding via Lateral ThinkingResearch Project

Improving Qwen2.5-VL's ability to understand humor that requires divergent, multi-step reasoning.

  • Trained a humor judge on selection, ranking, and rewriting tasks.
  • Built a multi-agent generation framework with expander, generator, and supervisor roles.
  • Combined LoRA supervised fine-tuning with a second DPO stage and judge-guided preference-data updates.
  • Improved performance by about 35% over Qwen2.5-VL-7B, 6% over GPT-4o, and 3% over the referenced state-of-the-art method on the humor dataset.