← Back to all jobs

Student Researcher (Multimodal Interaction and World Model - Seed) – 2027 Start (PhD)

ByteDance · California · Posted 2026-09-04

Apply on the company site →

Job description

About the team The Seed Multimodal Interaction and World Model team is dedicated to developing models that have human-level multimodal understanding and interaction capabilities. The team is working to advance the exploration and development of multimodal assistant products. Responsibilities - Conduct research on multimodal foundation models and related systems. - Explore methods to improve model capabilities across modalities, including areas such as vision-language modeling, world modeling, and representation learning. - Design and prototype algorithms, models, or system components. - Collaborate with the team to advance research directions. Minimum Qualifications - Currently pursuing a PhD in computer science, mathematics, engineering, or a related field. - Strong programming skills and solid foundation in algorithms and data structures, proficient in Python or C/C++. - Demonstrated research track record with publications in conferences related to multimodal learning, machine learning, or artificial intelligence. Preferred Qualifications - Experience in multimodal modeling, vision-language modeling, world modeling, simulation, or 3D representations. Experience with latent space modeling or structured representations is a plus. - Strong problem-solving ability and experience conducting independent research. - Ability to collaborate effectively in a research environment.