Multimodal AI Researcher
Apple · Sunnyvale · Posted 2026-08-11
Job description
The Video Computer Vision organization is working on breakthrough technologies for future Apple products. Our team delivers cutting-edge AI, machine learning, computer vision and graphics algorithms that power technologies including human understanding, perception, digital humans, multimodal generative AI, and agents. Our algorithms ship across a range of Apple products, including iPhone and Apple Vision Pro, where our work has contributed to technologies like Personalized Spatial Audio, EyeSight, and Persona as well as future Apple products. We are an applied research group, we push the state of the art and then bring it to product. In this role, you will collaborate with world-class experts in AI, ML, Software, and Hardware to tackle fundamental challenges in human-centric solutions that will impact millions of users across Apple's ecosystem. Minimum Qualifications: BS and a minimum of 3 years relevant industry experience. Experience building models for multimodal perception systems. Experience working with LLMs and VLMs. Software engineering skills and proficiency in Python and PyTorch. Curiosity and willingness to learn new things in order to improve the quality of their solutions. Preferred Qualifications: MS or PhD in computer vision, computer graphics, machine learning, computer science, computer engineering or related fields. Experience in developing, training/tuning foundation models and multimodal LLMs. Experience with training and troubleshooting generative architectures such as diffusion, reinforcement learning, flow matching or normalizing flow at scale. Experience with real-time or streaming multimodal models. Experience with speech understanding and generation. Experience applying reinforcement learning to help post-train foundation models. Excellent communication and experience working with multi-functional teams. Self-motivated with proven track record to optimally prioritize and deliver tasks on schedule.