AI Vision Engineer
Nexxa · SF Bay area · Full-time · Posted 2026-08-28
Job description
[**Nexxa**](http://Nexxa.ai) is building the best AI systems for heavy industries — enabling machines, systems and operations to think, decide and act autonomously across manufacturing, large-scale infrastructure, logistics and legacy environments. Our mission is to translate deep technical breakthroughs into operational reality, solving some of the hardest systems-level problems in industry. ## Role Overview We are looking for an AI Vision Engineer to help design, build, and deploy next-generation computer vision systems across a diverse set of real-world industrial applications. This role is ideal for someone with a strong foundation in Computer Vision and Machine Learning who is excited about working across the full vision stack, including classical and deep-learning-based CV, vision-language models (VLMs), multimodal reasoning, and real-time inference at the edge and in the cloud. You will work closely with our AI and engineering teams to develop production-ready vision solutions, improve model performance, and help shape the next generation of intelligent visual systems. ## What You'll Do - Design, train, evaluate, and deploy computer vision models for real-world industrial applications. - Build and optimize CV pipelines for tasks such as object detection, segmentation, classification, OCR, tracking, and visual understanding. - Develop and fine-tune vision-language models (VLMs) for multimodal reasoning, visual question answering, and document understanding. - Design and optimize real-time inference pipelines for deployment on edge devices and in the cloud. - Build scalable data pipelines for image and video collection, annotation, augmentation, training, and evaluation. - Fine-tune and evaluate open-source vision and multimodal foundation models using modern training and inference frameworks. - Develop robust evaluation frameworks and benchmarks to measure model accuracy, robustness, latency, and business impact. - Optimize models for production constraints, including quantization, pruning, and hardware-accelerated inference. - Collaborate with product, engineering, and research teams to translate business requirements into technical vision solutions. - Contribute to architecture decisions, technical design reviews, and AI/CV best practices. - Stay current with the latest advancements in computer vision, multimodal AI, and autonomous systems. ## Required Qualifications ### Education Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Mathematics, Statistics, Artificial Intelligence, or a related technical field, or equivalent practical experience. ### Experience & Skills - 3+ years of industry experience in Computer Vision, Machine Learning, Applied AI, or related fields. - Demonstrated experience independently owning and delivering computer vision projects from concept to production. - Strong programming skills in Python. - Hands-on experience with PyTorch and modern deep learning workflows. - Experience developing and deploying computer vision models (detection, segmentation, classification, OCR) in production environments. - Experience working with image and video processing libraries such as OpenCV. - Experience with common CV/detection frameworks (e.g., YOLO, Detectron2, MMDetection, or similar). - Experience working with VLMs, multimodal models, or Generative AI applications. - Strong understanding of machine learning fundamentals, model evaluation, experimentation, and model deployment. - Experience working with Hugging Face Transformers and open-source AI ecosystems. - Familiarity with data annotation workflows, dataset curation, hyperparameter optimization, and inference optimization. - Experience building production-grade software and AI systems. - Strong analytical and problem-solving skills. - Excellent communication and collaboration skills. ## Preferred Qualifications - Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Computer Vision, or a related field. - Experience with OCR, document understanding, or visual reasoning systems. - Experience with Vision-Language Models (VLMs) and multimodal AI applications. - Familiarity with 3D vision, SLAM, or sensor fusion (e.g., camera + LiDAR) for industrial or robotics applications. - Experience with real-time inference optimization (TensorRT, ONNX Runtime, quantization, pruning). - Experience deploying models on edge hardware (e.g., NVIDIA Jetson, embedded systems). - Experience with LangChain, LangGraph, or agentic AI frameworks combining vision and language models. - Experience with vector databases and visual/semantic retrieval systems. - Experience deploying AI systems on AWS, GCP, or other cloud platforms. - Experience with Docker, Kubernetes, and MLOps workflows. - Experience with PostgreSQL and large-scale data systems. - Experience with model serving, distributed training, and inference optimization at scale. - Contributions to open-source projects, technical blogs, research publications, Kaggle competitions, or other demonstrable CV/AI work. - Experience working in startup or high-growth environments. ## What We're Looking For - A builder who enjoys taking vision systems from prototype to production. - Someone comfortable working across classical Computer Vision, deep learning, and multimodal Generative AI. - An engineer who is curious, adaptable, and eager to learn emerging vision and AI technologies. - A strong collaborator who can contribute across research, engineering, and product discussions. - Someone excited about solving challenging real-world problems using computer vision. - An individual who takes ownership, moves quickly, and thrives in an environment with significant autonomy. **Why Join** [](http://nexxa.ai/)[**Nexxa.AI**](http://Nexxa.AI)**?** - **Innovative Environment:** Play a critical role in transforming heavy industries through groundbreaking AI and automation technologies. - **Collaborative Culture:** Be part of a team that values in