← Back to all jobs

Senior Applied AI Infrastructure Engineer - NREC

Carnegie Mellon University · PA · Posted 2026-08-26

Apply on the company site →

Job description

At the National Robotics Engineering Center (NREC), it is our engineers and technicians who drive the breakthroughs that define our success. The members of our technical staff collaborate closely with leadership and multidisciplinary teams to design, build, and deploy sophisticated robotic solutions that address complex challenges in industrial, commercial and government sectors. Each project benefits from their expertise, creativity, and hands-on problem-solving, fueling progress and innovation across the organization. As part of our dedicated team, you will work alongside world-class robotics professionals committed to pushing the boundaries of technology and redefining ideas into solutions for real-world applications. We foster a culture of professionalism, respect, and collaboration, offering a flexible and encouraging environment where you can sharpen your skills, lead impactful projects, and take control of your career development. We are seeking a dynamic Senior Applied AI Infrastructure Engineer to lead and contribute to the evaluation, deployment, and integration of secure generative and agentic AI tools, LLMs, and support of self-hosted infrastructure across engineering workflows. This is an exciting opportunity for someone who thrives in a fast-paced and innovative setting. In this role, you will be instrumental in advancing internal AI-assisted workflows, infrastructure reliability, and secure multi-GPU model serving, ensuring our team delivers exceptional and groundbreaking results. Your primary responsibilities include: Evaluating generative and agentic AI tools and recommending practical approaches to engineering leadership. Supporting cloud-hosted AI tools where appropriate and locally hosted tools where project confidentiality or data-handling requirements prohibit cloud use. Designing, implementing, documenting, testing, and maintaining internally hosted AI services and supporting infrastructure. Deploying and operating large language models on shared GPU systems and smaller project- or team-specific platforms. Integrating AI tools with engineering systems such as source-code repositories, Jira, Confluence, Jenkins, internal documentation, and test infrastructure. Developing secure tool interfaces, APIs, Model Context Protocol servers, and sandboxed environments that allow AI agents to perform useful engineering tasks. Prototyping and evaluating AI-assisted workflows for software development, testing, documentation, requirements analysis, and other engineering activities. Helping engineers use supported AI tools effectively across software, embedded, FPGA, mechanical, electrical, and other technical workflows. Developing internal documentation, examples, training materials, and reusable configurations for recommended tools and practices. Measuring the reliability and usefulness of AI-assisted workflows, including the quality of generated code, test results, review effort, and failure modes. Surveying emerging tools and techniques and implementing promising approaches where they provide practical value. Following best practices for team software development, including peer review, automated testing, version control, issue tracking, security review, and integrated documentation. Required Qualifications: B.S. in Computer Science, Computer Engineering, Electrical Engineering, or a related technical discipline, or equivalent experience. 5+ years of professional software engineering, machine-learning infrastructure, DevOps, platform engineering, or developer-tools experience. Strong Python programming skills. Linux development and system-administration experience. Familiarity with large language models, retrieval-augmented generation, tool-using agents, or AI-assisted software-development workflows. Strong technical communication and documentation skills. 3 or more of the following: Experience deploying and maintaining software services. Experience with containers and reproducible deployment tools such as Docker. Experience integrating software systems through APIs, command-line tools, authentication mechanisms, or similar interfaces. Experience with modern software engineering practices, including version control, code review, testing, CI/CD, logging, and troubleshooting. Ability to evaluate new technologies, communicate technical tradeoffs, and make practical recommendations. We especially want to hear from you if you have experience or qualifications in ANY of the following areas: Self-hosted LLM inference frameworks such as vLLM, TensorRT-LLM, llama.cpp, Ollama, NVIDIA NIM, or similar tools Multi-GPU systems, model serving, resource scheduling, or inference performance optimization Model Context Protocol servers or other structured interfaces for AI tool use Integration with Jira, Confluence, Jenkins, Git-based repositories, artifact repositories, or internal knowledge systems Coding agents that can modify code, run builds and tests, and prepare pull requests Sandboxed code execution, container isolation, secrets management, access control, or audit logging Evaluation of LLM applications, coding assistants, agents, or retrieval systems Retrieval-augmented generation, document ingestion, embeddings, reranking, or code indexing Cloud AI services and data-sensitive or disconnected AI deployments Embedded software, FPGA development, robotics, simulation, or hardware-in-the-loop testing GPU-based machine learning infrastructure Developing internal technical documentation, training, examples, or reusable engineering workflows Machine learning, computer vision, or robotics applications Other Requirements: Successful pre-employment background check {remove if not applicable} Sponsorship: Applicants for this position must be currently legally authorized to work for CMU in the United States. CMU will not sponsor or take over the sponsorship of an employment visa for this opportunity. Carnegie Mellon is not a qualifying employer for the STEM OPT benefit: only the 12-month OPT may be