Principal Core Infrastructure Engineer
Oracle · TN · Posted 2026-08-27
Job description
Oracle Cloud Infrastructure builds and operates large-scale cloud services in a distributed, multi-tenant environment. The SPLAT team owns critical platform services that provide secure API routing, service registration, traffic management, private connectivity, authentication, and operational controls for OCI services. SPLAT sits in the request path for more than 300 OCI control planes and processes hundreds of billions of API requests each month. Our engineering challenges span high-throughput Java services, distributed systems, networking, security, observability, capacity management, deployment automation, and production reliability. We are looking for a hands-on technical leader who can own substantial systems and cross-service initiatives. You will lead architecture and delivery, write and review production code, resolve complex operational problems, influence partner teams, and develop other engineers. The role requires strong distributed-systems judgment and the ability to make progress when requirements, ownership, or failure domains are not initially clear. Responsibilities: Architecture and DeliveryLead significant systems and initiatives from problem definition and design through implementation, rollout, adoption, and production validation.Translate scalability, security, reliability, and business requirements into clear technical designs and execution plans.Make sound tradeoffs involving availability, consistency, latency, throughput, durability, cost, and operational complexity.Design for partial failures, retries, duplicate requests, mixed-version deployments, dependency degradation, and regional disruption.Write and review secure, maintainable, well-tested Java code.Define service contracts, compatibility requirements, migration plans, validation strategies, and rollback criteria.Scalability and Operational ExcellenceEstablish capacity models, performance objectives, scaling strategies, and load-testing plans for high-throughput services.Design effective throttling, load shedding, backpressure, caching, concurrency, and failure-recovery mechanisms.Define useful service indicators, objectives, metrics, alarms, dashboards, runbooks, and deployment safeguards.Lead complex incident investigations and convert recurring failures or manual procedures into automation and preventive engineering improvements.Serve as a technical escalation point for problems that cross application, infrastructure, network, or organizational boundaries.Technical LeadershipProvide architectural direction in one or more critical areas such as routing, authentication, private connectivity, runtime performance, observability, or deployment infrastructure.Decompose broad initiatives so multiple engineers can own meaningful work while maintaining architectural consistency.Mentor engineers through design, code review, delivery, and incident response.Raise engineering quality through reusable systems, tools, standards, and operational practices.Contribute to hiring and help identify architectural investments, platform gaps, and reliability risks for the team roadmap.Cross-Team ExecutionAlign SPLAT and partner teams on technical decisions, responsibilities, dependencies, and rollout plans.Communicate complex designs, tradeoffs, risks, and progress clearly to engineers and leaders.Make progress under ambiguity by separating facts, assumptions, reversible decisions, and external dependencies.Adjust direction when production evidence or new technical information invalidates earlier assumptions.Use modern development and AI-assisted tools responsibly to improve engineering quality and productivity.Minimum QualificationsBachelor’s degree in Computer Science, Computer Engineering, or a related field, or equivalent practical experience.8+ years of experience designing, building, and operating production backend or platform services.Strong development experience in Java or another modern object-oriented language.Strong understanding of distributed systems, concurrency, fault tolerance, and production operations.Experience leading substantial technical initiatives across multiple engineers or teams.Demonstrated ability to diagnose complex production issues and improve service reliability.Strong written and verbal communication skills.Preferred QualificationsExperience with high-throughput, low-latency services, HTTP, networking, proxies, or API gateways.Experience with authentication, authorization, TLS, certificates, private connectivity, or multi-tenant security.Experience with throttling, load shedding, caching, and capacity planning.Experience with cloud infrastructure, Kubernetes, infrastructure as code, CI/CD, and deployment automation.Experience improving a broader engineering organization through mentoring, shared tooling, or technical standards.SkillsSoftware Design and DevelopmentBackend Programming LanguagesDistributed SystemsSystem Design Qualifications: Disclaimer: Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements. Range and benefit information provided in this posting are specific to the stated locations only US: Hiring Range in USD from: $114,600 to $234,600 per annum. May be eligible for bonus, equity, and compensation deferral. Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business. Candidates are typically placed into the range based on the preceding factors as well as internal peer equity. Oracle US offers a comprehensive benefits package which includes the following: 1. Medical, dental, and vision insurance, including expert medical opinion 2. Short term disability and long term disability 3. Life insurance and AD&D 4. Supplemental life insurance (Employee/Spouse/Child) 5. He