Staff Network Engineer, App Platform
Scale AI · CA · Posted 2026-09-03
Job description
Scale GP (Scale Generative AI Platform) is Scale’s enterprise AI platform, providing APIs and infrastructure for knowledge retrieval, inference, evaluation, agents, and more. We deploy SGP across AWS, Azure, and GCP, often directly into customer-controlled cloud environments across highly regulated industries including healthcare, financial services, telecom, and retail. As SGP has grown in scale and complexity, networking has become a critical architectural discipline of its own. Today, teams routinely encounter the same hard problems — VPC design, routing, ingress and egress, private connectivity, service exposure, address-space constraints, firewall policies, and cross-environment communication — but solve them differently depending on the deployment. We’re looking for a Senior Network Engineer to establish the architectural standards for how SGP connects to customer infrastructure and how traffic moves throughout the platform. You’ll own network architecture across a large and rapidly growing fleet of Kubernetes environments spanning AWS, Azure, and GCP, with many deployed inside customer-owned cloud accounts and networks that we do not control. The constraints vary significantly: a commercial deployment behind a customer-managed API gateway; a hub-and-spoke enterprise network where the customer assigns our address space; a GovCloud environment protected by default-deny firewall policies; or a fully air-gapped deployment with no external connectivity. The goal is not to create a bespoke network architecture for every customer. Your job is to define one clear, secure, and supportable networking model for SGP — and establish the small set of defensible variations required to operate across different clouds, customer architectures, and compliance regimes. What You'll Do • Own network architecture for the SGP platform: define and enforce network standards across cloud environments (AWS, Azure, GCP) and customer deployments • Design and review VPC architectures, peering, DNS, load balancing, and CDN/edge configurations (e.g., Cloudflare), grounding root-cause and tradeoff discussions in data and domain depth • Bring strong technical judgment to VPN, routing, access, ingress/egress, and traffic-flow decisions • Review proposed networking solutions from platform, product, and forward-deployed teams; evaluate what should and should not be introduced into the platform boundary • Define the standard connectivity pattern between Scale's control plane and customer data planes (VPC peering, PrivateLink/Private Service Connect, site-to-site VPN, reverse tunnels) — and converge today's per-customer designs onto it • Own the customer-boundary delivery blueprint: how code, images, and traffic cross into a customer tenant — CI/CD mirroring across org boundaries (including customer-side Azure DevOps), artifact scan gates, private registries, ingress — so new engagements configure a pattern instead of designing one • Own engineer access into customer and internal environments (Tailscale/Teleport/bastion-class decisions), replacing per-engineer VPN sprawl and hand-rolled tunnels with something auditable • Own the Kubernetes traffic layer: Istio/Envoy mesh and ingress, the per-cloud CNI matrix, and a portable NetworkPolicy contract that constrains egress for thousands of short-lived agent-sandbox pods running untrusted code • Reduce rework by preventing one-off implementations from becoming long-term platform burden • Partner with security engineering on network segmentation, zero-trust access, and compliance requirements in regulated customer environments • Debug complex connectivity, latency, and traffic-flow issues across hybrid and multi-cloud deployments • Document network architecture, standards, and runbooks so adjacent teams can operate confidently What We're Looking For • 5+ years of network engineering experience, including designing and operating production networks in cloud environments • Deep expertise in cloud networking on at least two of AWS, Azure, and GCP: VPC design, peering, Transit Gateway/hub-and-spoke topologies, private connectivity (PrivateLink, Private Service Connect), and DNS • Kubernetes networking depth — CNI, ingress, NetworkPolicy, service mesh (Istio/Envoy) — this is where most of our real incidents live • Strong fundamentals in TCP/IP, BGP, routing, firewalls, VPN (site-to-site and client), and TLS • Experience with edge/CDN and traffic-management platforms such as Cloudflare • Has defined a network standard or reference architecture that other teams adopted, and enforced it through design review • Comfortable as the sole domain owner: able to collect requirements across five-plus live environments you didn't design, then converge them without breaking any • Has made build-vs-adopt networking calls with vendor-support or contractual consequences in a customer's cloud • Infrastructure-as-code proficiency (Terraform preferred) • Familiarity with network security and compliance requirements in regulated industries (healthcare, finance, government), including environments across multiple compliance regimes (commercial, FedRAMP/GovCloud, air-gapped), is a plus • Excellent communication skills — able to explain networking tradeoffs to both technical and non-technical audiences and influence decisions without direct authority Why This Role Matters Without a dedicated owner, adjacent teams are covering a domain that needs specialized expertise. You will be the point of accountability for network architecture as SGP grows — raising the quality of every deployment, unblocking confident decisions, and keeping the platform boundary clean as we scale. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work locati