Data Center Manager
Backblaze · AZ · Posted 2026-07-22
Job description
About Backblaze Backblaze provides reliable, high-availability cloud storage trusted by consumers, SMBs, enterprises, and developers in over 150 countries. Backblaze B2 Cloud Storage supports data-intensive workloads including backup, media, analytics, and modern AI pipelines. Our teams focus on building durable, scalable systems with a strong emphasis on developer experience and operational efficiency. Role Overview The Data Center Operations Manager leads the day-to-day management of the organization's infrastructure, personnel, and vendor relationships across one or more data center sites or regions. The Manager works cross-functionally across engineering, security, finance, and operations teams to align infrastructure priorities with broader organizational goals. This role is responsible for maintaining mission-critical systems at peak reliability while ensuring operations remain secure, compliant, and cost-effective. The Manager is the primary owner of uptime, operational integrity, and continuous improvement across the team. KEY RESPONSIBILITIES • Team Leadership & People Management ▸ Recruit, onboard, mentor, and evaluate a diverse team of data center technicians and inventory controllers. ▸ Foster an inclusive environment that encourages collaboration, safety, and continuous professional development. ▸ Set and track individual performance goals, growth plans, and career pathways for all direct reports. ▸ Partner with HR on DEI initiatives, succession planning, and workforce development. ▸ Delegate responsibilities to ensure workload balance and effective utilization of team skills. • Project Management ▸ Define and manage project scope, timelines, budgets, and resources in collaboration with stakeholders. ▸ Lead concurrent projects including hardware refreshes, capacity expansions, and migrations. ▸ Maintain project documentation including MOPs, change control records, and post-mortems. ▸ Coordinate internal teams, vendors, and contractors to meet milestones; report status to leadership on a defined cadence. • Risk Mitigation & Business Continuity ▸ Maintain and regularly test Business Continuity Plans (BCP) and Disaster Recovery Plans (DRP). ▸ Conduct risk assessments to identify threats to availability, integrity, and confidentiality of data assets. ▸ Enforce incident response protocols, ensuring rapid containment, root-cause analysis, and corrective action. ▸ Manage physical access controls and coordinate with cybersecurity teams on information security alignment. • Compliance & Regulatory Adherence ▸ Ensure adherence to applicable standards including SOC 2, ISO 27001, HIPAA, PCI-DSS, GDPR, and NIST. ▸ Maintain audit-ready documentation, evidence repositories, and control libraries; support internal and third-party assessments. ▸ Enforce policies governing data handling, access control, change management, and incident reporting. ▸ Ensure staff complete required compliance training and maintain role-relevant certifications. • Equipment Maintenance & Infrastructure Support ▸ Manage preventive and corrective maintenance programs for servers, storage arrays, and networking hardware. ▸ Oversee hardware installation, decommissioning, and lifecycle replacement in line with policy. ▸ Maintain vendor and OEM relationships to ensure timely support, warranty compliance, and SLA adherence. ▸ Monitor environmental conditions via DCIM tools and escalate anomalies through appropriate channels. ▸ Enforce cable management, labeling, and documentation standards; maintain 24/7 readiness through on-call and escalation procedures. ▸ Own vendor SLA governance across all colocation providers, ensuring contractual obligations are met and leading escalations and remediation when they are not. • Data Analysis & Performance Reporting ▸ Track and report KPIs including uptime (SLA/SLO), PUE, capacity utilization, and incident rates. ▸ Use DCIM, ITSM, and monitoring platforms (Zabbix, Grafana) to generate operational insights and capacity forecasts. ▸ Deliver executive dashboards and reports that translate technical metrics into business narratives. ▸ Perform root-cause and trend analysis on incidents; benchmark performance and identify optimization opportunities. • Inventory Coordination & Asset Management ▸ Maintain an auditable asset register in CMDB/ITAM platforms (Netbox, NetSuite) across all sites. ▸ Coordinate procurement, receiving, staging, and deployment of hardware and consumable supplies. ▸ Manage spare parts pools and lifecycle disposition, including secure data sanitization (NIST 800-88) and e-waste compliance. ▸ Conduct regular physical inventory audits and reconcile discrepancies against system records. ▸ Collaborate with procurement and finance to forecast hardware needs and manage purchase orders. QUALIFICATIONS & REQUIREMENTS Education ▸ Bachelor's degree in Information Technology, Computer Science, or a related field required, or equivalent professional experience in data center operations. Experience ▸ 10+ years of progressive experience in data center operations, IT, or infrastructure management. ▸ 5+ years in a supervisory or management role overseeing technical teams. ▸ Proven track record managing projects and vendor relationships in mission-critical environments. Certifications (Preferred) ▸ PMP | ITIL 4 Foundation or higher | CompTIA Server+ or Network+ Core Competencies ▸ Strong leadership and team-building skills with a commitment to diversity, equity, and inclusion. ▸ Analytical mindset with attention to operational detail and the ability to drive data-informed decisions. ▸ Clear communicator — able to translate technical concepts for non-technical stakeholders. ▸ Experience with DCIM, ITSM (Jira, Confluence), asset management (NetSuite, NetBox), and monitoring tools (Zabbix, Grafana). ▸ Effective at managing competing priorities. WORKING CONDITIONS ▸ Occasional after-hours, weekend, or on-call availability required for critical incidents and maintenance windows. ▸ Work perfor