Data Center Facility Operations Reliability Engineer
Ever wondered what it takes to keep the digital world running smoothly, day in and day out? At Meta, our data centers are the beating heart of our global services, and their seamless operation relies on an incredible team of experts. We are looking for an experienced and highly motivated Reliability Engineer to join our Asset Management & Reliability team within Facility Operations. This is your chance to make a profound impact, safeguarding the stability and efficiency of our critical infrastructure. If you thrive on solving complex technical challenges and are passionate about preventing issues before they arise, then read on. We offer a fully remote, worldwide opportunity, inviting the best talent from across the globe to contribute to our mission.
Overview
As a Data Center Facility Operations Reliability Engineer, you will be a pivotal member of our Asset Management & Reliability team, embedded directly within Facility Operations. Your work will focus on identifying and proactively managing asset reliability risks throughout every stage of the end to end asset lifecycle for our extensive Data Center Operations. This role demands a strategic thinker who can navigate complex technical landscapes, contribute to roadmap development for reliability initiatives, and solve intricate problems that ensure uninterrupted operations across both new and existing Meta sites globally. Managing diverse stakeholders spread across various time zones will be crucial for the success of individual projects and our overarching asset management and reliability program.
Key Responsibilities
- Take the lead in performing operational Failure Mode and Effects Analysis (FMEA), then develop and implement robust maintenance strategies.
- Oversee and govern the content of our Global Maintenance Library, making sure stringent Maintenance Change Management processes are consistently followed.
- Drive Corrective Maintenance initiatives by conducting thorough Failure Mode Analysis to pinpoint root causes and put effective preventive measures in place.
- Evaluate all relevant regulatory and compliance requirements, establishing governance and guaranteeing strict adherence across our operations.
Requirements
- Demonstrated experience in reliability engineering or a related field, preferably within data center, critical infrastructure, or complex facility operations.
- A strong understanding of asset lifecycle management, risk assessment, and maintenance best practices.
- Proven ability to conduct Failure Mode and Effects Analysis (FMEA) and root cause analysis.
- Experience in developing and implementing maintenance strategies and procedures.
- Excellent communication and stakeholder management skills, capable of working effectively with teams across different time zones and disciplines.
- Self motivation and a proactive approach to problem solving and continuous improvement.
- Technical proficiency in evaluating regulatory and compliance standards relevant to facility operations.
What You'll Gain
Joining Meta means becoming part of a leading technology company that impacts billions of lives worldwide. You will have the unique opportunity to work on cutting edge infrastructure, collaborate with brilliant minds, and contribute directly to the reliability of systems that power global communication. This fully remote position offers unparalleled flexibility, allowing you to thrive professionally from anywhere in the world. Grow your expertise, take ownership of critical initiatives, and advance your career within a culture that values innovation, impact, and continuous learning.
How to Apply
Click the apply button below to view the full job details and submit your application directly through the employer's official page.

