
Job Overview
Location
U.S. Remote
Job Type
Full-time
Category
DevOps
Date Posted
June 30, 2026
Full Job Description
📋 Description
- • Lead on-call escalations for critical data center incidents, triaging issues virtually with deep system knowledge to guide field teams without overwhelming them.
- • Travel to data center sites 50%+ of the time to respond to live incidents, conduct post-incident reviews on-site, and implement operational improvements directly in the field.
- • Own end-to-end root cause analysis (RCA) for significant operational events, ensuring corrective actions eliminate entire classes of failure—not just isolated incidents.
- • Analyze patterns across fleet-wide incidents and RCAs to identify the highest-value learnings, prioritizing actions that deliver maximum reliability impact while avoiding scope creep.
- • Transfer proven operational practices and fixes from one data center campus to another, ensuring solutions become standardized across the entire fleet before failures recur.
- • Design, write, and enforce the operational assessment standard, auditing each campus against it and feeding findings directly into the corrective action loop.
- • Build operational frameworks from scratch, including assessment, audit, qualification, and training programs, rather than inheriting legacy systems.
- • Operate at scale equivalent to a G7 nation’s electricity consumption, managing infrastructure targeting 100 GW by the end of the decade.
- • Function in a live construction environment, maintaining flawless operations while data centers are actively being built and scaled.
- • Define and evolve the company’s operational model as it scales 100x, setting standards that will be adopted by thousands of future team members.
- • Maintain extreme ownership of operational outcomes, often expanding scope beyond core responsibilities to ensure issues are resolved end-to-end.
- • Apply first-principles thinking to challenge existing assumptions, rejecting analogies and prioritizing the best solution regardless of origin.
- • Drive velocity by acting with urgency and intensity, contributing to long hours and high-pressure environments focused on accelerating AI infrastructure deployment.
- • Cultivate a culture of love for the mission, recognizing that building civilization-scale AI infrastructure is the most critical technical challenge of our time.
🎯 Requirements
- • Proven experience running live critical data center operations and leading teams of operators under high-stakes conditions.
- • Demonstrated ability to triage complex incidents remotely, knowing precisely when to escalate and when to empower field teams to resolve issues independently.
- • Extensive track record of authoring root cause analyses that eliminate systemic failure classes, not just individual incidents, and tracking corrective actions to full closure.
- • Experience auditing operational sites against self-authored standards and holding consistent performance bars across geographically dispersed locations.
- • History of traveling extensively between data center sites to implement improvements, transfer best practices, and leave each location better than found.
- • Proven ability to build operational assessment, audit, or training programs from scratch without relying on inherited systems.
🏖️ Benefits
- • Competitive total compensation package including salary and equity.
- • Retirement or pension plan aligned with local norms.
- • Health, dental, and vision insurance.
- • Generous PTO policy aligned with local norms.
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About FluidStack Inc.
FluidStack Inc. operates a distributed cloud platform that aggregates under-utilized GPUs in data centers and individual machines worldwide, renting them on-demand to AI researchers, startups, and enterprises for training and inference workloads. The company automates deployment, security, and billing, offering prices up to 80% below traditional hyperscalers while providing instant access to high-end NVIDIA A100, H100, and consumer GPUs through a single API and web console. Headquartered in London, FluidStack targets machine-learning engineers who need scalable, low-cost compute without long-term commitments, claiming thousands of active nodes and customers including Fortune 500 enterprises and leading research labs.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

Web.com Group, Inc.
2 months ago

Public Cloud Group AG
1 month ago

PAR Technology Corporation
3 days ago

Haast Technologies Inc.
2 months ago