
Job Overview
Location
San Francisco, CA
Job Type
Full-time
Category
Software Engineering
Date Posted
June 19, 2026
Full Job Description
đź“‹ Description
- • As a Production Engineer, Compute (GPU) at FluidStack Inc., you will play a critical role in building and operating one of the world’s largest GPU fleets to power frontier AI infrastructure, directly enabling faster deployment of aligned AI systems that expand human freedom.
- • Day to day, you will own compute fleet health end to end by building metrics pipelines and unified health views, design and automate GPU repair and qualification workflows, manage Redfish/BMC tooling for firmware telemetry, and orchestrate live compute migrations at GW scale while ensuring reliability and scalability.
- • You will join a high-velocity, first-principles-driven team focused on rethinking every layer of the AI compute stack — from power acquisition and data center design to hardware operation — where ownership, speed, and technical excellence are paramount.
- • In this role, you will develop deep expertise in hyperscale GPU infrastructure, automation at the silicon-to-orchestration layer, and production reliability engineering, while shaping the operational foundation for civilization-scale AI compute.
🎯 Requirements
- • You treat toil as a bug and actively eliminate manual steps in repair workflows through automation.
- • You have an instinct for hardware and can reason about failure modes at the firmware and silicon level.
- • You are fluent with AI tooling including LLM APIs, MCP servers, and agentic frameworks, and use tools like Claude Code or Cursor daily.
- • You have shipped production automation that other teams depend on and are comfortable working in any language with AI coding assistance.
- • Bonus: Experience with hardware lifecycle management, RMA automation, BMC/Redfish/IPMI, GPU qualification/burn-in, workflow engines (Temporal/Cadence), or metrics pipelines (Prometheus/Grafana).
- • You move toward ambiguity, learn quickly in unfamiliar domains, and thrive in high-intensity, ownership-driven environments.
🏖️ Benefits
- • Competitive total compensation package (salary + equity) with a base salary range of $175,000 - $300,000 per year.
- • Retirement or pension plan aligned with local norms.
- • Health, dental, and vision insurance.
- • Generous PTO policy in line with local standards.
- • Opportunity to work on civilization-scale AI compute infrastructure with high autonomy and end-to-end ownership.
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About FluidStack Inc.
FluidStack Inc. operates a distributed cloud platform that aggregates under-utilized GPUs in data centers and individual machines worldwide, renting them on-demand to AI researchers, startups, and enterprises for training and inference workloads. The company automates deployment, security, and billing, offering prices up to 80% below traditional hyperscalers while providing instant access to high-end NVIDIA A100, H100, and consumer GPUs through a single API and web console. Headquartered in London, FluidStack targets machine-learning engineers who need scalable, low-cost compute without long-term commitments, claiming thousands of active nodes and customers including Fortune 500 enterprises and leading research labs.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

xAI Corporation
3 months ago

Ramp Business Corporation
2 months ago

Hims & Hers Health, Inc.
2 months ago
2 months ago
