
Job Overview
Location
San Francisco, CA - US
Job Type
Full-time
Category
Software Engineering
Date Posted
June 26, 2026
Full Job Description
đź“‹ Description
- • Own the day-to-day management, allocation, and strategic bin-packing optimization of Crusoe Cloud’s high-performance GPU fleet to maximize utilization and minimize fragmentation across accelerated computing clusters.
- • Act as the primary technical capacity partner to Sales, Customer Success, and Solutions Engineering, translating commercial customer pipelines into physical hardware requirements within data centers.
- • Collaborate with Fleet Management, Infrastructure Engineering, and Data Center Operations to ensure maximum uptime and seamless performance for AI/ML customer workloads.
- • Design and implement utilization modeling systems to track real-time GPU cluster headroom, workload densities, and allocation velocities to prevent supply bottlenecks.
- • Synthesize commercial demand signals and large-scale AI training/inference architectural trends to inform strategic hardware placement and scheduling decisions.
- • Partner with Core Software Engineering to identify manual allocation processes and translate them into scalable, automated programmatic scheduling and visualization tools.
- • Maintain deep alignment between Crusoe’s technical infrastructure and Go-To-Market engine to support the company’s mission of accelerating AI infrastructure with an energy-first approach.
- • Monitor and optimize GPU resource allocations across distributed clusters, ensuring efficient use of NVIDIA H100/B200 ecosystems and other accelerated compute hardware.
- • Communicate complex infrastructure constraints and capacity limitations clearly to non-technical business stakeholders, including Sales and Customer Success teams.
- • Translate commercial requirements from GTM teams into precise hardware specifications and scheduling timelines for engineering and operations teams.
- • Contribute to strategic forecasting by analyzing trends in AI workloads and their impact on GPU demand, cluster density, and infrastructure scaling needs.
- • Ensure alignment between physical data center layouts and customer workload requirements to reduce idle capacity and improve overall system efficiency.
- • Support environmental sustainability goals by optimizing energy-intensive AI compute operations through efficient resource utilization and reduced waste.
- • Participate in cross-functional planning cycles to align capacity projections with product roadmaps, customer commitments, and infrastructure expansion timelines.
- • Maintain accurate documentation of GPU allocation policies, optimization strategies, and capacity constraints for internal audit and operational continuity.
- • Proactively identify bottlenecks in GPU supply chains or scheduling workflows and propose data-driven solutions to improve throughput and customer satisfaction.
- • Work within a vertically integrated AI infrastructure environment that spans from energy generation to cloud compute, ensuring capacity planning reflects the full stack.
- • Contribute to the development of metrics and dashboards that provide visibility into GPU utilization, allocation efficiency, and forecast accuracy across global data centers.
- • Balance competing demands from multiple high-priority customers while maintaining fairness, transparency, and predictability in resource allocation.
- • Stay informed on advancements in GPU architectures, AI training frameworks, and cloud-native compute patterns to continuously refine capacity planning methodologies.
🎯 Requirements
- • 3+ years of direct experience in Infrastructure Capacity Planning, Technical Product Management, or Systems Engineering with heavy exposure to machine-level resource scaling.
- • Experience working within a hyperscaler cloud environment (e.g., AWS, GCP, Azure, Oracle Cloud) or a specialized, large-scale AI/accelerated compute cloud fabric.
- • Solid foundational understanding of GPU topologies (e.g., NVIDIA H100/B200 ecosystems).
- • Proven ability to communicate complex infrastructure and physical layout constraints clearly to business stakeholders like Sales and Customer Success, while translating commercial requirements back to hardware engineers.
- • Bachelor’s or Master's degree in Computer Engineering, Computer Science, Operations Research, Industrial Engineering, Data Science, or an equivalent quantitative field.
🏖️ Benefits
- • Competitive compensation and equity packages with Restricted Stock Units
- • Paid time off, paid holidays & leave of absence programs
- • Comprehensive health, dental & vision insurance
- • Employer contributions to HSA account
- • Paid parental leave
- • 401(k) Retirement plan with company match up to 4% of salary
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About Crusoe Energy Systems LLC
Crusoe Energy Systems is building Crusoe Cloud, an AI cloud platform that provides managed AI services and AI data center infrastructure. They cater to businesses seeking to accelerate AI solution development with optimized models and high-performance computing. The company utilizes environmentally aligned power sources, including wind, solar, and natural gas, to power its data centers. With features like managed Kubernetes and Slurm, Crusoe simplifies operations and ensures reliability with 24/7 support. Crusoe is expanding its reach, including a strategic European expansion with its first data center in Norway. Crusoe recently raised $1.375 billion at a valuation above $10 billion.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

NVIDIA Corporation
3 months ago

UiPath, Inc.
2 months ago

Samsara Inc.
2 months ago

Crusoe Energy Systems LLC
2 months ago