
Job Overview
Location
Remote - United States
Job Type
Full-time
Category
Engineering
Date Posted
July 9, 2026
Full Job Description
đź“‹ Description
- • Lead the engineering team responsible for the day-to-day implementation, scaling, and operation of AI compute clusters.
- • Translate engineering roadmaps and technical requirements from the Director of AI Infrastructure into detailed project plans and execution milestones.
- • Drive delivery of cluster deployments, hardware bring-up, node configuration, and integration with orchestration and scheduling systems.
- • Ensure cluster reliability, uptime, and performance through monitoring, automation, and continuous operational improvements.
- • Oversee lifecycle operations for bare metal and GPU fleets, including provisioning, configuration management, firmware/driver updates, and hardware validation.
- • Manage incident response for GPU and cluster infrastructure, ensuring timely resolution and root-cause analysis.
- • Work closely with AI/ML, SRE, Networking, and Hardware Engineering teams to ensure cluster capabilities meet training and inference needs.
- • Coordinate with Product to confirm technical requirements, feature readiness, and delivery timelines.
- • Support integrations across networking, storage, scheduler, and resource orchestration components.
- • Improve tooling and automation for cluster provisioning, observability, configuration management, and large-scale fleet operations.
- • Contribute to the development and refinement of multi-tenant scheduling, workload management, and orchestration systems in partnership with senior technical staff.
- • Identify performance bottlenecks and propose engineering-level optimizations.
- • Coach and mentor engineers, fostering a high-performance, detail-oriented engineering culture.
- • Support career development, expectations, and performance management for team members.
- • Help refine engineering processes, including code reviews, testing standards, documentation, and operational runbooks.
🎯 Requirements
- • 6–10 years of experience in infrastructure engineering, HPC, large-scale systems, or similar fields.
- • Strong understanding of AI compute infrastructure, including GPU/CPU clusters, distributed training architectures, and high-performance networking (InfiniBand/RDMA).
- • Experience running production bare metal, GPU, or hardware fleet operations at meaningful scale.
- • Hands-on expertise with Linux systems, Kubernetes or Slurm, provisioning tools (Terraform, Ansible), observability platforms, and networking fundamentals.
- • Proven track record in cluster operations, hardware bring-up, distributed systems, or ML workload support.
- • Experience leading engineering teams or pods, with the ability to manage execution while staying close to technical work.
- • Ability to communicate effectively with cross-functional engineering teams and translate strategy into actionable engineering tasks.
- • Strong execution mindset with the ability to prioritize, deliver, and adapt in a fast-paced environment.
🏖️ Benefits
- • Excellent Medical Benefits w/ 100% company-paid premiums for employee only plan + 100% company-paid dental & vision premiums.
- • 401(k) plan that matches 100% up to 4% with immediate vesting.
- • Professional Development Reimbursement of $2,500 each year.
- • 11 Holidays + Paid Time Off Accrual + Rollover Plan + take your birthday off.
- • Commitment matters to Vultr! Increased PTO at 3 year & 10 year anniversary + 1 month paid sabbatical every 5 years + Anniversary Bonus each year.
- • $500 first year remote office setup + $400 each following year for new equipment.
- • Internet reimbursement up to $75 per month.
- • Gym membership reimbursement up to $50 per month.
- • Company-paid Wellable subscription.
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About The Constant Company, LLC
The Constant Company, LLC operates the Vultr cloud infrastructure brand, providing on-demand compute, storage, bare-metal, and managed Kubernetes services from 32 global data centers. Founded in 2014, the company targets developers, SaaS businesses, and enterprises with hourly billing, API-driven provisioning, and standardized hardware. Services include virtual machines, block storage, load balancers, object storage, managed databases, and cloud GPUs, all accessible through a unified control panel and REST API. Vultr emphasizes price-performance, global reach, and rapid deployment for web applications, CI/CD workflows, and edge workloads without long-term contracts.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

NETGEAR, Inc.
12 days ago

Unilever PLC
23 days ago

Silver.com LLC
13 days ago

Latamcent
2 months ago