The Constant Company, LLC logo

Senior Manager, AI Infrastructure Operations

Job Overview

Location

Remote - United States

Job Type

Full-time

Category

Engineering

Date Posted

July 9, 2026

Full Job Description

đź“‹ Description

  • • Lead the engineering team responsible for the day-to-day implementation, scaling, and operation of AI compute clusters.
  • • Translate engineering roadmaps and technical requirements from the Director of AI Infrastructure into detailed project plans and execution milestones.
  • • Drive delivery of cluster deployments, hardware bring-up, node configuration, and integration with orchestration and scheduling systems.
  • • Ensure cluster reliability, uptime, and performance through monitoring, automation, and continuous operational improvements.
  • • Oversee lifecycle operations for bare metal and GPU fleets, including provisioning, configuration management, firmware/driver updates, and hardware validation.
  • • Manage incident response for GPU and cluster infrastructure, ensuring timely resolution and root-cause analysis.
  • • Work closely with AI/ML, SRE, Networking, and Hardware Engineering teams to ensure cluster capabilities meet training and inference needs.
  • • Coordinate with Product to confirm technical requirements, feature readiness, and delivery timelines.
  • • Support integrations across networking, storage, scheduler, and resource orchestration components.
  • • Improve tooling and automation for cluster provisioning, observability, configuration management, and large-scale fleet operations.
  • • Contribute to the development and refinement of multi-tenant scheduling, workload management, and orchestration systems in partnership with senior technical staff.
  • • Identify performance bottlenecks and propose engineering-level optimizations.
  • • Coach and mentor engineers, fostering a high-performance, detail-oriented engineering culture.
  • • Support career development, expectations, and performance management for team members.
  • • Help refine engineering processes, including code reviews, testing standards, documentation, and operational runbooks.

🎯 Requirements

  • • 6–10 years of experience in infrastructure engineering, HPC, large-scale systems, or similar fields.
  • • Strong understanding of AI compute infrastructure, including GPU/CPU clusters, distributed training architectures, and high-performance networking (InfiniBand/RDMA).
  • • Experience running production bare metal, GPU, or hardware fleet operations at meaningful scale.
  • • Hands-on expertise with Linux systems, Kubernetes or Slurm, provisioning tools (Terraform, Ansible), observability platforms, and networking fundamentals.
  • • Proven track record in cluster operations, hardware bring-up, distributed systems, or ML workload support.
  • • Experience leading engineering teams or pods, with the ability to manage execution while staying close to technical work.
  • • Ability to communicate effectively with cross-functional engineering teams and translate strategy into actionable engineering tasks.
  • • Strong execution mindset with the ability to prioritize, deliver, and adapt in a fast-paced environment.

🏖️ Benefits

  • • Excellent Medical Benefits w/ 100% company-paid premiums for employee only plan + 100% company-paid dental & vision premiums.
  • • 401(k) plan that matches 100% up to 4% with immediate vesting.
  • • Professional Development Reimbursement of $2,500 each year.
  • • 11 Holidays + Paid Time Off Accrual + Rollover Plan + take your birthday off.
  • • Commitment matters to Vultr! Increased PTO at 3 year & 10 year anniversary + 1 month paid sabbatical every 5 years + Anniversary Bonus each year.
  • • $500 first year remote office setup + $400 each following year for new equipment.
  • • Internet reimbursement up to $75 per month.
  • • Gym membership reimbursement up to $50 per month.
  • • Company-paid Wellable subscription.

Skills & Technologies

Node.js
Kubernetes
Terraform
Linux
DevOps
Senior
Remote
$150k-160k

Ready to Apply?

You will be redirected to an external site to apply.

AI Job Fit Analysis
Pro

See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.

The Constant Company, LLC logo
The Constant Company, LLC
Visit Website

About The Constant Company, LLC

The Constant Company, LLC operates the Vultr cloud infrastructure brand, providing on-demand compute, storage, bare-metal, and managed Kubernetes services from 32 global data centers. Founded in 2014, the company targets developers, SaaS businesses, and enterprises with hourly billing, API-driven provisioning, and standardized hardware. Services include virtual machines, block storage, load balancers, object storage, managed databases, and cloud GPUs, all accessible through a unified control panel and REST API. Vultr emphasizes price-performance, global reach, and rapid deployment for web applications, CI/CD workflows, and edge workloads without long-term contracts.

Get more remote jobs like this

Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.

Newsletter

Weekly remote jobs and featured talent.

No spam. Only curated remote roles and product updates. You can unsubscribe anytime.

Similar Opportunities

Dubai
Full-time
Expires Sep 14, 2026
Python
REST
Senior
+1 more

12 days ago

Remote - Munro, Argentina
Full-time
Expires Sep 2, 2026
Remote

23 days ago

Argentina
Full-time
Expires Sep 12, 2026
Python
TypeScript
AWS
+4 more

13 days ago

Expired
Argentina
Full-time
Expired Jul 27, 2026
Python
JavaScript
TypeScript
+4 more

2 months ago