
Job Overview
Location
Remote - United States
Job Type
Full-time
Category
Engineering
Date Posted
August 8, 2026
Full Job Description
đź“‹ Description
- • We are seeking an AI Cluster Architect to create and refine large-scale GPU cluster architectures within strict power and infrastructure limits.
- • The architect must understand how different GPU SKUs, NICs, switches, and fabrics interact at scale, including their individual and aggregate power and thermal characteristics.
- • The role requires deep experience navigating heterogeneous environments, multiple generations of hardware, and end user requirements.
- • The architect will evaluate multi-plane, rail-optimized, and tiered fabric designs across technologies like InfiniBand, RoCE, and SpectrumX to ensure the networking architecture supports the intended GPU count without overrunning facility limits or switch radix and/or topology constraints.
- • The role balances customer-specific requirements for compute, storage, and service density, ensuring that the final cluster design maintains acceptable levels of GPU and fabric performance, while maximizing the number of usable GPUs within the total power budget.
- • The architect will model and validate power consumption across the full cluster bill of materials (GPUs, CPUs, NICs, switches, fabric components, storage, and facility limits).
- • The architect will gather, interpret, and maintain detailed SKU-level power and thermal specifications for GPUs, NICs, switches, DPUs, storage, and server platforms.
- • The architect will develop power-aware cluster configuration templates and capacity-planning models that can scale across sites with varying constraints and allow for quick iteration and ideation.
- • The architect will provide guidance on future-proofing, including the ability to incorporate next-gen GPUs, NICs, or fabrics.
- • The architect will collaborate with vendors on novel fabric architectures that enable large-scale cluster deployments (100k+ GPUs).
🎯 Requirements
- • 7+ years designing or building large-scale HPC, AI, or hyperscale GPU clusters.
- • Expert understanding of GPU and accelerator system design, including node topology, PCIe/NVLink/NVSwitch/ROCm, and NIC-to-GPU affinity considerations.
- • Strong familiarity with InfiniBand, RoCE, and SpectrumX networking, including multi-tier, multi-plane, Clos/dragonfly variants, and large-radix switch design.
- • Demonstrated experience modeling power draw and thermal characteristics of servers, GPUs, NICs, switches, optics, and storage systems.
- • Ability to design networks that maintain full non-blocking performance or intentionally introduce over/under-subscription while understanding impacts on workload performance.
- • Proven ability to gather and analyze vendor SKU-level specifications and incorporate them into scalable cluster architectures.
- • Experience balancing customer-driven requirements for compute, storage, and service density in combination with overall GPU count.
🏖️ Benefits
- • Excellent Medical Benefits w/ 100% company-paid premiums for employee only plan + 100% company-paid dental & vision premiums.
- • 401(k) plan that matches 100% up to 4% with immediate vesting.
- • Professional Development Reimbursement of $2,500 each year.
- • 11 Holidays + Paid Time Off Accrual + Rollover Plan + take your birthday off.
- • Commitment matters to Vultr! Increased PTO at 3 year & 10 year anniversary + 1 month paid sabbatical every 5 years + Anniversary Bonus each year.
- • $500 first year remote office setup + $400 each following year for new equipment.
- • Internet reimbursement up to $75 per month.
- • Gym membership reimbursement up to $50 per month.
- • Company-paid Wellable subscription.
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About The Constant Company, LLC
The Constant Company, LLC operates the Vultr cloud infrastructure brand, providing on-demand compute, storage, bare-metal, and managed Kubernetes services from 32 global data centers. Founded in 2014, the company targets developers, SaaS businesses, and enterprises with hourly billing, API-driven provisioning, and standardized hardware. Services include virtual machines, block storage, load balancers, object storage, managed databases, and cloud GPUs, all accessible through a unified control panel and REST API. Vultr emphasizes price-performance, global reach, and rapid deployment for web applications, CI/CD workflows, and edge workloads without long-term contracts.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

NETGEAR, Inc.
29 days ago

Unilever PLC
1 month ago

Silver.com LLC
1 month ago

Latamcent
3 months ago