
Job Overview
Location
San Francisco
Job Type
Full-time
Category
Engineering
Date Posted
July 10, 2026
Full Job Description
đź“‹ Description
- • This customer-facing role combines deep technical expertise with hands-on implementation.
- • You'll be instrumental in customer architecture and design, infrastructure deployment and optimization, and production operations and support.
- • Partner with clients to understand workload requirements and design optimal GPU cluster architectures.
- • Create technical proposals and capacity planning for clusters ranging from 100 to 10,000+ GPUs.
- • Develop deployment strategies for LLM training, inference, and HPC workloads.
- • Present architectural recommendations to technical and executive stakeholders.
- • Deploy and configure orchestration systems including SLURM and Kubernetes for distributed workloads.
- • Implement high-performance networking with InfiniBand, RoCE, and NVLink interconnects.
- • Optimize GPU utilization, memory management, and inter-node communication.
- • Configure parallel filesystems (Lustre, BeeGFS, GPFS) for optimal I/O performance.
- • Tune system performance from kernel parameters to CUDA configurations.
- • Serve as primary technical escalation point for customer infrastructure issues.
- • Diagnose and resolve complex problems across the full stack - hardware, drivers, networking, and software.
- • Implement monitoring, alerting, and automated remediation systems.
- • Provide 24/7 on-call support for critical customer deployments.
- • Create runbooks and documentation for customer operations teams.
- • Work directly with customers pushing the boundaries of AI, from startups training foundation models to enterprises deploying massive inference infrastructure.
- • Collaborate with our world-class engineering team while having direct impact on systems powering the next generation of AI breakthroughs.
🎯 Requirements
- • 3+ years hands-on experience with GPU clusters and HPC environments.
- • Deep expertise with SLURM and Kubernetes in production GPU settings.
- • Proven experience with InfiniBand configuration and troubleshooting.
- • Strong understanding of NVIDIA GPU architecture, CUDA ecosystem, and driver stack.
🏖️ Benefits
- • Cash Compensation Range of $150-300k plus Equity Incentives.
- • Opportunity to work with world-class engineering team.
- • Direct impact on systems powering the next generation of AI breakthroughs.
- • Opportunity to collaborate with customers pushing the boundaries of AI.
- • Opportunity to work on large-scale deployments and contribute to open-source HPC/AI infrastructure projects.
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About Prime Intellect, Inc.
San Francisco–based startup building decentralized AI infrastructure that lets researchers pool compute and data to collaboratively train large models. Founded in 2023, the company offers open-source protocols and cloud orchestration tools that aggregate GPUs across providers, coordinate distributed training, and cryptographically verify contributions so participants share ownership and future rewards of the resulting models.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

NETGEAR, Inc.
11 days ago

Unilever PLC
23 days ago

Silver.com LLC
13 days ago

Latamcent
2 months ago