
Job Overview
Location
San Francisco Office (Fremont St)
Job Type
Full-time
Category
Software Engineer
Date Posted
July 16, 2026
Full Job Description
đź“‹ Description
- • Design, build, and maintain scalable control plane services, operators, and custom controllers for Kubernetes.
- • Develop automation for cluster lifecycle management (provisioning, upgrades, patching, and deletion).
- • Develop internal tools, APIs, and command-line interfaces (CLIs) that enable customers and ML and AI teams to deploy and monitor inference services effectively.
- • Write resilient systems that gracefully handle failure across large-scale distributed environments.
- • Define and implement Service-Level Objectives (SLOs) and Service-Level Indicators (SLIs) for Kubernetes services, workloads, and the platform.
- • Drive into systems at a low level to solve unique cluster problems and write up the findings.
- • Assist customers with high-level Kubernetes questions and integrations with applications, storage, and authentication.
- • Assist with initial cluster buildouts and validation to help identify failed hardware before customer delivery.
- • Work closely with our HPC Ops and Datacenter Ops teams on issues that require lower-level expertise or cross-functional solutions.
- • Participate in a well-managed, sustainable on-call rotation.
🎯 Requirements
- • Have a Bachelor’s degree or foreign equivalent in Computer Engineering, Computer Science, Electrical Engineering, or related field.
- • Have 5 years of progressive experience as Software Engineer or related occupation.
- • Must have at least 1 year of prior work experience in each of the following:
- • Distributed systems software development as applied to computer networking infrastructure.
- • Full-stack development in multiple software languages including Python, Go, Java, or C/C++.
- • Software development using cloud-native technologies including Docker, Kubernetes, Helm, Etcd, and gRPC.
- • Development and CI/CD workflow using tools including Gerrit, Git, and Jenkins.
- • Working with and developing open-source projects.
- • Tuning Kubernetes configuration and working with Operators, CRDs, CSI, and CNI.
🏖️ Benefits
- • Generous cash & equity compensation.
- • Health, dental, and vision coverage for you and your dependents.
- • Wellness and commuter stipends for select roles.
- • 401k Plan with 2% company match (USA employees).
- • Flexible paid time off plan that we all actually use.
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About Lambda Inc.
Lambda Inc. provides cloud-based GPU clusters and workstations for artificial-intelligence research and development. The company designs and operates high-performance hardware infrastructure optimized for machine-learning workloads, offering on-demand access to NVIDIA GPUs, pre-configured deep-learning software stacks, and scalable storage. Customers include AI labs, universities, and enterprises training large language and computer-vision models. Founded in 2012, Lambda is headquartered in San Francisco and maintains data centers across North America and Europe.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

Silver.com LLC
2 months ago

Anyone AI Inc.
3 months ago

Nexus Mutual
2 months ago
2 months ago
