Lambda Inc. logo

Senior Platform Engineer - Core Infrastructure

Job Overview

Location

San Francisco Office (Fremont St)

Job Type

Full-time

Category

Engineering

Date Posted

July 21, 2026

Full Job Description

đź“‹ Description

  • • We are seeking a Senior Platform Engineer - Core Infrastructure to join our team at Lambda Inc. in San Francisco, San Jose, or Seattle office location.
  • • As a Senior Platform Engineer, you will be responsible for building and scaling our cloud offering, including the Lambda website, cloud APIs and systems, and internal tooling for system deployment, management, and maintenance.
  • • You will architect, deploy, and operate Kubernetes clusters across AWS and Lambda's bare-metal datacenters.
  • • You will build and maintain automation for cluster lifecycle management, including provisioning, upgrades, and scaling.
  • • You will own the reliability, performance, and security of Kubernetes workloads in production.
  • • You will implement observability, logging, and alerting for clusters and critical workloads.
  • • You will partner with product teams to design scalable, cloud-native services and CI/CD pipelines.
  • • You will set the standards for resource management, networking, and RBAC across the platform.
  • • You will lead incident response, root-cause analysis, and post-mortems for platform issues.
  • • You will mentor engineers and raise the bar for platform engineering across the org.
  • • You will have 5+ years in Platform, Infrastructure, or SRE roles, including running Kubernetes in production at scale.
  • • You will have deep knowledge of Kubernetes internals and day-2 operations, including upgrades, scaling, and troubleshooting.
  • • You will have strong skills with Helm, Kustomize, or similar, and GitOps-based delivery.
  • • You will have proficiency with infrastructure-as-code (Terraform, Pulumi, or equivalent).
  • • You will have a solid grounding in networking, service meshes, and container runtimes.
  • • You will have hands-on experience with observability stacks (Prometheus, Grafana, OpenTelemetry).
  • • You will have strong coding skills in Go or Python for automation and tooling.
  • • You will have practical security experience, including network policies, secrets management, and image scanning.
  • • Nice to Have: Experience with multi-cluster, multi-cloud, or hybrid environments.
  • • Nice to Have: Knowledge of GPU scheduling, HPC workloads, or ML/AI infrastructure.
  • • Nice to Have: Experience with workflow orchestration / durable execution frameworks (Temporal, Cadence, or Argo Workflows).
  • • Nice to Have: Exposure to cost optimization and capacity planning for large clusters.
  • • Nice to Have: Contributions to CNCF or Kubernetes open-source projects.
  • • Nice to Have: CKA/CKS certification.
  • • Salary Range Information: The annual salary range for this position has been set based on market data and other factors.
  • • About Lambda: Founded in 2012, with 500+ employees, and growing fast.
  • • About Lambda: Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove.
  • • About Lambda: We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG.
  • • About Lambda: Our values are publicly available: https://lambda.ai/careers.
  • • About Lambda: We offer generous cash & equity compensation.
  • • About Lambda: Health, dental, and vision coverage for you and your dependents.
  • • About Lambda: Wellness and commuter stipends for select roles.
  • • About Lambda: 401k Plan with 2% company match (USA employees).
  • • About Lambda: Flexible paid time off plan that we all actually use.
  • • Equal Opportunity Employer: Lambda is an Equal Opportunity employer.
  • • Equal Opportunity Employer: Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

🎯 Requirements

  • • 5+ years in Platform, Infrastructure, or SRE roles, including running Kubernetes in production at scale.
  • • Deep knowledge of Kubernetes internals and day-2 operations (upgrades, scaling, troubleshooting).
  • • Strong skills with Helm, Kustomize, or similar, and GitOps-based delivery.
  • • Proficiency with infrastructure-as-code (Terraform, Pulumi, or equivalent).
  • • Solid grounding in networking, service meshes, and container runtimes.
  • • Hands-on experience with observability stacks (Prometheus, Grafana, OpenTelemetry).
  • • Strong coding skills in Go or Python for automation and tooling.
  • • Practical security experience: network policies, secrets management, and image scanning.

🏖️ Benefits

  • • Generous cash & equity compensation.
  • • Health, dental, and vision coverage for you and your dependents.
  • • Wellness and commuter stipends for select roles.
  • • 401k Plan with 2% company match (USA employees).
  • • Flexible paid time off plan that we all actually use.

Skills & Technologies

Python
Go
AWS
Kubernetes
Terraform
DevOps
Senior
Hybrid

Ready to Apply?

You will be redirected to an external site to apply.

AI Job Fit Analysis
Pro

See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.

Lambda Inc. logo
Lambda Inc.
Visit Website

About Lambda Inc.

Lambda Inc. provides cloud-based GPU clusters and workstations for artificial-intelligence research and development. The company designs and operates high-performance hardware infrastructure optimized for machine-learning workloads, offering on-demand access to NVIDIA GPUs, pre-configured deep-learning software stacks, and scalable storage. Customers include AI labs, universities, and enterprises training large language and computer-vision models. Founded in 2012, Lambda is headquartered in San Francisco and maintains data centers across North America and Europe.

Get more remote jobs like this

Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.

Newsletter

Weekly remote jobs and featured talent.

No spam. Only curated remote roles and product updates. You can unsubscribe anytime.

Similar Opportunities

Dubai
Full-time
Expires Sep 14, 2026
Python
REST
Senior
+1 more

8 days ago

Remote - Munro, Argentina
Full-time
Expires Sep 2, 2026
Remote

20 days ago

Argentina
Full-time
Expires Sep 12, 2026
Python
TypeScript
AWS
+4 more

10 days ago

Expires soon
Argentina
Full-time
Expires Jul 27, 2026 (Soon)
Python
JavaScript
TypeScript
+4 more

2 months ago