
Job Overview
Location
San Francisco
Job Type
Full-time
Category
Software Engineering
Date Posted
June 24, 2026
Full Job Description
đź“‹ Description
- • The Staff Security Reliability Engineer is a senior technical owner responsible for designing, building, and operating reliable, secure, and scalable infrastructure that underpins identity, access, endpoint, and shared platform services across OpenAI, ensuring uptime, safety, and recoverability in mission-critical environments.
- • Day-to-day responsibilities include designing and operating infrastructure across on-prem, hybrid, and shared environments; establishing standardized infrastructure-as-code patterns using Terraform, Chef, and Ansible; owning the full lifecycle of critical platforms; building observability and incident response mechanisms; automating high-toil workflows; and translating operational learnings into durable technical standards.
- • The Infrastructure Engineering team within IT focuses on replacing bespoke, one-off infrastructure with standardized, repeatable systems to increase reliability and operational leverage as OpenAI scales, supporting critical R&D and internal services.
- • In this role, you will deepen your expertise in Site Reliability Engineering, identity and access management, infrastructure-as-code, and secure platform operations while influencing cross-functional teams and setting architectural standards that enable faster, safer innovation across the organization.
🎯 Requirements
- • 10+ years of hands-on experience operating and architecting mission-critical infrastructure in high-reliability environments
- • Experience as a senior technical owner for designing and maturing complex on-prem, hybrid, or cloud-integrated systems, setting durable architectural patterns used by multiple teams
- • Experience applying Site Reliability Engineering principles at scale using observability, automation, and incident learnings to reduce risk and operational toil
- • Experience operating infrastructure for R&D, specialized labs, manufacturing, or other safety-critical environments where uptime and recoverability are essential
- • Experience with fleet, endpoint, or virtual desktop platforms such as FleetDM, Chef, or Azure Virtual Desktop
- • Experience partnering closely with identity or security engineering teams on hardened, policy-enforced infrastructure at scale
🏖️ Benefits
- • Opportunity to work on mission-critical infrastructure that powers OpenAI’s internal services and R&D environments
- • High-leverage role with significant influence on security, reliability, and scalability across the organization
- • Work in a collaborative, innovative environment focused on safely deploying AI for the benefit of humanity
- • In-office presence in San Francisco HQ with access to cutting-edge technology and expert teams
- • Commitment to equal opportunity and reasonable accommodations for applicants with disabilities
- • Background checks administered in compliance with local fair chance ordinances, including San Francisco and California Fair Chance Act
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About OpenAI, Inc.
OpenAI is a San Francisco-based artificial intelligence research and deployment company founded in 2015. It develops large-scale AI models such as GPT, DALL-E, and Codex, providing cloud APIs and consumer applications like ChatGPT. Originally established as a non-profit, it later created a capped-profit subsidiary to attract capital while maintaining its mission to ensure artificial general intelligence benefits all of humanity.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

Lendable Ltd
2 months ago

Latamcent
2 months ago

Upvest GmbH
3 months ago

Handshake Technologies, Inc.
3 months ago