
Job Overview
Location
Remote in the United States or Canada (East Coast/ET)
Job Type
Full-time
Category
Software Engineering
Date Posted
July 2, 2026
Full Job Description
đź“‹ Description
- • Design, build, and operate shared platform foundations including GCP infrastructure, Kubernetes, networking, routing, CI/CD, and observability systems that engineers rely on daily.
- • Diagnose and troubleshoot complex distributed systems operating at high request volumes, including Content Lake handling ~75,000 requests per second and ~4.5 million per minute.
- • Ensure comprehensive observability across the platform by analyzing system behavior, improving dashboards, alert severity, and paging standards to enhance incident response.
- • Contribute to modernization efforts of edge, caching, and gateway layers by migrating to Fastly and tightening observability across the entire stack.
- • Improve deployment reliability by creating golden paths, production readiness checks, safe rollouts, and automation to reduce cognitive load on engineering teams.
- • Raise the reliability bar through structured on-call practices, incident response improvements, and feedback loops that make production systems healthier over time.
- • Participate in an on-call rotation to respond to incidents and support the rollout of developer on-call practices across teams.
- • Mentor engineers through code reviews, design reviews, and pairing sessions to elevate technical standards and foster SRE culture.
- • Work closely with development teams to design infrastructure that enables real-time content authoring, processing, and distribution globally.
- • Operate and optimize core technologies including Kubernetes, Prometheus, Elasticsearch, PostgreSQL, NATS, Kong, Fastly, and Google Cloud Platform.
- • Maintain high availability and security for customer-facing systems used by global brands such as SKIMS, Figma, Riot Games, Anthropic, COMPLEX, Nordstrom, Arc’teryx, and Morningbrew.
- • Embrace an open but considered approach to adopting new technologies while maintaining stability and performance at scale.
- • Build systems and standards that reduce toil, improve incident communication, and create sustainable operational practices under high-pressure conditions.
- • Collaborate with a globally distributed team with reasonable overlap with European engineering hours to support 24/7 global operations.
🎯 Requirements
- • 5+ years of experience as part of an SRE on-call rotation
- • Experience with Kubernetes for orchestrating, scaling, and managing containerized applications in cloud-based environments
- • Experience building CI/CD pipelines
- • Experience with an observability stack (e.g., Prometheus)
- • Experience managing scalable, highly available, cloud-based applications with high request volume and customer-facing uptime expectations
- • Comfortable working across CDNs, edge, gateways, and caching layers, or eager to deepen expertise in these areas
🏖️ Benefits
- • Comprehensive health plans and perks
- • Competitive stock options program and location-based salary
- • Positive, flexible, and trust-based work environment that encourages long-term professional and personal growth
- • A healthy work-life balance that accommodates individual and family needs
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About Sanity, Inc
Sanity, Inc. offers a headless content platform that treats content as structured data, enabling teams to author, manage, and deliver content via APIs to any digital channel. The product includes an open-source editing environment (Sanity Studio), a hosted Content Lake for real-time storage and delivery, media asset management, and developer tools for custom schemas, integrations, and automation. Sanity targets product, engineering, and editorial teams that need flexible content models, collaborative editing, and programmatic publishing across web, mobile, and app experiences. The company provides enterprise features for security, performance, and governance and integrates with common developer frameworks and services.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities
27 days ago

Vic.ai Inc.
28 days ago
3 months ago
2 months ago

