Calix, Inc. logo

Staff Site Reliability Operations Engineer

Job Overview

Location

Remote - USA

Job Type

Full-time

Category

Software Engineering

Date Posted

June 18, 2026

Full Job Description

📋 Description

  • Lead the global platform reliability strategy and next-generation observability architecture on Google Cloud Platform (GCP) using the Grafana Labs telemetry stack (Grafana, Mimir, Loki, Tempo, Beyla).
  • Architect, optimize, and troubleshoot full-stack network infrastructure spanning Layer 1 through Layer 7, including physical/fiber transport, BGP/OSPF routing, TCP/QUIC transport tuning, TLS termination, DNS architecture, and application protocols like HTTP/3 and gRPC.
  • Design, scale, secure, and manage production-grade Google Kubernetes Engine (GKE) clusters with multi-cluster networking, custom controllers, and GitOps workflows.
  • Tune and maintain high-throughput Apache Kafka event streams to ensure low-latency, high-availability data delivery across distributed systems.
  • Ensure performance, scalability, and disaster recovery readiness of large-scale data ecosystems including PostgreSQL, AlloyDB, and BigQuery.
  • Implement AIOps methodologies using machine learning for time-series anomaly detection, log clustering, and correlation to reduce alert fatigue and predict infrastructure bottlenecks.
  • Integrate AIOps insights with Grafana workflows to automate incident triage, accelerate root-cause analysis, and trigger auto-remediation scripts.
  • Champion the long-term technical roadmap for distributed infrastructure engineering and cloud-native observability standards across the organization.
  • Mentor senior and junior engineers in advanced debugging techniques, distributed systems thinking, and intelligent operations within a globally distributed team.
  • Provision and manage multi-region GCP cloud architectures exclusively using HashiCorp Terraform for Infrastructure as Code.
  • Build custom infrastructure tooling, Kubernetes operators, and data integration scripts using high proficiency in Go and Python.
  • Optimize and maintain the unified observability platform by scaling Grafana Enterprise/Cloud, Prometheus/Mimir, Loki, and Tempo at enterprise scale.
  • Apply deep Linux internals knowledge, eBPF-based monitoring, kernel-level networking, and packet analysis tools (Wireshark, tcpdump) for system-level diagnostics.
  • Leverage Google Cloud architectural best practices including Cloud SDN, Cloud Armor, Interconnect, IAM, and cost optimization strategies.
  • Deliver results with high autonomy in a 100% remote engineering environment, collaborating across multiple time zones with a distributed workforce.

🎯 Requirements

  • 8+ years in SRE, Production Engineering, or Distributed Systems infrastructure roles
  • Deep technical expertise across all OSI layers (L1-L7), including physical infrastructure, BGP/OSPF, TCP/QUIC, TLS, DNS, HTTP/3, and gRPC
  • Expert-level mastery of Google Kubernetes Engine (GKE) internals, multi-cluster networking, and GitOps workflows
  • Proven track record managing high-throughput Apache Kafka pipelines and large-scale data environments (PostgreSQL, AlloyDB, BigQuery)
  • Deep hands-on experience deploying and scaling the Grafana stack (Grafana, Mimir, Loki, Tempo) in production
  • Advanced, production-scale expertise in HashiCorp Terraform for provisioning multi-region GCP architectures
  • High proficiency in Go and Python for building infrastructure tooling and automation scripts

🏖️ Benefits

  • Eligible for bonus compensation
  • Comprehensive benefits package (details available via Calix careers page link)
  • Competitive annual base pay range: $136,000–$231,000 USD (all other US locations), $156,400–$265,700 USD (San Francisco Bay Area)
  • 100% remote work with flexibility to work from anywhere in the United States or Canada

Skills & Technologies

Python
Go
Fiber
PostgreSQL
GCP
Senior
Remote
Degree Required

Ready to Apply?

You will be redirected to an external site to apply.

AI Job Fit Analysis
Pro

See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.

Calix, Inc. logo
Calix, Inc.
Visit Website

About Calix, Inc.

Calix, Inc. provides cloud and software platforms, systems and services for broadband service providers worldwide. The company offers revenue-generating cloud solutions, network management, subscriber experience, and analytics software that enable operators to deploy gigabit services. Its portfolio includes Calix Cloud, Calix Services, and Calix Systems that support fiber, copper, and coax networks. Founded in 1999 and headquartered in San Jose, California, Calix serves communications service providers, municipalities, and utilities, helping them simplify operations, reduce costs, and deliver enhanced broadband experiences to residential and business customers.

Get more remote jobs like this

Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.

Newsletter

Weekly remote jobs and featured talent.

No spam. Only curated remote roles and product updates. You can unsubscribe anytime.

Similar Opportunities

Expired
United States (Remote)
Full-time
Expired Jun 13, 2026
Fiber
Remote
Degree Required

4 months ago

Expired
Serbia
Full-time
Expired Jun 13, 2026
Python
Java
Remote

4 months ago

London
Full-time
Expires Sep 19, 2026
Onsite

11 days ago

Expired
China
Full-time
Expired Jun 25, 2026
JavaScript
TypeScript
React
+5 more

3 months ago