The Site Reliability Engineering (SRE) team architects, builds, and maintains the rock-solid infrastructure that applications rely on. At the Senior Level, you own reliability, performance, and cost outcomes for the systems under your area end-to-end, not just executing well-defined tasks, but deciding between tradeoffs, scoping ambiguous problems, and driving process and system improvements that span teams. You'll work closely with development, security, and product teams, and mentor other engineers as a technical point of reference for the team.
What You Will Do
Own the availability, performance, scalability, and security of production systems end-to-end, across cloud (AWS/GCP) and on-premises environments.
Scope and lead medium-to-large infrastructure initiatives: gather requirements, prioritize by business impact, and communicate impact to stakeholders.
Lead structured incident investigation, isolating server, database, and application layers with a metrics-first approach, including on-call during high-traffic events.
What Are We Looking For
At least 4 years of experience in either SRE, DevOps, MLOps, or platform engineering, including senior-level scope at a high-traffic company.
Deep expertise in one major cloud provider (preferably AWS), with a proven ability to ramp up on the other quickly.Production experience & expertise with Kubernetes & Linux fundamentals
Data lowongan bersumber dari jobstreet. Tombol “Lamar” mengarahkan Anda ke halaman aslinya.