Get to Know the Role
The Lead Infrastructure Engineer – Kubernetes & Service Mesh is a lead technical authority within the MEKS platform team. This role is responsible for designing, building, and operating the high-scale container and service mesh infrastructure. This infrastructure powers critical backend services across Grab. Reporting directly to the Platform Engineering Manager, this hands-on lead role focuses entirely on deep technical execution, platform architecture, system reliability, and developer experience. You will serve as the primary architect and technical mentor for the platform, driving multi-cluster AWS EKS strategies, Istio service mesh implementations, and cross-team developer enablement. This is an onsite position based in the Jakarta office.
Service Mesh Implementation involves designing and deploying production-grade Istio service mesh infrastructure. This infrastructure consists of a data plane and control plane. It requires managing various aspects, including mTLS, traffic management, canary/blue-green deployments, rate limiting, and circuit breaking, to optimize the service mesh.
Architect and manage large-scale AWS EKS clusters, driving multi-cluster/multi-region topologies, node-pool optimization, custom scheduler strategies, and automated cluster lifecycles.
Build and maintain self-service APIs, GitOps workflows, and automation that simplify onboarding, deployments, and traffic management for internal product engineering teams.
Partner directly with application teams to guide architectural choices, streamline workload migrations to MEKS, and accelerate time-to-market.
Create comprehensive platform documentation, golden path templates, runbooks, and reference architectures to foster engineering self-sufficiency.
Reliability, Observability & Incident Response
Implement telemetry, metrics, distributed tracing (OpenTelemetry/Jaeger), and centralized logging across clusters, Envoy proxies, and application workloads.
Establish and uphold technical SLOs/SLAs, driving proactive capacity management, resource right-sizing, and performance tuning for peak availability and optimal infrastructure efficiency.
Act as a senior technical point of escalation for complex platform outages, lead root-cause analyses , and execute corrective actions to prevent recurrence.
Collaborate with Security, SRE, Networking, and Cloud Infrastructure teams to ensure platform security, regulatory compliance, and seamless integration.
Data lowongan bersumber dari jobstreet. Tombol “Lamar” mengarahkan Anda ke halaman aslinya.