Skill diminta
AWSCI/CDDeliveryGitHub ActionsGitLab CIKubernetesLeadershipTerraformTroubleshooting
Deskripsi
Job Description
- Own and continuously improve cloud infrastructure running business-critical workloads.
- Design, implement, and maintain secure, scalable, and highly available platforms on AWS.
- Lead Kubernetes platform operations, including cluster management, scalability, reliability, and security improvements.
- Drive infrastructure automation using Terraform and Infrastructure as Code best practices.
- Build and improve CI/CD platforms to enable safe, reliable, and efficient software delivery.
- Establish observability standards through monitoring, logging, alerting, and incident management practices.
- Lead incident response, root cause analysis, and postmortem processes for critical production issues.
- Partner with engineering teams to improve deployment strategies, operational readiness, and application reliability.
- Improve developer experience through automation, platform engineering initiatives, and self-service tooling.
- Lead cloud cost optimization and capacity planning initiatives.
- Define and enforce infrastructure, security, and operational best practices across the organization.
- Mentor engineers and provide technical leadership in DevOps, cloud, and platform engineering domains.
- Minimum Qualifications
- 5+ years of hands-on experience in DevOps, Platform Engineering, Site Reliability Engineering (SRE), Cloud Engineering, or related roles.
- Strong hands-on experience designing, operating, and improving production workloads on AWS.
- Proven experience managing cloud services such as EKS, EC2, RDS/Aurora, VPC, IAM, Route53, S3, CloudFront, and related AWS services.
Strong experience operating and troubleshooting Kubernetes-based production environments, including cluster administration, networking, autoscaling, ingress management, and security best practices.
Experience building, maintaining, and optimizing CI/CD platforms using GitHub Actions, GitLab CI/CD, or similar technologies.