Skill diminta
AWSCI/CDCommunicationDatadogDockerGitGitHub ActionsGitLab CIGrafanaJenkinsKubernetesLinuxMicroservicesPrometheusPythonRabbitMQRedisTerraformTroubleshooting
Deskripsi
Location
Menteng, Central Jakarta (WFO)
Job Description
We are looking for a Site Reliability Engineer (SRE) to build, maintain, and improve the reliability, scalability, and performance of our production systems. You will work closely with Software Engineers, DevOps, QA, and Product teams to ensure our services remain highly available and resilient.
Responsibilities
- Maintain and improve the reliability, availability, and scalability of production systems.
- Design, implement, and manage cloud infrastructure on AWS.
- Build and maintain CI/CD pipelines to support automated deployments.
- Automate infrastructure provisioning using Infrastructure as Code (IaC).
- Monitor application and infrastructure performance using observability tools.
- Respond to production incidents, troubleshoot system issues, and perform root cause analysis (RCA).
- Optimize system performance, reliability, and operational efficiency.
- Implement backup, disaster recovery, and security best practices.
- Collaborate with engineering teams to improve system architecture and deployment processes.
- Develop automation scripts to eliminate repetitive operational tasks.
- Participate in on-call rotation and incident response when required.
Requirements
- Bachelor's degree in Computer Science, Information Technology, or a related field.
- Minimum
- 3 years
- of experience as a Site Reliability Engineer, DevOps Engineer, or Infrastructure Engineer.
- Strong experience with
- AWS Cloud Services
- (EC2, ECS/EKS, RDS, S3, IAM, VPC, CloudWatch).
- Strong knowledge of Linux system administration.
- Experience with Docker and Kubernetes.
- Experience implementing Infrastructure as Code using Terraform.
- Experience with CI/CD tools such as GitLab CI, GitHub Actions, or Jenkins.
- Familiarity with monitoring and observability tools such as Prometheus, Grafana, ELK Stack, Loki, or Datadog.
- Strong scripting skills using Bash, Python, or Go.
- Understanding of networking concepts (TCP/IP, DNS, Load Balancer, Reverse Proxy).
- Experience troubleshooting production systems under high traffic.
- Good understanding of system security and cloud best practices.
- Strong analytical and problem-solving skills.
- Excellent communication and collaboration skills.
- Preferred Qualifications
- AWS Certification is a plus.
- Experience with microservices architecture.
- Experience working in startup, e-commerce, fintech, or OTA environments.
- Familiarity with Redis, Kafka, RabbitMQ, or other distributed systems.
- Knowledge of cost optimization and cloud infrastructure management.
- Experience implementing SLI, SLO, and SLA.
- Experience with incident management and postmortem documentation.
- Technical Skills
Cloud
AWS
Container
Docker, Kubernetes
IaC
Terraform
CI/CD
GitLab CI, GitHub Actions, Jenkins
Monitoring
Prometheus, Grafana, ELK Stack, Loki, Datadog
OS
Linux
Scripting
Bash, Python, Go
Version Control
Git