Skill diminta
AWSAnsibleAzureCI/CDCommunicationDockerGCPGrafanaKubernetesLinuxPrometheusPythonTerraform
Deskripsi
- Key responsibilities
- Design, implement and maintain monitoring, alerting and observability solutions across our infrastructure and applications
- Automate operational tasks and processes to reduce manual workload and improve system reliability
- Manage system performance tuning, capacity planning and resource optimisation
- Respond to and resolve incidents with a focus on root cause analysis and prevention of future occurrences
- Develop and maintain runbooks, documentation and standard operating procedures for system management
- Collaborate with development teams to improve deployment processes and system architecture from a reliability perspective
- Implement infrastructure as code (IaC) practices and version control for all infrastructure configurations
- Participate in on-call rotations and provide timely incident response and escalation support
- Conduct post-incident reviews and drive continuous improvement initiatives
- Evaluate and recommend new tools, technologies and methodologies to enhance system reliability
- What we're looking for
Proven experience as a Site Reliability Engineer, Systems Administrator or similar role in a production environment managing complex systems and infrastructure
- Strong knowledge of Linux/Unix operating systems and shell scripting (Bash, Python or similar)
- Hands-on experience with cloud platforms such as AWS, Google Cloud Platform or Microsoft Azure
- Proficiency with containerisation technologies including Docker and Kubernetes
- Experience with infrastructure as code tools such as Terraform, Ansible or CloudFormation
- Solid understanding of networking concepts including TCP/IP, DNS, load balancing and firewalls
- Familiarity with monitoring and logging tools such as Prometheus, Grafana, ELK Stack or similar solutions
- Strong problem-solving skills with the ability to troubleshoot complex system issues
- Excellent communication and collaboration skills, with the ability to work effectively in cross-functional teams
- Knowledge of CI/CD pipelines and deployment automation is highly desirable
- Relevant certifications such as AWS Solutions Architect, Kubernetes Administrator or similar are preferred
- Apply now
If you are an experienced IT SRE with a passion for reliability engineering and system optimisation, we would love to hear from you. Please submit your CV, a cover letter outlining your relevant experience, and any portfolio or reference materials that demonstrate your capabilities. Join PT ABISHAR TECHNOLOGIES INDONESIA and make a meaningful impact on our infrastructure and operations.