Skill diminta
AnsibleCI/CDCommunicationGrafanaKubernetesLinuxMySQLPostgreSQLPrometheusTerraformTroubleshooting
Deskripsi
About the role
We're looking for a Site Reliability Engineer to run infra operations for our live projects and products with real traffic, focusing on early troubleshooting and realtime reporting.
- What you'll do
- Run day-to-day infra operations and realtime system monitoring per shift schedule
- Do early troubleshooting and report system status/incidents in realtime to relevant teams
- Monitor workloads on Kubernetes, containers, and Linux-based servers
- Handle basic network configuration and cloud computing services for hosting and operations
- Build and monitor Grafana dashboards - prometheus stack
- Monitor and do basic troubleshooting on databases (PostgreSQL, MySQL, ClickHouse)
- Required qualifications (basic level)
- Minimum Bachelor's degree (S1) in IT, Computer Science, or related field
- Open to fresh graduates with internship experience in relevant fields
- Basic understanding of Kubernetes, Containerization, Networking, Linux, Grafana, and Cloud computing
- Database PostgreSQL, MySQL, ClickHouse
- Good communication (able to report system status clearly and in realtime)
- Willing to work shifts
- Nice to have
- Terraform / Ansible
- Mikrotik
- Comfortable using AI tools
CI/CD