Site Reliability Engineer

  •  Job Reference: 161328
  •  Industry: Information and Communications Technology
  •  brand-id: R22108699
  •  Brand Name: 02C3423

We are looking for a Site Reliability Engineer (SRE) with experience in platform engineering, DevOps, and production operations. The role involves building reliable systems, automating infrastructure, and ensuring observability across mission-critical applications. You will work closely with development, infrastructure, and operations teams to design scalable solutions, improve system resilience, and define operational best practices.

Key Responsibilities

  • System reliability: Ensure high availability, performance, and resilience of production systems.

  • Infrastructure automation: Build and maintain CI/CD pipelines, automate deployments, and manage containerized workloads.

  • Observability & monitoring: Implement logging, metrics, tracing, alerting, and dashboards using modern monitoring tools.

  • Incident response: Integrate systems with enterprise monitoring, SIEM, and incident management workflows; participate in on-call rotations.

  • Operational excellence: Define and implement runbooks, escalation paths, and production support models.

  • Collaboration: Work with cross-functional teams in Agile environments to deliver reliable and secure systems.

  • Continuous improvement: Conduct root cause analysis, implement permanent fixes, and drive automation initiatives to reduce manual effort.

Skills Required

  • 5+ years of experience as an SRE, DevOps Engineer, Platform Engineer, Infrastructure Engineer, or Production Engineer.

  • Solid knowledge of Linux administration, Shell scripting, Git, CI/CD, Docker, and Kubernetes.

  • Hands-on experience with observability platforms (Grafana, Prometheus, ELK, CloudWatch, Azure Monitor, etc.).

  • Experience integrating systems with enterprise monitoring, alerting, SIEM, or incident response workflows.

  • Proven ability to define and implement runbooks, operational procedures, escalation paths, and production support models.

  • Familiarity with cloud platforms (AWS, Azure, GCP) and infrastructure-as-code tools (Terraform, Ansible).

  • Excellent problem-solving skills, ability to troubleshoot complex systems, and experience in Agile/DevOps practices.

  • Excellent communication skills and ability to collaborate with cross-functional stakeholders.