This job is no longer available
This job expired on 19/07/2026. It no longer accepts applications.
Senior Site Reliability Engineer
Cloudbeds
Job description
About the role
Cloudbeds is looking for a Senior Site Reliability Engineer to ensure the reliability and performance of its hospitality platform that processes millions of transactions worldwide. You will design scalable AWS solutions and champion automation, resilience, and continuous improvement across a fully remote engineering team.
Key responsibilities
- Design and implement reliable, scalable AWS architecture.
- Maintain and support high‑load Kubernetes (EKS) clusters and related components.
- Support CI/CD pipelines using ArgoCD and GitHub Actions.
- Automate deployments with Terraform infrastructure‑as‑code.
- Develop and enhance observability and monitoring using Grafana, Prometheus, DataDog, and CloudWatch.
- Participate in incident management, root‑cause analysis, and performance optimization.
- Collaborate with development and security teams to establish best practices and meet reliability targets.
- Provide infrastructure support rotation and guidance to other engineering groups.
Required profile
- 5+ years of DevOps or SRE experience within the AWS ecosystem.
- 5+ years of hands‑on experience with Kubernetes (EKS) and Helm charts.
- Proven experience building CI/CD pipelines with ArgoCD and GitHub Actions.
- Strong background in Terraform for infrastructure‑as‑code.
- Expertise in observability tools such as Grafana, Prometheus, DataDog, and CloudWatch.
- Solid incident management, troubleshooting, performance analysis, and root‑cause investigation skills.
Required skills
- AWS
- Kubernetes (EKS)
- Helm
- ArgoCD
- GitHub Actions
- Terraform
- Grafana
- Prometheus
- DataDog
- CloudWatch
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Peru.
Salaries by job title
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Cloudbeds