Lead Site Reliability Engineer
Hace 18 horas
Colombia
EPAM Systems
Jornada completa
Gratis con email o Google
Guarda esta oferta y sigue tu búsqueda
Crea una cuenta gratis para guardar empleos, crear alertas y volver a esta oferta desde tu panel.
Gratis con email o Google
Al continuar, aceptas nuestros Términos & Política de Privacidad.
EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.
We are looking for a hands-on Lead Site Reliability Engineer to maintain, enhance, and support a Java services ecosystem while partnering closely with a backend engineering team. You will strengthen reliability, observability, and on-call practices across critical services.
Responsibilities
• Provide on-call support for Java backend identity services during business hours
• Troubleshoot complex production issues using logs and telemetry to identify root causes
• Prepare and deploy patches to address issues in cloud infrastructure
• Implement reliability improvements for key identity services through practical code and configuration changes
• Build and refine metrics and dashboards to enable rapid assessment of platform health
• Monitor SLOs across backend services and drive remediation when error rates increase
• Create and improve runbooks to standardize operational responses across services Requirements
• 5+ years of experience in Site Reliability Engineering or DevOps for distributed systems
• Strong experience with Amazon Web Services in production environments
• Strong experience with Amazon DynamoDB and Amazon ElastiCache operations
• Proven experience with observability and troubleshooting in distributed systems using logs and telemetry
• Hands-on experience with Git-based workflows
• Hands-on experience with Gradle in Java service environments
• Leadership skills to guide reliability improvements and support operational decision-making
• Incident response skills to communicate operational issues clearly and concisely in writing
• Fast learning ability to absorb information quickly and apply it during on-call support
• SLO management skills to track, evaluate, and improve reliability through repeatable processes
• English proficiency: B2 (Upper-Intermediate) Nice to have
• Kubernetes
• Terraform
• Grafana
• Apache Kafka
• New Relic
We offer
• International projects with top brands
• Work with global teams of highly skilled, diverse peers
• Healthcare benefits
• Employee financial programs
• Paid time off and sick leave
• Upskilling, reskilling and certification courses
• Unlimited access to the LinkedIn Learning library and 22,000+ courses
• Global career opportunities
• Volunteer and community involvement opportunities
• EPAM Employee Groups
• Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
• Provide on-call support for Java backend identity services during business hours
• Troubleshoot complex production issues using logs and telemetry to identify root causes
• Prepare and deploy patches to address issues in cloud infrastructure
• Implement reliability improvements for key identity services through practical code and configuration changes
• Build and refine metrics and dashboards to enable rapid assessment of platform health
• Monitor SLOs across backend services and drive remediation when error rates increase
• Create and improve runbooks to standardize operational responses across services Requirements
• 5+ years of experience in Site Reliability Engineering or DevOps for distributed systems
• Strong experience with Amazon Web Services in production environments
• Strong experience with Amazon DynamoDB and Amazon ElastiCache operations
• Proven experience with observability and troubleshooting in distributed systems using logs and telemetry
• Hands-on experience with Git-based workflows
• Hands-on experience with Gradle in Java service environments
• Leadership skills to guide reliability improvements and support operational decision-making
• Incident response skills to communicate operational issues clearly and concisely in writing
• Fast learning ability to absorb information quickly and apply it during on-call support
• SLO management skills to track, evaluate, and improve reliability through repeatable processes
• English proficiency: B2 (Upper-Intermediate) Nice to have
• Kubernetes
• Terraform
• Grafana
• Apache Kafka
• New Relic
We offer
• International projects with top brands
• Work with global teams of highly skilled, diverse peers
• Healthcare benefits
• Employee financial programs
• Paid time off and sick leave
• Upskilling, reskilling and certification courses
• Unlimited access to the LinkedIn Learning library and 22,000+ courses
• Global career opportunities
• Volunteer and community involvement opportunities
• EPAM Employee Groups
• Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.