Cloud Platform Technical Lead ID92209
Guarda esta oferta y sigue tu búsqueda
Crea una cuenta gratis para guardar empleos, crear alertas y volver a esta oferta desde tu panel.
Al continuar, aceptas nuestros Términos & Política de Privacidad.
AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.
WHY JOIN US
If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you
ABOUT THE ROLE
We are looking for a Cloud Platform Technical Lead to own the reliability and orchestration of an enterprise data platform in a regulated healthcare environment.
MUST HAVES
- 6+ years of professional experience in Cloud Engineering, Platform Engineering, DevOps or Site Reliability Engineering
- Deep hands‑on expertise with AWS production environments , including cloud architecture, security, networking fundamentals, access management, monitoring, capacity, and operational troubleshooting.
- Advanced experience with Kubernetes and Amazon EKS , including workload deployment, cluster and application troubleshooting, observability, scaling, upgrades, access, and reliability.
- Hands‑on experience operating Argo Workflows or a comparable orchestration platform supporting production data workloads.
- Strong experience with Infrastructure as Code, preferably Terraform , and with source‑controlled configuration, CI/CD, release automation, and rollback practices.
- Strong experience designing and operating observability, logging, monitoring, alerting, and incident‑routing solutions for distributed production platforms.
- Demonstrated ability to lead major incidents and cross‑functional troubleshooting across infrastructure, applications, data pipelines, and analytics layers.
- Experience operating production platforms with defined service levels, escalation paths, runbooks, change controls, release processes, and on‑call responsibilities.
- Strong working knowledge of modern data platforms, including Snowflake, S3‑based data lakes, SQL, dbt, managed ingestion tools such as Fivetran or HVR, and custom data pipelines .
- Understanding of data quality, freshness, lineage, schema evolution, pipeline dependencies, backfills, and recovery procedures.
- Proficiency in Python, Shell, Bash, or comparable languages for automation and operational tooling.
- Ability to make well‑reasoned architecture and operational decisions, communicate tradeoffs, estimate work, identify risks, and guide teams through change.
- Proven experience mentoring engineers, reviewing technical work, delegating ownership, and improving engineering processes across a distributed team.
- Strong stakeholder‑management and communication skills, including the ability to collect requirements, explain technical risks, and present recommendations to client leaders.
- Strong written and verbal English communication skills.
- Availability to work within the LatAm service window of approximately 9:00 AM to 6:00 PM Eastern Time and participate in an agreed senior escalation and on‑call rotation .
NICE TO HAVES
- Experience with data observability platforms such as SYNQ and operational tooling such as Splunk, PagerDuty, Opsgenie, or comparable solutions.
- Experience implementing dependency‑aware alerting, selective auto‑remediation, impact analysis, or other advanced reliability practices.
- Experience leading a managed service, platform operations team, Site Reliability Engineering function, or follow‑the‑sun support model.
- Experience in healthcare, financial services, or another regulated environment.
WHAT YOU WILL DO
- Own the technical direction and end‑to‑end reliability of the managed Data Platform across AWS, Amazon EKS, Kubernetes, Argo Workflows, Snowflake, S3, Fivetran and custom ingestion pipelines, and Tableau dependencies.
- Lead the technical transition into managed services, including platform discovery, dependency mapping, risk identification, knowledge transfer, shadowing, reverse shadowing, and readiness validation by platform tower.
- Establish and evolve cloud architecture principles, Infrastructure as Code standards, CI/CD and release practices, observability patterns, operational controls, and platform engineering priorities.
- Provide technical leadership for AWS infrastructure, Kubernetes and EKS operations, Argo Workflows, deployment automation, secrets and access management, platform capacity, and production reliability.
- Translate business, service, security, and platform needs into actionable technical requirements, implementation plans, and a prioritized continuous‑improvement backlog.
- Guide cross‑platform decisions involving Snowflake, db