Senior Site Reliability Engineer

Hace 1 semana

, Colombia OpsGenius Trabajo remoto Jornada completa
Senior Site Reliability Engineer (Azure)

Location:
Remote, LATAM

Job Type
Full-Time, with benefits OpsGenius is a boutique US-based cloud operations firm. We take over the operational side of large production platforms so engineering teams can get back to building. We are hiring a Senior SRE to join a small, senior team standing up the operational foundation for a large multi-tenant SaaS platform on Azure, running at significant scale across many environments and a high-volume API surface. The platform has scaled quickly, and you will be one of the people building the operational foundation alongside it: observability, on-call, and deployment practices. This is hands-on production ownership, not a ticket queue. WHAT YOU WILL DO Build out observability: SLO definitions, alert rationalization, synthetic checks, and dashboards using Azure Monitor, Application Insights, and potentially Datadog Stand up and participate in an on-call rotation, including paging tooling, escalation paths, and incident documentation Modernize deployments: blue/green or canary patterns with validated rollback, automated smoke tests, and Infrastructure as Code with Terraform Perform Azure SQL performance work: query optimization, indexing strategy, and execution plan analysis across a large multi-tenant estate Own backup and restore strategy for Azure SQL, including point-in-time recovery testing and periodic restore drills Tune elastic pool sizing and evaluate DTU versus vCore tradeoffs for cost and performance Optimize Cosmos DB RU consumption and partition key design for high-throughput workloads Operate containerized workloads on Azure Container Apps, alongside App Service and Service Bus Write operational runbooks (incident triage, rollback, backup restore, secret rotation, certificate renewal) that any on-call engineer can follow Contribute to Azure cost optimization: right-sizing, autoscaling tuning, tagging, and cost reporting Support security and reliability hardening over time, including IAM reviews, backup restore drills, and DR exercises Collaborate daily with a US-based team during overlapping hours WHAT WE ARE LOOKING FOR Required 7+ years in SRE, DevOps, or cloud infrastructure roles, including direct ownership of production systems under an on-call rotation Deep, hands-on Azure experience. AWS-primary backgrounds with light Azure exposure will not be a fit Hands-on Azure SQL DBA

experience:
indexing, query tuning, HA/DR (failover groups, geo-replication), and backup/restore. This is not a generalist SRE role; real database ownership is expected Comfort reading query execution plans and diagnosing performance regressions at the database level, not just infrastructure monitoring Terraform or equivalent IaC in production Experience with observability tooling: Application Insights, Azure Monitor, Datadog, or similar Strong written and spoken English; daily communication with a US team and occasional client stakeholders Work schedule overlapping US Eastern hours (roughly 9 to 5 Eastern is ideal) Nice to Have Experience in HIPAA, PHI, or other regulated environments. A background check to healthcare-industry standards is required prior to production access Cosmos DB RU optimization and partition key design at high throughput Kubernetes experience Prior work in a multi-tenant SaaS environment ON-CALL EXPECTATIONS This role includes a pager-based on-call rotation covering SEV-1 and SEV-2 incidents, shared with the rest of the SRE team. On-call is a core part of the role. Expect it to be light in the first month and ramp as the team takes over production responsibility. COMPLIANCE AND ACCESS All personnel are named and approved by the client before any access is provisioned Background checks to healthcare-industry standards are completed prior to production access Production access is provisioned through the client identity provider with MFA and time-bound elevation BENEFITS Full-time employment Paid time off Supplemental health insurance Learning credits for training and certification Performance incentives Regular one on ones and ongoing career support Fully remote WHY THIS ROLE You are building the observability, on-call, and deployment practices, not inheriting someone else's Small senior team, direct access to senior engineers, no layers of process Interesting scale problems: multi-tenant data at volume and real performance challenges, in an environment that has to stay up while you harden it