Senior/Specialist SRE, Colombia

Hace 2 días

, Colombia CI&T Jornada completa $ 72 - $ 96 Por obra

We are tech transformation specialists, uniting human expertise with AI to create scalable tech solutions.

With over 8,000 CI&Ters around the world, we’ve built partnerships with more than 1,000 clients during our 30 years of history. Artificial Intelligence is our reality.

Your mission

The Observability & Monitoring Specialist is responsible for enhancing the visibility, performance, and health of SCT’s application landscape. This role focuses on bridging existing monitoring gaps across 100+ applications by leveraging modern observability solutions to ensure proactive issue detection and rapid incident resolution.

Working as a core part of the infrastructure team, this specialist will optimize our monitoring stack—specifically Splunk, LogicMonitor, and AppDynamics—to transform raw data into actionable insights, ensuring that development and operations teams have the telemetry needed to maintain high-availability systems.

Key Responsibilities

  • Advanced Observability & Monitoring Strategy: Identify and remediate visibility gaps across a landscape of 100+ applications, ensuring full‑stack monitoring from infrastructure to the end‑user experience. Design and implement modern monitoring patterns to move from reactive alerting to proactive anomaly detection.
  • Platform Management (Splunk, LogicMonitor, AppDynamics): Optimize log aggregation, create complex dashboards and advanced queries (SPL); manage system‑level monitoring; configure APM to track business transactions and identify code‑level bottlenecks.
  • Reliability Engineering & Performance Analysis: Correlate data across multiple platforms to provide a holistic view of system health and performance. Partner with application owners to define meaningful SLI and SLO.
  • Automation & Modernization: Automate the deployment and configuration of monitoring agents across Windows and Linux environments, advocate for modern observability solutions such as OpenTelemetry.
  • Deployment & Incident Response Support: Provide technical support during high‑priority incidents by leveraging AppDynamics and Splunk for rapid root‑cause analysis. Create and maintain operational dashboards and runbooks.
  • Continuous Improvement & Governance: Audit the existing monitoring environment to eliminate redundant alerts and ensure compliance with ITGC and security standards. Conduct knowledge‑sharing sessions.

Professional Expectations

  • Fluent English skills: to interact with multicultural team and American client daily.
  • Data‑driven mindset using metrics to influence decisions.
  • Strong collaboration skills, bridging infrastructure stability and application performance.
  • Proactive identification of trends in system behavior to prevent outages.

Benefits

  • Maternity and Parental leaves
  • Mobile services subsidy
  • Sick pay – Life insurance
  • CI&T University
  • Colombian Holidays
  • Paid Vacations
  • Many others

Collaboration is our superpower, diversity unites us, and excellence is our standard.

We value diverse identities and life experiences, fostering a diverse, inclusive, and safe work environment. We encourage applications from diverse and underrepresented groups to our job positions.

If you like it, just apply and good luck