Direkt zum Inhalt

Site Reliability Engineer (SRE) EMS Production & Observability

00070847661

Site Reliability Engineer (SRE)

EMS Production & Observability

Job Description

Role Overview

We are looking for an experienced Site Reliability Engineer (SRE) to support the reliability, availability, performance, and operational readiness of EMS production services.

The engineer will be responsible for production support and incident resolution, enhancing ELF, building actionable operational dashboards, improving monitoring and alerting, automating repetitive operational activities, and proactively preparing the platform for high-volume/high-visibility events such as the US Open 2026.

This role requires someone who can move beyond reactive production support and use observability, automation, capacity planning, trend analysis, and engineering improvements to identify and address reliability risks before they become customer-impacting incidents.

Key Responsibilities

Production Reliability & Incident Management

  • Own and support EMS production incidents, including investigation, triage, mitigation, recovery, and root-cause analysis.

  • Participate in production support/on-call rotations and provide timely response to critical incidents.

  • Troubleshoot application, infrastructure, API, integration, latency, capacity, and availability issues across the EMS ecosystem.

  • Drive incidents through resolution by collaborating with application engineering, infrastructure, platform, networking, database, and dependent service teams.

  • Conduct Root Cause Analysis (RCA) for significant incidents and ensure corrective/preventive actions are implemented.

  • Identify recurring production problems and convert operational fixes into permanent engineering solutions.

Observability, Monitoring & Dashboards

  • Design and maintain real-time operational dashboards providing visibility into EMS health and customer-impacting conditions.

  • Define and monitor key SLIs/SLOs around availability, latency, throughput, errors, and service dependencies.

  • Build meaningful alerting based on customer/service impact rather than relying solely on infrastructure thresholds.

  • Correlate application logs, metrics, events, and infrastructure telemetry to accelerate issue detection and troubleshooting.

  • Develop dashboards for transaction/request volume, success and failure rates, response time and latency, API/service availability, infrastructure and application health, dependency health, error trends and top failure reasons, capacity/utilization, and event-specific traffic and performance.

  • Continuously tune alerts to reduce noise and improve signal-to-noise ratio and Mean Time to Detect (MTTD).

ELF Enhancement & Automation

  • Enhance ELF capabilities to improve production monitoring, operational efficiency, troubleshooting, and reliability.

  • Identify manual production-support activities and automate them using scripting, APIs, CI/CD, or appropriate platform capabilities.

  • Build automated health checks, diagnostics, remediation, and operational reporting.

  • Develop reusable operational tools and runbooks that enable faster diagnosis and recovery.

  • Work with engineering teams to incorporate reliability and observability requirements into new releases.

Major Event Readiness

  • Lead proactive reliability planning for major business events such as US Open 2026 and other anticipated high-volume periods.

  • Analyze historical and projected traffic patterns to identify potential capacity or performance bottlenecks.

  • Establish event-specific dashboards, alerts, and operational thresholds.

  • Perform capacity planning, load/performance validation, and dependency readiness assessments.

  • Identify critical failure scenarios and develop mitigation and recovery plans before the event.

  • Conduct pre-event readiness reviews, failure simulations/game days, and operational rehearsals.

  • Establish clear event-day monitoring, escalation paths, communication channels, and incident-response procedures.

  • Produce post-event analysis covering performance, incidents, trends, and opportunities for improvement.

Reliability Engineering

The SRE will help establish and continuously improve engineering practices including SLIs, SLOs and error budgets; availability and latency measurement; capacity and performance management; automated remediation; incident management and postmortems; operational readiness reviews; dependency monitoring; resilience testing; and disaster/recovery preparedness.

A key expectation is to use production data to answer questions such as: What can fail? How will we detect it? How quickly can we isolate it? What is the customer impact? Can we automatically recover? How do we prevent it from recurring?

Required Qualifications

  • Strong experience in Site Reliability Engineering, Production Engineering, DevOps, or Production Support for enterprise-scale applications.

  • Strong hands-on production troubleshooting and incident-management experience.

  • Experience building production monitoring, alerting, and observability dashboards.

  • Strong understanding of application and infrastructure metrics, logs, distributed services, and API monitoring.

  • Experience with cloud infrastructure and containerized environments.

  • Knowledge of Kubernetes, CI/CD, Infrastructure as Code, and automation.

  • Scripting/programming experience with Python, Bash, PowerShell, or equivalent.

  • Understanding of networking concepts including DNS, load balancing, firewalls, routing, and connectivity troubleshooting.

  • Experience performing RCA and driving preventive engineering actions.

  • Ability to analyze traffic, latency, errors, and capacity trends and translate findings into engineering improvements.

  • Strong communication skills and ability to coordinate effectively during high-severity production incidents.

Preferred Qualifications

Experience with Azure, AKS/Kubernetes, Terraform, GitHub Actions, Grafana/Prometheus, Splunk or equivalent observability platforms, Application Performance Monitoring (APM), synthetic monitoring, automated remediation, chaos/game-day testing, and high-volume event readiness would be valuable.

Experience supporting highly available, customer-facing platforms where reliability during major events and peak traffic periods is critical is strongly preferred.


Über Cognizant  
Cognizant (NASDAQ: CTSH) i ist ein Technologiedienstleister und Entwickler von KI-Lösungen. Wir schlagen die Brücke zwischen KI-Investitionen und echtem unternehmerischem Mehrwert, indem wir ganzheitliche Full-Stack-KI-Lösungen für unsere Kunden entwickeln. Mit unserer fundierten Branchen-, Prozess- und Engineering-Expertise integrieren wir die spezifischen Anforderungen von Unternehmen passgenau in Technologiesysteme. So entfalten wir das menschliche Potenzial, erzielen greifbare Ergebnisse und sichern globalen Unternehmen in einer sich rasant wandelnden Welt den entscheidenden Vorsprung. Erfahren Sie mehr unter cognizant.ai oder @cognizant.

Zusätzliche Informationen zur Beschäftigung
Die Vergütungsinformationen sind zum Zeitpunkt der Veröffentlichung dieser Stellenausschreibung korrekt. Cognizant behält sich das Recht vor, diese Informationen jederzeit unter Beachtung der geltenden gesetzlichen Bestimmungen zu ändern.

Bewerberinnen und Bewerber können verpflichtet sein, an Vorstellungsgesprächen persönlich oder per Videokonferenz teilzunehmen. Darüber hinaus kann es erforderlich sein, bei jedem Gespräch einen gültigen staatlichen Lichtbildausweis vorzulegen.

Cognizant ist ein Arbeitgeber mit Chancengleichheit. Ihre Bewerbung und Kandidatur werden nicht aufgrund von Rasse, Hautfarbe, Geschlecht, Religion, Glaubensbekenntnis, sexueller Orientierung, Geschlechtsidentität, nationaler Herkunft, Behinderung, genetischen Informationen, Schwangerschaft, Veteranenstatus oder sonstiger durch bundes‑, landes‑ oder kommunalrechtliche Vorschriften geschützter Merkmale berücksichtigt oder abgelehnt.

Leistungen, die Ihnen helfen, sich zu entfalten und weiterzuentwickeln

Unser Benefits‑Programm ist auf Sie zugeschnitten, damit Sie ein erfülltes, ausgewogenes und gesundes Leben führen können.

eine blaue Strichzeichnung einer Pflanze mit Blättern

Finanzielle Absicherung

Wir überprüfen regelmäßig Marktdaten, um sicherzustellen, dass die Vergütung den Wert widerspiegelt, den Sie einbringen. Unsere Benefits gehen über das Gehalt hinaus und können unter anderem betriebliche Altersvorsorge, finanzielle Weiterbildung und mehr umfassen.

Stay Healthy Midnight Blue RGB

Körperliche und mentale Gesundheit

Wir ermöglichen Ihnen, Ihr Wohlbefinden in den Mittelpunkt zu stellen – durch bezahlte Auszeiten, flexible Arbeitsmodelle, wo möglich, Gesundheitsleistungen, Beratung, unser Mental‑Health‑Allyship‑Programm und vieles mehr.

Build The Career You Want Midnight Blue RGB

Ihre Karriere, Ihr Weg

Mit über 350.000 Positionen bei Cognizant haben Sie die Möglichkeit, neue Technologien, Branchen und Standorte kennenzulernen – und die Fähigkeiten aufzubauen, die Sie für die Weiterentwicklung Ihrer Karriere benötigen.

Making A Meaningful Impact Midnight Blue RGB

Wirkung in der realen Welt

Denken Sie an die größten Marken, auf die Sie sich verlassen. Wahrscheinlich verlassen sie sich auf uns, um ihr Geschäft zu stärken. Hier setzen Sie mutige Ideen in Lösungen um, die das Leben von Menschen überall verbessern.

Noch nicht die passende Stelle gefunden?

Erhalten Sie die neuesten Updates zu Stellenangeboten, Recruiting‑Events und Unternehmensnews – individuell auf Sie zugeschnitten!

Bleiben Sie informiert