跳到主要內容

Site Reliability Engineer (SRE) EMS Production & Observability

00070847661

Site Reliability Engineer (SRE)

EMS Production & Observability

Job Description

Role Overview

We are looking for an experienced Site Reliability Engineer (SRE) to support the reliability, availability, performance, and operational readiness of EMS production services.

The engineer will be responsible for production support and incident resolution, enhancing ELF, building actionable operational dashboards, improving monitoring and alerting, automating repetitive operational activities, and proactively preparing the platform for high-volume/high-visibility events such as the US Open 2026.

This role requires someone who can move beyond reactive production support and use observability, automation, capacity planning, trend analysis, and engineering improvements to identify and address reliability risks before they become customer-impacting incidents.

Key Responsibilities

Production Reliability & Incident Management

  • Own and support EMS production incidents, including investigation, triage, mitigation, recovery, and root-cause analysis.

  • Participate in production support/on-call rotations and provide timely response to critical incidents.

  • Troubleshoot application, infrastructure, API, integration, latency, capacity, and availability issues across the EMS ecosystem.

  • Drive incidents through resolution by collaborating with application engineering, infrastructure, platform, networking, database, and dependent service teams.

  • Conduct Root Cause Analysis (RCA) for significant incidents and ensure corrective/preventive actions are implemented.

  • Identify recurring production problems and convert operational fixes into permanent engineering solutions.

Observability, Monitoring & Dashboards

  • Design and maintain real-time operational dashboards providing visibility into EMS health and customer-impacting conditions.

  • Define and monitor key SLIs/SLOs around availability, latency, throughput, errors, and service dependencies.

  • Build meaningful alerting based on customer/service impact rather than relying solely on infrastructure thresholds.

  • Correlate application logs, metrics, events, and infrastructure telemetry to accelerate issue detection and troubleshooting.

  • Develop dashboards for transaction/request volume, success and failure rates, response time and latency, API/service availability, infrastructure and application health, dependency health, error trends and top failure reasons, capacity/utilization, and event-specific traffic and performance.

  • Continuously tune alerts to reduce noise and improve signal-to-noise ratio and Mean Time to Detect (MTTD).

ELF Enhancement & Automation

  • Enhance ELF capabilities to improve production monitoring, operational efficiency, troubleshooting, and reliability.

  • Identify manual production-support activities and automate them using scripting, APIs, CI/CD, or appropriate platform capabilities.

  • Build automated health checks, diagnostics, remediation, and operational reporting.

  • Develop reusable operational tools and runbooks that enable faster diagnosis and recovery.

  • Work with engineering teams to incorporate reliability and observability requirements into new releases.

Major Event Readiness

  • Lead proactive reliability planning for major business events such as US Open 2026 and other anticipated high-volume periods.

  • Analyze historical and projected traffic patterns to identify potential capacity or performance bottlenecks.

  • Establish event-specific dashboards, alerts, and operational thresholds.

  • Perform capacity planning, load/performance validation, and dependency readiness assessments.

  • Identify critical failure scenarios and develop mitigation and recovery plans before the event.

  • Conduct pre-event readiness reviews, failure simulations/game days, and operational rehearsals.

  • Establish clear event-day monitoring, escalation paths, communication channels, and incident-response procedures.

  • Produce post-event analysis covering performance, incidents, trends, and opportunities for improvement.

Reliability Engineering

The SRE will help establish and continuously improve engineering practices including SLIs, SLOs and error budgets; availability and latency measurement; capacity and performance management; automated remediation; incident management and postmortems; operational readiness reviews; dependency monitoring; resilience testing; and disaster/recovery preparedness.

A key expectation is to use production data to answer questions such as: What can fail? How will we detect it? How quickly can we isolate it? What is the customer impact? Can we automatically recover? How do we prevent it from recurring?

Required Qualifications

  • Strong experience in Site Reliability Engineering, Production Engineering, DevOps, or Production Support for enterprise-scale applications.

  • Strong hands-on production troubleshooting and incident-management experience.

  • Experience building production monitoring, alerting, and observability dashboards.

  • Strong understanding of application and infrastructure metrics, logs, distributed services, and API monitoring.

  • Experience with cloud infrastructure and containerized environments.

  • Knowledge of Kubernetes, CI/CD, Infrastructure as Code, and automation.

  • Scripting/programming experience with Python, Bash, PowerShell, or equivalent.

  • Understanding of networking concepts including DNS, load balancing, firewalls, routing, and connectivity troubleshooting.

  • Experience performing RCA and driving preventive engineering actions.

  • Ability to analyze traffic, latency, errors, and capacity trends and translate findings into engineering improvements.

  • Strong communication skills and ability to coordinate effectively during high-severity production incidents.

Preferred Qualifications

Experience with Azure, AKS/Kubernetes, Terraform, GitHub Actions, Grafana/Prometheus, Splunk or equivalent observability platforms, Application Performance Monitoring (APM), synthetic monitoring, automated remediation, chaos/game-day testing, and high-volume event readiness would be valuable.

Experience supporting highly available, customer-facing platforms where reliability during major events and peak traffic periods is critical is strongly preferred.


关于高知特 (Cognizant)
高知特(Cognizant)(纳斯达克代码:CTSH)作为一家AI Builder和相关技术服务提供商,致力于通过打造全栈AI解决方案,帮助企业将人工智能投资转化为实际价值。公司凭借深厚的行业经验、流程优化和工程技术专长,将企业独特的业务场景融入科技系统,赋能组织释放人才潜能,推动切实成果,并帮助全球企业在瞬息万变的环境中保持领先。如需了解更多详情,敬请访问 cognizant.ai 或关注@cognizant。

补充雇佣信息
薪酬信息截至本职位发布之日为准。Cognizant 保留在适用法律允许的范围内随时修改该信息的权利。
申请人可能需要通过现场面试或视频会议的方式参加面试。此外,候选人在每次面试时可能需要出示其当前所在州或政府签发的有效身份证件。
Cognizant 是一家提供平等就业机会的雇主。在招聘过程中,您的申请和候选资格不会因种族、肤色、性别、宗教、信仰、性取向、性别认同、国籍、残疾、遗传信息、怀孕、退伍军人身份或任何其他受联邦、州或地方法律保护的特征而受到影响。

帮助您蓬勃发展与成长的福利

我们的福利计划以您为中心打造——帮助您享受充实、平衡且健康的生活。
有葉子的植物的藍色線條圖

财务健康

我们会定期审查市场数据,确保薪酬体现您所带来的价值。您的福利不仅限于薪资,还可能包括退休计划、财务教育等。

Stay Healthy Midnight Blue RGB

身心健康

我们通过带薪休假、在条件允许下的灵活工作安排、医疗保障计划、心理咨询、心理健康盟友计划等,赋能您将身心健康放在首位。

Build The Career You Want Midnight Blue RGB

您的职业发展,由您做主

在 Cognizant 提供的 35 万多个岗位中,您将有机会探索新的技术、行业和工作地点,并打造推动职业发展的关键技能。

Making A Meaningful Impact Midnight Blue RGB

现实世界的影响力

想想您所依赖的那些知名品牌。很可能,他们也依赖我们来帮助强化其业务。在这里,您将把大胆的想法转化为改善全球生活的解决方案。

还没有找到合适的机会吗?

获取为您量身定制的最新职位机会、招聘活动和公司新闻!

掌握最新动态