About the Role
As a Senior Site Reliability Engineer you will make an impact by leading the operational excellence, reliability, and continuous improvement of AI-powered digital platforms supporting healthcare payer operations. You will oversee hybrid-cloud services leveraging Large Language Models (LLMs), MLOps practices, cloud-native technologies, and modern engineering frameworks to deliver secure, scalable, and compliant solutions that enhance member and provider experiences.
You will be a valued member of the Technology & Engineering team, collaborating closely with business stakeholders, product teams, platform engineers, AI specialists, and operations teams to ensure service stability, innovation, and regulatory compliance.
In This Role, You Will:
- Own end-to-end service accountability for payer-focused AI, LLM, and ML-enabled platforms, ensuring high availability, performance, and compliance across hybrid environments.
- Lead operational governance and continuous improvement initiatives aligned with enterprise service management best practices.
- Oversee MLOps processes supporting LLM-powered applications, including model deployment, monitoring, retraining, validation, and rollback strategies.
- Coordinate the deployment and lifecycle management of containerized services utilizing Kubernetes to deliver scalable, resilient, and highly available solutions.
- Implement and govern Infrastructure as Code (IaC) practices using Terraform to provision and manage cloud and on-premises resources consistently and securely.
- Standardize configuration management through Ansible automation to improve operational efficiency and reduce service disruptions.
- Support the development, deployment, and operational management of Node.js-based services and APIs that integrate AI and machine learning capabilities.
- Establish best practices for Git-based version control, release management, code reviews, and repository governance across application and infrastructure teams.
- Drive operational excellence across Linux environments, including security hardening, patch management, performance optimization, and system monitoring.
- Guide Python-based development supporting data pipelines, AI orchestration, automation frameworks, and analytics workloads.
- Collaborate with business stakeholders, product owners, and healthcare domain experts to translate complex payer requirements into reliable technology services.
- Lead incident management, root cause analysis, problem management, and service restoration activities to minimize business impact.
- Monitor platform health through metrics, logs, traces, and observability tools while continuously improving service reliability and resilience.
- Foster a culture of knowledge sharing, operational excellence, automation, and continuous improvement across distributed teams.
Work Model
We believe hybrid work is the way forward as we strive to provide flexibility wherever possible. Based on this role's business requirements, this is a hybrid position requiring attendance at a client or Cognizant office based on project needs.
The working arrangements for this role are accurate as of the date of posting and may change according to client and business requirements.
What You Need to Have to Be Considered
- Strong experience with Kubernetes and GCP (GKE)
- Strong experience in IaC (Terraform), Helm, and GitHub Actions
- Proficiency in Python, Ansible, Node.js
- Strong experience with Prometheus and Grafana observability stack
- Solid understanding of Linux systems and networking fundamentals
- Experience in incident management, on-call support, and production triage
- Hands-on experience with automation and CI/CD pipelines
- Strong understanding of AI/ML concepts and AIOps practices (model lifecycle, monitoring, or AI-driven alerting)
These Will Help You Stand Out
- Google Cloud Architect Certification
- Certified Kubernetes Administrator (CKA)
- Experience in Java/J2EE, Spring Boot
- Experience supporting or operating ML/AI platforms or pipelines (MLOps)
- Exposure to AIOps tools, anomaly detection, or predictive analytics systems
- Experience with large-scale distributed systems and microservices architecture
- Experience with GPU-based workloads or ML infrastructure on GCP
- Knowledge of Kubeflow, Vertex AI, or ML pipelines
- Experience integrating AI-driven automation into monitoring and incident response
Benefits:
Cognizant offers the following benefits for this position, subject to applicable eligibility requirements:
- Medical/Dental/Vision/Life Insurance.
- Paid holidays plus Paid Time Off.
- 401(k) plan and contributions.
- Long-term/Short-term Disability.
- Paid Parental Leave.
- Employee Stock Purchase Plan.
Why Cognizant?
At Cognizant, we're engineering modern businesses through innovation, technology, and human-centered solutions. You'll work with talented professionals, cutting-edge technologies, and industry-leading clients while helping shape the future of AI-driven healthcare solutions.
We're excited to meet people who share our mission and can make an impact in a variety of ways. Don't hesitate to apply, even if you don't meet every requirement. We value diverse experiences, transferable skills, and a passion for innovation.
Acerca de Cognizant
Cognizant (Nasdaq: CTSH) es un creador de soluciones de IA y proveedor de servicios tecnológicos que conecta la inversión en IA con el valor empresarial mediante el desarrollo de soluciones de IA full‑stack para sus clientes. Su profundo conocimiento de la industria, junto con su experiencia en procesos e ingeniería, permite incorporar el contexto único de cada organización en sistemas tecnológicos que amplifican el potencial humano, generan resultados tangibles y mantienen a las empresas a la vanguardia en un entorno en constante cambio. Más información en cognizant.ai o @cognizant.
Información adicional sobre el empleo
La información sobre la compensación es exacta en la fecha de publicación de este anuncio. Cognizant se reserva el derecho de modificar esta información en cualquier momento, de conformidad con la legislación aplicable.
Es posible que se solicite a los solicitantes que asistan a entrevistas de forma presencial o mediante videoconferencia. Asimismo, durante cada entrevista, los candidatos podrán estar obligados a presentar un documento de identidad válido emitido por el estado o por el gobierno.
Cognizant es un empleador que ofrece igualdad de oportunidades. Su solicitud y candidatura no se evaluarán en función de la raza, el color, el sexo, la religión, el credo, la orientación sexual, la identidad de género, el origen nacional, la discapacidad, la información genética, el embarazo, la condición de veterano ni cualquier otra característica protegida por las leyes federales, estatales o locales.







