Lead SRE Architect
Location : Remote
*Please note, this role is not able to offer visa transfer or sponsorship now or in the future*
Job Description –
Role Summary
We are seeking an experienced Lead Site Reliability Engineering (SRE) Architect to improve the reliability, stability, and operational excellence of a large-scale hybrid infrastructure environment. The architect will assess the current infrastructure landscape, analyze recurring incidents, identify systemic weaknesses, and define strategic improvements to reduce outages and enhance platform resilience.
The role requires a strong infrastructure architect with expertise across compute, networking, databases, middleware, Unix/Linux, virtualization, cloud platforms(Azure), and SRE practices.
Key Responsibilities
· Infrastructure Assessment & Architecture
· Understand the existing enterprise infrastructure landscape across Network, Compute, Database, Middleware, Unix/Linux, Virtualization (Nutanix), and Azure Cloud.
· Perform comprehensive infrastructure health assessments and gap analysis.
· Identify architectural bottlenecks, technical debt, and operational risks.
· Develop a strategic roadmap to improve infrastructure reliability, scalability, availability, and operational efficiency.
Site Reliability Engineering
· Review historical P1/P2 incidents, outages, and recurring operational issues.
· Conduct deep root cause analysis to identify systemic failure patterns.
· Recommend architectural and operational improvements to prevent repeat incidents.
· Define reliability engineering best practices across infrastructure domains.
· Drive implementation of proactive monitoring, observability, automation, and self-healing capabilities.
· Reliability & Operational Excellence
· Define Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability metrics.
· Recommend improvements in monitoring, alerting, capacity management, backup, disaster recovery, and high availability.
· Improve incident response, problem management, and change management processes.
· Establish governance for infrastructure reliability and operational maturity.
· Stakeholder Management
· Work closely with Infrastructure Delivery, Operations, Engineering, and Customer stakeholders.
· Present findings, recommendations, and transformation roadmaps to leadership.
· Provide architectural guidance during major incidents and infrastructure modernization initiatives.
Required Skills
· 15+ years of enterprise infrastructure experience.
· Strong knowledge across:
· Network
· Compute
· Unix/Linux
· Windows Infrastructure
· Database Platforms
· Middleware
· Storage
· Strong expertise in Nutanix.
· Good understanding of Microsoft Azure in hybrid environments.
· Experience with enterprise monitoring and observability platforms.
· Strong knowledge of ITIL Incident, Problem, and Change Management.
· Experience performing RCA for major production incidents.
· Understanding of High Availability, Disaster Recovery, Capacity Planning, and Performance Engineering.
· Excellent stakeholder management and communication skills.
Preferred Skills
· Experience in large enterprise utility or critical infrastructure environments.
· Knowledge of DevOps, Infrastructure as Code, and automation tools.
· Experience with SRE practices, error budgets, reliability metrics, and operational excellence frameworks.
· Certifications in Azure, Nutanix, ITIL, or enterprise architecture are desirable.
· Expected Deliverables
· Current-state infrastructure assessment.
· Enterprise infrastructure gap analysis.
· Reliability maturity assessment.
· RCA review of historical critical incidents.
· Prioritized recommendations to eliminate recurring P1/P2 incidents.
· Infrastructure modernization roadmap.
· Reliability improvement roadmap.
· Executive-level architecture recommendations.
· Knowledge transfer and governance framework.
Salary and Other Compensation:
The Salary for this position will be 130,000 to 150,000 depending on experience and other qualifications of the successful candidate.
This position is also eligible for Cognizant’s discretionary annual incentive program, based on performance and subject to the terms of Cognizant’s applicable plans.
Benefits: Cognizant offers the following benefits for this position, subject to applicable eligibility requirements:
- Medical/Dental/Vision/Life Insurance
- Paid holidays plus Paid Time Off
- 401(k) plan and contributions
- Long-term/Short-term Disability
- Paid Parental Leave
- Employee Stock Purchase Plan
Disclaimer: The compensation, and benefits information is accurate as of the date of this posting. Cognizant reserves the right to modify this information at any time, subject to applicable law.
About Cognizant:
Cognizant (Nasdaq: CTSH) is an AI Builder and technology services provider, bridging the gap between AI investment and enterprise value by building full-stack AI solutions for our clients. Our deep industry, process and engineering expertise enables us to build an organization’s unique context into technology systems that amplify human potential, drive tangible outcomes and keep global enterprises ahead in a fast-changing world. See how at cognizant.ai or @cognizant.
Additional employment information
Compensation information is accurate as of the date of this posting. Cognizant reserves the right to modify this information at any time, subject to applicable law.
Applicants may be required to attend interviews in person or by video conference. In addition, candidates may be required to present their current state or government issued ID during each interview.
Cognizant is an equal opportunity employer. Your application and candidacy will not be considered based on race, color, sex, religion, creed, sexual orientation, gender identity, national origin, disability, genetic information, pregnancy, veteran status or any other characteristic protected by federal, state or local laws.












