Operations Support Engineer
Role Overview
We are seeking an experienced Operations Support Engineer to provide L1, L2, and L3 application support for internal enterprise platforms and digital products. This role is responsible for ensuring the reliability, stability, security, and performance of production and non-production environments while supporting ongoing business and technology operations.
The successful candidate will work closely with development teams, product managers, architects, security teams, business users, and external vendors to identify and resolve issues, implement operational improvements, and maintain high service availability. This position plays a critical role in managing incidents, optimizing system performance, and supporting cloud-based and AI-driven solutions.
Key Responsibilities
- Monitor and analyze application and infrastructure performance across production and non-production environments to ensure optimal system availability and efficiency.
- Develop data-driven strategies and implement continuous improvement initiatives in collaboration with application teams, solution architects, security teams, and other stakeholders.
- Partner with business users to establish and support cloud-based environments for AI and analytics use cases.
- Manage application, infrastructure, and security incidents; perform root cause analysis and coordinate resolution efforts with internal teams and external vendors.
- Ensure incidents and service requests are resolved within agreed service levels and provide timely escalation when required.
- Develop and maintain operational procedures, support documentation, system configurations, and audit-compliant process guides.
- Prepare operational reports, analyze performance metrics, and communicate findings to stakeholders and management.
- Manage day-to-day operational activities and ensure business continuity.
- Coordinate and oversee support teams, including external vendors, to provide 24x7 operational support coverage.
- Assist in implementing operational best practices, automation, monitoring, and security controls across systems and environments.
Required Skills & Experience
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline.
- Proven experience as an Operations Engineer, Production Support Engineer, Site Reliability Engineer (SRE), or a similar IT operations role.
- Experience implementing incident management, problem management, and change management processes using ITSM platforms.
- Strong understanding of security controls, privileged access management, and environment access governance.
- Experience implementing application and infrastructure monitoring solutions, including Application Performance Monitoring (APM) tools.
- Familiarity with cloud-native monitoring and logging services.
- Experience identifying and implementing automation solutions to reduce manual effort, operational risk, and system downtime.
- Knowledge of infrastructure automation and scripting tools such as Terraform, Ansible, or similar technologies.
- Hands-on experience with Agile methodologies, DevOps practices, CI/CD pipelines, test-driven development, and information security standards.
- Strong analytical and troubleshooting skills with the ability to resolve complex technical issues.
- Excellent communication and stakeholder management skills with the ability to translate technical concepts for non-technical audiences.
- Ability to work effectively within high-performing, cross-functional teams.
Preferred Qualifications
- Experience managing and supporting cloud infrastructure and services on platforms such as AWS, Microsoft Azure, or Google Cloud Platform.
- Relevant cloud certifications are highly desirable.
- Experience working with AI/ML platforms and cloud-based AI solutions.
- Understanding of Retrieval-Augmented Generation (RAG), workflow automation, tool integration, and data lake architectures is advantageous.
- Experience supporting enterprise-scale applications in a 24x7 operational environment.
Key Competencies
- Problem-solving and critical thinking
- Operational excellence mindset
- Continuous improvement orientation
- Strong ownership and accountability
- Collaboration and stakeholder engagement
- Innovation and adaptability
- Customer-focused approach
关于高知特 (Cognizant)
高知特(Cognizant)(纳斯达克代码:CTSH)作为一家AI Builder和相关技术服务提供商,致力于通过打造全栈AI解决方案,帮助企业将人工智能投资转化为实际价值。公司凭借深厚的行业经验、流程优化和工程技术专长,将企业独特的业务场景融入科技系统,赋能组织释放人才潜能,推动切实成果,并帮助全球企业在瞬息万变的环境中保持领先。如需了解更多详情,敬请访问 cognizant.ai 或关注@cognizant。
补充雇佣信息
薪酬信息截至本职位发布之日为准。Cognizant 保留在适用法律允许的范围内随时修改该信息的权利。
申请人可能需要通过现场面试或视频会议的方式参加面试。此外,候选人在每次面试时可能需要出示其当前所在州或政府签发的有效身份证件。
Cognizant 是一家提供平等就业机会的雇主。在招聘过程中,您的申请和候选资格不会因种族、肤色、性别、宗教、信仰、性取向、性别认同、国籍、残疾、遗传信息、怀孕、退伍军人身份或任何其他受联邦、州或地方法律保护的特征而受到影响。







