Role Overview
We are seeking seasoned professionals with deep expertise in operating and managing High-Performance Computing (HPC) platforms. The ideal candidate will have hands-on experience in designing, deploying, and maintaining HPC clusters, storage systems, and networking infrastructure, leveraging industry-leading tools and technologies.
Key Responsibilities
· HPC Infrastructure Management
o Operate and maintain HPC clusters based on CentOS, RHEL, and hardware platforms like HPE and NVIDIA DGX.
o Ensure optimal performance, scalability, and reliability of compute resources.
· Storage Administration
o Manage large-scale storage systems including Dell Isilon, VAST Storage, Lustre, and GPFS.
o Implement data lifecycle management and optimize storage performance for HPC workloads.
· Networking
o Configure and maintain InfiniBand-based networking for low-latency, high-bandwidth communication.
o Troubleshoot network performance issues and ensure secure connectivity.
· Cluster and Job Scheduling
o Administer cluster management tools such as Bright Cluster Manager, Altair Grid Manager, and IBM LSF.
o Optimize job scheduling and resource allocation for diverse workloads.
· Monitoring and Automation
o Implement monitoring solutions using Zabbix, Grafana, and ELK Stack.
o Automate provisioning and configuration using Cobbler, Chef, Ansible, and AWS ParallelCluster.
· Performance Tuning & Troubleshooting
o Conduct performance benchmarking and tuning for HPC workloads.
o Diagnose and resolve hardware/software issues across compute, storage, and network layers.
· Security & Compliance
o Ensure HPC environment adheres to security best practices and compliance standards.
Required Skills & Qualifications
· Technical Expertise
o Strong knowledge of Linux OS (CentOS, RHEL) and HPC hardware platforms (HPE, NVIDIA DGX).
o Hands-on experience with parallel file systems (Lustre, GPFS) and enterprise storage solutions.
o Proficiency in InfiniBand networking and high-speed interconnects.
o Familiarity with job schedulers and cluster management tools (IBM LSF, Bright Cluster Manager, Altair Grid Manager).
· Automation & Scripting
o Expertise in Ansible, Chef, Cobbler, and scripting languages (Bash, Python).
o Experience with AWS ParallelCluster or similar cloud-based HPC solutions.
· Monitoring & Logging
o Practical experience with Zabbix, Grafana, and ELK Stack for system health and performance monitoring.
· Soft Skills
o Strong problem-solving and analytical skills.
o Ability to work in a fast-paced environment and lead technical teams.
o Excellent communication and documentation skills.
Preferred Qualifications
· Exposure to AI/ML workloads on HPC clusters.
· Experience with containerization (Docker, Singularity) in HPC environments.
· Knowledge of security hardening for HPC systems.
Education
· Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.
#LI-LK1
关于高知特 (Cognizant)
高知特(Cognizant)(纳斯达克代码:CTSH)作为一家AI Builder和相关技术服务提供商,致力于通过打造全栈AI解决方案,帮助企业将人工智能投资转化为实际价值。公司凭借深厚的行业经验、流程优化和工程技术专长,将企业独特的业务场景融入科技系统,赋能组织释放人才潜能,推动切实成果,并帮助全球企业在瞬息万变的环境中保持领先。如需了解更多详情,敬请访问 cognizant.ai 或关注@cognizant。
补充雇佣信息
薪酬信息截至本职位发布之日为准。Cognizant 保留在适用法律允许的范围内随时修改该信息的权利。
申请人可能需要通过现场面试或视频会议的方式参加面试。此外,候选人在每次面试时可能需要出示其当前所在州或政府签发的有效身份证件。
Cognizant 是一家提供平等就业机会的雇主。在招聘过程中,您的申请和候选资格不会因种族、肤色、性别、宗教、信仰、性取向、性别认同、国籍、残疾、遗传信息、怀孕、退伍军人身份或任何其他受联邦、州或地方法律保护的特征而受到影响。







