Role Overview
We are seeking seasoned professionals with deep expertise in operating and managing High-Performance Computing (HPC) platforms. The ideal candidate will have hands-on experience in designing, deploying, and maintaining HPC clusters, storage systems, and networking infrastructure, leveraging industry-leading tools and technologies.
Key Responsibilities
· HPC Infrastructure Management
o Operate and maintain HPC clusters based on CentOS, RHEL, and hardware platforms like HPE and NVIDIA DGX.
o Ensure optimal performance, scalability, and reliability of compute resources.
· Storage Administration
o Manage large-scale storage systems including Dell Isilon, VAST Storage, Lustre, and GPFS.
o Implement data lifecycle management and optimize storage performance for HPC workloads.
· Networking
o Configure and maintain InfiniBand-based networking for low-latency, high-bandwidth communication.
o Troubleshoot network performance issues and ensure secure connectivity.
· Cluster and Job Scheduling
o Administer cluster management tools such as Bright Cluster Manager, Altair Grid Manager, and IBM LSF.
o Optimize job scheduling and resource allocation for diverse workloads.
· Monitoring and Automation
o Implement monitoring solutions using Zabbix, Grafana, and ELK Stack.
o Automate provisioning and configuration using Cobbler, Chef, Ansible, and AWS ParallelCluster.
· Performance Tuning & Troubleshooting
o Conduct performance benchmarking and tuning for HPC workloads.
o Diagnose and resolve hardware/software issues across compute, storage, and network layers.
· Security & Compliance
o Ensure HPC environment adheres to security best practices and compliance standards.
Required Skills & Qualifications
· Technical Expertise
o Strong knowledge of Linux OS (CentOS, RHEL) and HPC hardware platforms (HPE, NVIDIA DGX).
o Hands-on experience with parallel file systems (Lustre, GPFS) and enterprise storage solutions.
o Proficiency in InfiniBand networking and high-speed interconnects.
o Familiarity with job schedulers and cluster management tools (IBM LSF, Bright Cluster Manager, Altair Grid Manager).
· Automation & Scripting
o Expertise in Ansible, Chef, Cobbler, and scripting languages (Bash, Python).
o Experience with AWS ParallelCluster or similar cloud-based HPC solutions.
· Monitoring & Logging
o Practical experience with Zabbix, Grafana, and ELK Stack for system health and performance monitoring.
· Soft Skills
o Strong problem-solving and analytical skills.
o Ability to work in a fast-paced environment and lead technical teams.
o Excellent communication and documentation skills.
Preferred Qualifications
· Exposure to AI/ML workloads on HPC clusters.
· Experience with containerization (Docker, Singularity) in HPC environments.
· Knowledge of security hardening for HPC systems.
Education
· Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.
#LI-LK1
コグニザントについて
コグニザント(NASDAQ: CTSH)は、AI Builderおよびテクノロジーサービスプロバイダーとして、お客様にフルスタックのAIソリューションを構築することで、AI投資と企業価値を結ぶ架け橋となっています。業界、ビジネスプロセス、エンジニアリングに関する当社の深い専門知識を活かし、組織固有のビジネス環境をテクノロジー・システムに組み込みます。これにより、人間の可能性を最大限に引き出し、確かな成果を実現するとともに、急速に変化する世界においてグローバル企業が常に一歩先を行くための支援を行っています。 詳細については、cognizant.ai をご覧ください。
雇用に関する追加情報
本募集に記載されている報酬情報は、掲載日時点で正確なものです。Cognizantは、適用される法令に従い、いつでも本情報を変更する権利を留保します。
応募者は、対面またはビデオ会議による面接への参加を求められる場合があります。また、各面接の際に、現在有効な州政府または政府発行の身分証明書の提示を求められる場合があります。
Cognizantは機会均等雇用主です。応募および選考において、人種、肌の色、性別、宗教、信条、性的指向、性自認、国籍、障がい、遺伝情報、妊娠、退役軍人の地位、その他連邦法・州法・地方自治体の法律により保護されるいかなる特性に基づく差別も行いません。







