Job Summary
Design and optimize large scale data processing solutions as a senior developer using Spark and AWS based services in a hybrid work arrangement. Apply advanced expertise in Spark optimization AWS Glue EMR Redshift S3 IAM and CloudWatch to build reliable data pipelines and analytical platforms. Collaborate with cross functional teams in day shifts to deliver secure efficient and business aligned data products that support strategic decision making.
Responsibilities
Design and implement scalable data processing pipelines using Apache Spark and AWS Glue to efficiently transform and prepare data for analytics and reporting across the organization.Optimize Spark jobs by tuning configurations improving partition strategies and refining data formats to consistently reduce runtime control resource consumption and improve overall reliability of data workflows.Build and manage AWS EMR clusters to run complex distributed data processing workloads while maintaining cost efficiency and ensuring high performance for critical business processes.Develop robust data ingestion and storage solutions using Amazon S3 as the central data lake enforcing folder structures and retention practices that support governance reusability and efficient querying.Create and maintain analytical data models in Amazon Redshift that enable fast and accurate reporting ensuring proper distribution keys sort keys and compression settings are applied for query performance optimization.Configure and maintain AWS CloudWatch monitoring for data pipelines and compute resources setting up detailed metrics and alarms that enable proactive detection and resolution of operational issues before they impact business stakeholders.Administer AWS IAM policies for data engineering components to ensure secure access control applying fine grained permissions that protect sensitive data while allowing teams to work efficiently within approved boundaries.Collaborate with product owners analysts and engineering peers during hybrid work days to gather requirements refine technical designs and deliver high quality data solutions that directly support strategic business initiatives.Review and refactor existing data processing code to improve readability performance and maintainability mentoring less experienced developers by sharing best practices in Spark optimization and cloud native development.Document data flows transformation logic and operational procedures in clear technical references that enable future enhancements onboarding of new team members and transparent communication with audit and compliance functions.Troubleshoot complex production incidents in day shift schedules by analyzing logs metrics and data anomalies applying structured problem solving to restore service quickly and prevent recurrence through targeted improvements.Contribute to continuous improvement of the data platform by evaluating new AWS features and data engineering techniques recommending adoption paths that increase reliability scalability and positive impact on business and society.Align daily development activities with the company purpose by building data solutions that support ethical decision making enable better customer experiences and improve operational efficiency in a sustainable manner.
Qualifications
Possess around ten to twelve years of professional experience as a data focused developer working on large scale distributed systems with a strong track record of delivering stable and performant solutions.Demonstrate deep hands on expertise in Apache Spark optimization including experience with partitioning strategies caching approaches broadcast joins and efficient use of data formats such as parquet.Show advanced practical knowledge of AWS Glue for building serverless data integration jobs including authoring scripts configuring crawlers and managing glue data catalogs for structured and semi structured data.Apply strong working experience with AWS EMR clusters for running Spark workloads including cluster sizing auto scaling configurations and integration with S3 based data lakes.Exhibit solid experience with Amazon Redshift for designing analytical schemas tuning queries and managing workload patterns to serve reporting and dashboarding needs quickly and reliably.Use Amazon S3 in depth for organizing data repositories applying lifecycle policies and supporting data governance and access patterns suitable for large enterprise environments.Implement secure and compliant solutions with AWS IAM by defining roles and policies that restrict access appropriately for data compute and monitoring resources in a complex cloud environment.Employ AWS CloudWatch proficiently for logging metrics and alerts creating actionable dashboards and notifications that help sustain high availability and timely response to anomalies.Communicate effectively with cross functional teams in a hybrid work setting demonstrating clear verbal and written skills that support collaborative design estimation and delivery of data solutions.Adapt quickly to evolving business requirements and technology advancements in the data engineering space showing curiosity and commitment to continuous learning and improvement.
关于高知特 (Cognizant)
高知特(Cognizant)(纳斯达克代码:CTSH)作为一家AI Builder和相关技术服务提供商,致力于通过打造全栈AI解决方案,帮助企业将人工智能投资转化为实际价值。公司凭借深厚的行业经验、流程优化和工程技术专长,将企业独特的业务场景融入科技系统,赋能组织释放人才潜能,推动切实成果,并帮助全球企业在瞬息万变的环境中保持领先。如需了解更多详情,敬请访问 cognizant.ai 或关注@cognizant。
补充雇佣信息
薪酬信息截至本职位发布之日为准。Cognizant 保留在适用法律允许的范围内随时修改该信息的权利。
申请人可能需要通过现场面试或视频会议的方式参加面试。此外,候选人在每次面试时可能需要出示其当前所在州或政府签发的有效身份证件。
Cognizant 是一家提供平等就业机会的雇主。在招聘过程中,您的申请和候选资格不会因种族、肤色、性别、宗教、信仰、性取向、性别认同、国籍、残疾、遗传信息、怀孕、退伍军人身份或任何其他受联邦、州或地方法律保护的特征而受到影响。