Job Summary
Design and optimize large scale data processing solutions as a senior developer using Spark and AWS based services in a hybrid work arrangement. Apply advanced expertise in Spark optimization AWS Glue EMR Redshift S3 IAM and CloudWatch to build reliable data pipelines and analytical platforms. Collaborate with cross functional teams in day shifts to deliver secure efficient and business aligned data products that support strategic decision making.
Responsibilities
Design and implement scalable data processing pipelines using Apache Spark and AWS Glue to efficiently transform and prepare data for analytics and reporting across the organization.Optimize Spark jobs by tuning configurations improving partition strategies and refining data formats to consistently reduce runtime control resource consumption and improve overall reliability of data workflows.Build and manage AWS EMR clusters to run complex distributed data processing workloads while maintaining cost efficiency and ensuring high performance for critical business processes.Develop robust data ingestion and storage solutions using Amazon S3 as the central data lake enforcing folder structures and retention practices that support governance reusability and efficient querying.Create and maintain analytical data models in Amazon Redshift that enable fast and accurate reporting ensuring proper distribution keys sort keys and compression settings are applied for query performance optimization.Configure and maintain AWS CloudWatch monitoring for data pipelines and compute resources setting up detailed metrics and alarms that enable proactive detection and resolution of operational issues before they impact business stakeholders.Administer AWS IAM policies for data engineering components to ensure secure access control applying fine grained permissions that protect sensitive data while allowing teams to work efficiently within approved boundaries.Collaborate with product owners analysts and engineering peers during hybrid work days to gather requirements refine technical designs and deliver high quality data solutions that directly support strategic business initiatives.Review and refactor existing data processing code to improve readability performance and maintainability mentoring less experienced developers by sharing best practices in Spark optimization and cloud native development.Document data flows transformation logic and operational procedures in clear technical references that enable future enhancements onboarding of new team members and transparent communication with audit and compliance functions.Troubleshoot complex production incidents in day shift schedules by analyzing logs metrics and data anomalies applying structured problem solving to restore service quickly and prevent recurrence through targeted improvements.Contribute to continuous improvement of the data platform by evaluating new AWS features and data engineering techniques recommending adoption paths that increase reliability scalability and positive impact on business and society.Align daily development activities with the company purpose by building data solutions that support ethical decision making enable better customer experiences and improve operational efficiency in a sustainable manner.
Qualifications
Possess around ten to twelve years of professional experience as a data focused developer working on large scale distributed systems with a strong track record of delivering stable and performant solutions.Demonstrate deep hands on expertise in Apache Spark optimization including experience with partitioning strategies caching approaches broadcast joins and efficient use of data formats such as parquet.Show advanced practical knowledge of AWS Glue for building serverless data integration jobs including authoring scripts configuring crawlers and managing glue data catalogs for structured and semi structured data.Apply strong working experience with AWS EMR clusters for running Spark workloads including cluster sizing auto scaling configurations and integration with S3 based data lakes.Exhibit solid experience with Amazon Redshift for designing analytical schemas tuning queries and managing workload patterns to serve reporting and dashboarding needs quickly and reliably.Use Amazon S3 in depth for organizing data repositories applying lifecycle policies and supporting data governance and access patterns suitable for large enterprise environments.Implement secure and compliant solutions with AWS IAM by defining roles and policies that restrict access appropriately for data compute and monitoring resources in a complex cloud environment.Employ AWS CloudWatch proficiently for logging metrics and alerts creating actionable dashboards and notifications that help sustain high availability and timely response to anomalies.Communicate effectively with cross functional teams in a hybrid work setting demonstrating clear verbal and written skills that support collaborative design estimation and delivery of data solutions.Adapt quickly to evolving business requirements and technology advancements in the data engineering space showing curiosity and commitment to continuous learning and improvement.
About Cognizant
Cognizant (Nasdaq: CTSH) is an AI Builder and technology services provider, bridging the gap between AI investment and enterprise value by building full-stack AI solutions for our clients. Our deep industry, process and engineering expertise enables us to build an organization’s unique context into technology systems that amplify human potential, drive tangible outcomes and keep global enterprises ahead in a fast-changing world. See how at cognizant.ai or @cognizant.
Additional employment information
Compensation information is accurate as of the date of this posting. Cognizant reserves the right to modify this information at any time, subject to applicable law.
Applicants may be required to attend interviews in person or by video conference. In addition, candidates may be required to present their current state or government issued ID during each interview.
Cognizant is an equal opportunity employer. Your application and candidacy will not be considered based on race, color, sex, religion, creed, sexual orientation, gender identity, national origin, disability, genetic information, pregnancy, veteran status or any other characteristic protected by federal, provincial or local laws.