Job Summary
This role is for an experienced Architect responsible for designing and guiding implementation of modern LakeHouse data platforms using Spark job definition OneLake SQL and PySpark in a hybrid work model. The Architect will create scalable secure data solutions that support complex analytics and reporting needs with a focus on reliable pipelines and optimized query performance that help the organization deliver better data driven services to clients.
Responsibilities
- Design robust LakeHouse based data architectures that integrate diverse enterprise sources into OneLake while ensuring scalability reliability and alignment with organizational data strategy to support analytical and reporting workloads efficiently.
- Define end to end Spark job definition standards including coding conventions runtime configurations and resource optimization practices to ensure consistent performance and maintainability across large scale data processing pipelines.
- Develop efficient SQL and PySpark data transformation logic that supports complex business rules while maintaining data quality integrity and traceability for downstream analytics and operational systems.
- Oversee design of hybrid deployment patterns that enable secure access to the LakeHouse platform across on premises and cloud environments while aligning with enterprise security and compliance policies.
- Provide technical guidance to project teams on optimal use of OneLake capabilities for data ingestion storage partitioning and lifecycle management to deliver resilient and cost effective data solutions.
- Collaborate with stakeholders to translate analytical and reporting needs into detailed data models and pipeline designs using SQL and PySpark that enable timely and accurate insights for business decision making.
- Optimize Spark job execution by tuning cluster configurations adjusting partition strategies and refining transformation logic to reduce processing times and resource consumption while maintaining reliability.
- Implement standardized monitoring logging and alerting for LakeHouse workloads to proactively identify performance issues data quality problems and operational risks and drive timely resolution.
- Document architecture decisions data flow diagrams and technical standards for LakeHouse Spark OneLake SQL and PySpark implementations to support knowledge sharing and long term platform sustainability.
- Coordinate with data governance teams to embed metadata management access controls and audit capabilities within the LakeHouse environment to protect sensitive information and meet regulatory expectations.
- Review solution designs and code artifacts from project teams to ensure alignment with architectural principles coding best practices and nonfunctional requirements such as performance scalability and resilience.
- Drive continuous improvement of LakeHouse and Spark job definition practices by evaluating emerging tool capabilities frameworks and patterns and recommending pragmatic enhancements that deliver measurable value.
- Partner with business and product teams to identify opportunities where advanced analytics enabled by the LakeHouse platform can improve customer experiences operational efficiency and societal impact through better financial insights.
Qualifications
- Demonstrate proven architecture experience of at least ten years designing and delivering large scale data platforms with strong focus on distributed processing and enterprise integration.
- Show advanced proficiency in LakeHouse concepts including unified storage governance data modeling and workload management to design solutions that serve both analytical and operational use cases.
- Exhibit deep hands on expertise in Spark job definition and PySpark development including pipeline orchestration performance tuning and error handling for high volume data workloads.
- Possess strong SQL skills for building complex queries views and data models that support reporting dashboards and self service analytics while ensuring accuracy and consistency.
- Apply practical knowledge of OneLake features for organizing data zones managing security boundaries and optimizing storage strategies to achieve reliable and cost conscious solutions.
- Bring experience in hybrid work environments and be comfortable collaborating across distributed teams using remote and onsite engagement models while maintaining effective communication.
- Leverage retail banking domain exposure when available to better understand account transactions risk indicators and regulatory needs enabling data solutions that support responsible financial services.
- Utilize strong problem solving and analytical abilities to assess architectural tradeoffs propose clear options and recommend solutions that balance performance cost and long term maintainability.
Certifications Required
Preferred certifications include Azure Data Engineer Associate or Databricks Certified Data Engineer Professional or equivalent modern data architecture credential.
关于高知特 (Cognizant)
高知特(Cognizant)(纳斯达克代码:CTSH)作为一家AI Builder和相关技术服务提供商,致力于通过打造全栈AI解决方案,帮助企业将人工智能投资转化为实际价值。公司凭借深厚的行业经验、流程优化和工程技术专长,将企业独特的业务场景融入科技系统,赋能组织释放人才潜能,推动切实成果,并帮助全球企业在瞬息万变的环境中保持领先。如需了解更多详情,敬请访问 cognizant.ai 或关注@cognizant。
补充雇佣信息
薪酬信息截至本职位发布之日为准。Cognizant 保留在适用法律允许的范围内随时修改该信息的权利。
申请人可能需要通过现场面试或视频会议的方式参加面试。此外,候选人在每次面试时可能需要出示其当前所在州或政府签发的有效身份证件。
Cognizant 是一家提供平等就业机会的雇主。在招聘过程中,您的申请和候选资格不会因种族、肤色、性别、宗教、信仰、性取向、性别认同、国籍、残疾、遗传信息、怀孕、退伍军人身份或任何其他受联邦、州或地方法律保护的特征而受到影响。







