跳到主要內容

Data Scientist

00069454811



Job Summary

We are looking for two AI Data Engineers to build operate and continuously improve the data pipelines retrieval infrastructure and ML and LLMOps foundations that power our AI initiatives. These professionals will be responsible for turning reference architectures and data contracts into robust production-grade implementations that serve conversational AI assistants dashboard copilots autonomous agents RAG applications and predictive ML models.


Responsibilities

We are looking for two AI Data Engineers to build operate and continuously improve the data pipelines retrieval infrastructure and ML and LLMOps foundations that power our AI initiatives. These professionals will be responsible for turning reference architectures and data contracts into robust production-grade implementations that serve conversational AI assistants dashboard copilots autonomous agents RAG applications and predictive ML models.

The role involves data pipeline engineering including building testing and maintaining production pipelines in batch and real-time environments using Snowflake PySpark Delta Lake and Kafka. Engineers will implement data quality checks schema validation and alerting at every pipeline stage migrate legacy ETL and data warehouse systems to cloud-native AWS or Azure architectures with measurable latency and cost improvements and maintain CI/CD pipelines with automated testing deployment rollback and infrastructure as code using Terraform and GitHub Actions.

They will also work on RAG vector and retrieval infrastructure by building end-to-end retrieval systems covering document ingestion embedding pipelines vector store management with Pinecone FAISS ChromaDB or OpenSearch and hybrid retrieval layers. Responsibilities include implementing chunking metadata filtering and re-ranking tuned for precision recall and latency maintaining data freshness and index consistency and instrumenting with context relevance and faithfulness metrics.

In addition they will support semantic layer and knowledge infrastructure by implementing and maintaining business entity mappings ontologies and knowledge graphs with Neo4j building and versioning feature stores and semantic data contracts serving ML models and LLM applications and managing metadata data lineage and audit trail instrumentation across the platform.

The engineers will contribute to ML and LLMOps pipeline support by building ML data infrastructure for training curation feature engineering and MLflow experiment tracking supporting LLM fine-tuning workflows through corpus curation quality filtering and dataset formatting implementing automated evaluation pipelines for factual accuracy hallucination detection and regression tracking and maintaining production monitoring dashboards for pipeline health model metrics and alerting.

They will also develop agentic data infrastructure by building and maintaining data APIs tool schemas and memory or state stores for autonomous agents implementing agent observability to capture inputs retrieved context tool calls reasoning traces and outputs and maintaining text-to-SQL layers semantic query interfaces and context APIs for conversational AI consumers.

Governance security and data quality responsibilities include implementing role-based and attribute-based access PII detection and masking data classification and audit logging enforcing data contracts and schema governance with automated breaking-change detection and versioned migrations building data quality monitoring for completeness freshness and consistency with automated alerting and root-cause tooling and supporting compliance readiness through audit trails data provenance and regulatory documentation.

Candidates must have five to eight years of data engineering experience and at least two years of production AI ML or LLM-era data infrastructure experience. They should demonstrate proven expertise in building production pipelines at scale in batch and streaming environments with Snowflake and AWS or Azure deep knowledge of Python PySpark Snowflake Delta Lake Kafka and Spark Structured Streaming and hands-on experience with vector stores embedding pipelines and retrieval infrastructure in production RAG environments. Working knowledge of MLOps including MLflow CI/CD for AI automated evaluation and production monitoring along with strong grounding in data governance quality frameworks and compliance-aligned engineering is essential.

Technical skills required include expert-level proficiency in Python SQL PySpark Kafka Delta Lake AWS services such as S3 Glue Kinesis EKS and Redshift Docker Kubernetes GitHub Actions and Snowflake. Strong skills in LangChain LlamaIndex LLM APIs such as OpenAI Bedrock Claude and HuggingFace Pinecone FAISS ChromaDB OpenSearch MLflow FastAPI and Neo4j are expected. Solid skills in CI/CD pipelines CloudWatch Grafana data lineage platforms and MCP along with familiarity with LangGraph prompt engineering RLHF dataset preparation and LLM fine-tuning workflows are desirable.

The technology stack includes Delta Lake PySpark Kafka Spark Structured Streaming Snowflake AWS services such as S3 Glue EKS Bedrock Kinesis Redshift and Lambda Azure Kubernetes Docker Terraform GitHub Actions Jenkins MLflow LangChain LlamaIndex HuggingFace OpenAI AWS Bedrock Claude Pinecone FAISS ChromaDB OpenSearch Neo4j FastAPI Python SQL MCP LangGraph MLOps CI/CD Grafana and CloudWatch.


关于高知特 (Cognizant)
高知特(Cognizant)(纳斯达克代码:CTSH)作为一家AI Builder和相关技术服务提供商,致力于通过打造全栈AI解决方案,帮助企业将人工智能投资转化为实际价值。公司凭借深厚的行业经验、流程优化和工程技术专长,将企业独特的业务场景融入科技系统,赋能组织释放人才潜能,推动切实成果,并帮助全球企业在瞬息万变的环境中保持领先。如需了解更多详情,敬请访问 cognizant.ai 或关注@cognizant。

补充雇佣信息
薪酬信息截至本职位发布之日为准。Cognizant 保留在适用法律允许的范围内随时修改该信息的权利。
申请人可能需要通过现场面试或视频会议的方式参加面试。此外,候选人在每次面试时可能需要出示其当前所在州或政府签发的有效身份证件。
Cognizant 是一家提供平等就业机会的雇主。在招聘过程中,您的申请和候选资格不会因种族、肤色、性别、宗教、信仰、性取向、性别认同、国籍、残疾、遗传信息、怀孕、退伍军人身份或任何其他受联邦、州或地方法律保护的特征而受到影响。

帮助您蓬勃发展与成长的福利

我们的福利计划以您为中心打造——帮助您享受充实、平衡且健康的生活。
有葉子的植物的藍色線條圖

财务健康

我们会定期审查市场数据,确保薪酬体现您所带来的价值。您的福利不仅限于薪资,还可能包括退休计划、财务教育等。

Stay Healthy Midnight Blue RGB

身心健康

我们通过带薪休假、在条件允许下的灵活工作安排、医疗保障计划、心理咨询、心理健康盟友计划等,赋能您将身心健康放在首位。

Build The Career You Want Midnight Blue RGB

您的职业发展,由您做主

在 Cognizant 提供的 35 万多个岗位中,您将有机会探索新的技术、行业和工作地点,并打造推动职业发展的关键技能。

Making A Meaningful Impact Midnight Blue RGB

现实世界的影响力

想想您所依赖的那些知名品牌。很可能,他们也依赖我们来帮助强化其业务。在这里,您将把大胆的想法转化为改善全球生活的解决方案。

还没有找到合适的机会吗?

获取为您量身定制的最新职位机会、招聘活动和公司新闻!

掌握最新动态