Lead Data Engineer – PySpark, Palantir Foundry

Posted 2hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Lead Data Engineer building governed PySpark and Palantir Foundry pipelines for Logic20/20’s AI and analytics consulting practice. Improving releases, data infrastructure, and audit readiness for client projects.

Responsibilities:

  • Deliver client value and ensure high client satisfaction
  • Design, enhance, and maintain production-grade data pipelines supporting model output aggregation and downstream risk analysis
  • Establish and mature repository governance practices, including branching strategies, pull request standards, merge policies, release tagging, and version control workflows
  • Own release engineering practices for reproducible, traceable, and auditable production releases
  • Refactor and improve pipeline code for modularity, maintainability, scalability, and documentation quality
  • Develop configuration-driven pipeline patterns across environments and releases
  • Support testing, validation, benchmarking, and change management for critical data pipelines
  • Partner with data scientists, machine learning engineers, data engineers, product stakeholders, and other technical teams
  • Align teams on schemas, interfaces, inputs, and delivery expectations
  • Translate complex technical concepts into clear updates for technical and non-technical stakeholders
  • Improve structure and governance in codebases, repositories, and engineering workflows
  • Contribute to engineering best practices in a regulated, audit-sensitive delivery environment

Requirements:

  • 10-15+ years of data engineering, data science, machine learning engineering, and/or relevant experience using Python
  • Experience leading technical teams and overseeing enterprise-scale data initiatives
  • Strong expertise in PySpark, SQL, and cloud services
  • Ability to improve, refactor, or stabilize existing codebases and pipeline environments
  • Experience with cloud-optimized datasets, efficient partitioning strategies, and large-scale spatial operations
  • Understanding of machine learning model outputs flowing into downstream data pipelines, platforms, or production systems
  • Experience in highly regulated industries such as utilities, financial services, healthcare, insurance, or similar environments
  • Experience designing maintainable, scalable, and well-documented cloud-based data infrastructure or modern data platform environments
  • Experience supporting reproducibility, dataset versioning, release traceability, and audit readiness
  • Ability to define expected inputs, outputs, schemas, and interfaces across technical teams
  • Strong communication skills
  • Detail-oriented, governance-minded approach to engineering
  • Practical experience with Git-based workflows, code reviews, branching strategies, and release management
  • Experience with Palantir Foundry is highly preferred
  • Experience with GIS technologies and geospatial data platforms

Benefits:

  • Competitive base salary
  • Performance-based bonuses
  • Other incentives
  • Training and mentorship opportunities
  • Project opportunities for career development
  • Supportive, globally connected work environment