Senior Data Engineer – Clinical Platforms, Databricks

Posted 3hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Senior Data Engineer building Databricks-based clinical trial platforms at Muttdata. Designing compliant lakehouse pipelines and integrations for regulated clinical data.

Responsibilities:

  • Design, build, and optimize enterprise data pipelines, lakehouse storage layers, and data models using Databricks (PySpark, Spark SQL, Delta Lake) to power custom clinical application backends.
  • Collaborate with frontend developers, software architects, and clinical research teams to build API-driven endpoints, data ingestion engines, and query layers for proprietary clinical trial software.
  • Build performant, standards-compliant data structures to store EDC outputs, audit trails, device telemetry, and patient-reported outcomes, enabling rapid querying and downstream analytics.
  • Partner with Clinical QA and Validation teams to ensure database structures, data pipelines, and clinical data repositories comply with GxP, 21 CFR Part 11, HIPAA, and GDPR.
  • Implement real-time and batch ingestion jobs connecting legacy clinical systems, central labs, EHRs, and wearable devices into a unified Databricks Lakehouse architecture.
  • Monitor, troubleshoot, and optimize Spark jobs, Delta Lake tables, and query execution times to support high-throughput, low-latency clinical platform workflows.

Requirements:

  • 4+ years of hands-on experience building production data pipelines and lakehouse architectures using Databricks, Delta Lake, and Apache Spark (PySpark or Scala).
  • Demonstrated experience building, extending, or maintaining custom software applications for clinical trials (e.g., custom EDC, CTMS, Clinical Data Repositories, or eCOA/ePRO platforms).
  • Deep understanding of clinical data standards and regulatory environments, including CDISC (SDTM, ADaM, CDASH), 21 CFR Part 11, GxP validation, and ICH-GCP guidelines.
  • Strong experience with relational schema design, dimensional modeling, and unstructured data handling within Delta Lake environments.
  • Proficiency in Python, SQL, RESTful API integrations, CI/CD pipelines, Git, and automated testing frameworks.
  • Experience working in cloud environments (AWS preferred, Azure or GCP).
  • Advanced English to discuss technical requirements and solutions with clients in the United States

Benefits:

  • Remote-first culture – work from anywhere!
  • AWS, DBT, Google Cloud, Azure & Databricks certifications fully covered
  • In-Company English Lessons.
  • Birthday off + an extra vacation week (Mutt Week! 🏖️)
  • Referral bonuses – help us grow the team & get rewarded!
  • Maslow: Monthly credits to spend in our benefits marketplace.
  • Annual Mutters' Trip – an unforgettable getaway with the team!
  • Monthly Childcare Reimbursement – Because supporting families matters too