Senior Data Engineer – Clinical Platforms, Databricks
Posted 3hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior Data Engineer building Databricks-based clinical trial platforms at Muttdata. Designing compliant lakehouse pipelines and integrations for regulated clinical data.
Responsibilities:
- Design, build, and optimize enterprise data pipelines, lakehouse storage layers, and data models using Databricks (PySpark, Spark SQL, Delta Lake) to power custom clinical application backends.
- Collaborate with frontend developers, software architects, and clinical research teams to build API-driven endpoints, data ingestion engines, and query layers for proprietary clinical trial software.
- Build performant, standards-compliant data structures to store EDC outputs, audit trails, device telemetry, and patient-reported outcomes, enabling rapid querying and downstream analytics.
- Partner with Clinical QA and Validation teams to ensure database structures, data pipelines, and clinical data repositories comply with GxP, 21 CFR Part 11, HIPAA, and GDPR.
- Implement real-time and batch ingestion jobs connecting legacy clinical systems, central labs, EHRs, and wearable devices into a unified Databricks Lakehouse architecture.
- Monitor, troubleshoot, and optimize Spark jobs, Delta Lake tables, and query execution times to support high-throughput, low-latency clinical platform workflows.
Requirements:
- 4+ years of hands-on experience building production data pipelines and lakehouse architectures using Databricks, Delta Lake, and Apache Spark (PySpark or Scala).
- Demonstrated experience building, extending, or maintaining custom software applications for clinical trials (e.g., custom EDC, CTMS, Clinical Data Repositories, or eCOA/ePRO platforms).
- Deep understanding of clinical data standards and regulatory environments, including CDISC (SDTM, ADaM, CDASH), 21 CFR Part 11, GxP validation, and ICH-GCP guidelines.
- Strong experience with relational schema design, dimensional modeling, and unstructured data handling within Delta Lake environments.
- Proficiency in Python, SQL, RESTful API integrations, CI/CD pipelines, Git, and automated testing frameworks.
- Experience working in cloud environments (AWS preferred, Azure or GCP).
- Advanced English to discuss technical requirements and solutions with clients in the United States
Benefits:
- Remote-first culture – work from anywhere!
- AWS, DBT, Google Cloud, Azure & Databricks certifications fully covered
- In-Company English Lessons.
- Birthday off + an extra vacation week (Mutt Week! 🏖️)
- Referral bonuses – help us grow the team & get rewarded!
- Maslow: Monthly credits to spend in our benefits marketplace.
- Annual Mutters' Trip – an unforgettable getaway with the team!
- Monthly Childcare Reimbursement – Because supporting families matters too

















