Software Engineer – Agentic Data Pipelines

Posted 4hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Software engineer building agentic LLM systems for biomedical data acquisition, curation, and quality control. Supporting Iambic Therapeutics’ AI-driven drug discovery platform and Enchant model training.

Responsibilities:

  • Design, build, and maintain agentic systems that convert biomedical data-source pointers into reviewed, versioned datasets
  • Develop LLM-based pipelines for data cleaning, normalization, and formatting across molecular, genomic, clinical, and literature data
  • Implement automated quality-control workflows to detect anomalies, flag inconsistencies, and enforce data standards
  • Evaluate and iterate on agent architectures, prompting strategies, tool definitions, validation loops, and evaluation harnesses
  • Collaborate with ML scientists on the Enchant team to translate data requirements into scalable acquisition and processing systems
  • Monitor and maintain distributed data pipelines in production, diagnose failures, and improve robustness
  • Document data provenance, processing decisions, and quality metrics for reproducibility and auditing
  • Operate agents safely using sandboxed execution, least-privilege credentials, restricted network access, and audit logs
  • Raise potential security risks to the team
  • Contribute to a longer-term natural-language orchestrator for drug discovery inference, fine-tuning, virtual screens, and dataset analysis

Requirements:

  • Master’s degree in a computational STEM field, or a Bachelor's with 2+ years of industry experience
  • Strong Python engineering skills and experience building and maintaining production-quality software
  • Hands-on experience with LLM APIs such as Claude and GPT
  • Experience with agentic patterns including tool use, orchestration, and multi-step reasoning
  • Familiarity with biomedical or chemical data sources and formats such as PDB, UniProt, ChEMBL, SDF/MOL, and FASTA
  • Data engineering fundamentals including ETL design, data validation, and structured and unstructured data at scale
  • Hands-on experience with Python testing frameworks such as pytest fixtures and parametrization
  • Experience with agent orchestration frameworks and evaluation harnesses for LLM-generated code
  • Familiarity with cloud infrastructure and workflow orchestration such as AWS, Docker, and Kubernetes
  • Knowledge of multimodal biomedical data spanning small molecules, proteins, assays, images, ‘omics, and/or clinical records
  • Experience with large-scale dataset construction or curation for ML model training
  • Knowledge of agent security practices including sandboxing, scoped credentials, and prompt injection
  • Authorized to work in the United States; visa sponsorship status addressed in application

Benefits:

  • Company-paid healthcare
  • Flexible spending accounts
  • Voluntary life insurance
  • 401K matching
  • Uncapped vacation
  • Onsite gym
  • Dining facilities
  • Brand-new state-of-the-art facility
  • Easy access to great places to live and play