Senior Data Scientist
Posted 18hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior Data Scientist building production machine learning and AI solutions on Azure Databricks for BlueFlag’s Department of Veterans Affairs platform. Modernizing legacy workloads and leading responsible LLM adoption.
Responsibilities:
- Consult with internal clients to frame business problems as analytical or ML problems, define success metrics, and set realistic scope
- Build data products and workflows supporting critical operations from source data discovery through production deployment
- Migrate and modernize workloads from R, Stata, and SAS to Python and Databricks
- Develop, train, and validate classification, regression, forecasting, clustering, and anomaly-detection models
- Engineer features and build reusable, governed feature pipelines on large datasets using PySpark and SQL
- Track experiments, register models, and manage model versions and promotion through MLflow
- Deploy models for batch scoring and real-time serving; automate retraining with scheduled jobs and CI/CD pipelines
- Monitor production models for performance, data drift, and data quality; configure thresholds and alerts
- Lead client AI adoption, including LLM and agentic use cases such as RAG, document summarization, and classification
- Evaluate LLM and agent outputs for accuracy, groundedness, and safety
- Document models for governance and responsible AI review
- Build reference use cases with tutorials, reference code, and training
- Host office hours and pair with client analysts and data scientists to upskill them
- Present findings and model results to technical and non-technical audiences, including leadership
Requirements:
- Bachelor's degree in Engineering, Computer Science, Statistics, Mathematics, Systems, Business, or a related scientific or technical discipline, and 15+ years of experience (or commensurate experience)
- Proficient in Python (pandas, NumPy, SciPy, scikit-learn) and advanced SQL (window functions, CTEs, query tuning) for data analysis
- 2+ years of hands-on work on a leading cloud data platform such as Databricks, Azure, AWS, or GCP
- Experience across the end-to-end data science workflow, from finding and assessing datasets to production deployment, including experiment tracking and model management with MLflow or similar
- Solid grounding in traditional machine learning: supervised and unsupervised methods, gradient-boosted trees (XGBoost, LightGBM), model selection, cross-validation, and hyperparameter tuning
- Strong applied statistics: hypothesis testing, regression, sampling, and experimental design
- Experience working with large datasets in a distributed environment (Spark/PySpark)
- Sound model evaluation practice: picking the right metrics, handling class imbalance, avoiding leakage, and explaining model behavior (for example SHAP or feature importance)
- Working knowledge of large language models (LLMs) and agentic AI workflows, including prompt design and RAG patterns
- Version control with Git and collaborative development practices (code review, branching, testing)
- Ability to explain technical work to non-technical stakeholders and turn ambiguous requests into defined deliverables
- Must be a citizen of the United States
- Must be able to obtain a public trust clearance
- Must be eligible to work in the United States
- Desired: 5+ years as a data scientist, shipping multiple products that run in operation
- Desired: Working experience with Databricks in Azure, including Unity Catalog, Delta Lake, Databricks Jobs, and Databricks notebooks/Repos
- Desired: MLOps experience across the lifecycle, including model serving, governed feature tables, CI/CD, production monitoring, model testing, automated retraining, lineage, and auditability
- Desired: Experience with Azure AI services or Mosaic AI
- Desired: Prior experience shipping products that use LLMs or AI agents, including evaluation and guardrails
- Desired: 2+ years building visual insights with Power BI, Databricks AI/BI dashboards, or Tableau
- Desired: Prior experience with R, Stata, or SAS
- Desired: Experience refactoring R, Stata, or SAS codebases to Python
- Desired: Experience with VA or federal healthcare data and handling PHI/PII under federal privacy and security requirements
- Desired: Familiarity with federal AI governance and responsible AI practices
- Desired: Experience training or mentoring analysts and data scientists
- Desired: Master's or PhD in a quantitative field
Benefits:
- Competitive salary
- Generous annual leave and paid holidays
- Comprehensive group health and dental plans
- 401(k) with company match
- Life insurance and AD&D coverage
- Ongoing training and professional development opportunities
















