Overview
We are looking for an experienced Data Engineer to own and evolve the data platform that powers healthcare provider search, provider profiles, and appointment scheduling. You will take a hands on role with meaningful technical ownership across architecture, production operations, and engineering standards.
Project Overview:
This project is centered on building and evolving a healthcare provider data platform that enables accurate search, profile management, and appointment scheduling. It involves designing and operating reliable production pipelines and improving data quality, observability, and scalability across the platform. Project nickname: Provider Data Platform.
- Own end-to-end design, deployment, and operation of data pipelines.
- Lead the modernization of a large Apache Airflow environment and develop reusable DAGs, libraries, and standards.
- Build and maintain workflows for data processing, quality checks, provider updates, and data publishing.
- Develop automated data-quality frameworks and establish data contracts/source-of-truth rules.
- Ensure pipelines are observable, reliable, idempotent, and recoverable.
- Maintain and optimize Spark-based ETL jobs on AWS EMR, S3, and Elasticsearch.
- Improve performance, reliability, and cost efficiency of data workloads.
- Troubleshoot production data issues and lead root-cause analysis.
- Build monitoring, anomaly detection, and alerting tied to business outcomes.
- Improve deployment automation, configuration management, secrets handling, and infrastructure-as-code.
- Collaborate with product, engineering, operations, and data teams on scalable solutions.
- Contribute to architecture, technical planning, code reviews, and engineering standards.
- Mentor engineers and help drive decisions around scalability, maintainability, and technical debt.
- 5+ years of software/data engineering experience with production data systems.
- Strong Python and SQL skills.
- Strong experience with Apache Airflow and production data pipelines.
- Experience with AWS data infrastructure (S3, EMR, EC2, IAM, CloudWatch, etc.).
- Experience with Apache Spark or similar distributed processing frameworks.
- Experience with large datasets, Parquet, and data lakes.
- Experience with data quality, validation, reconciliation, and monitoring.
- Knowledge of software design, testing, Git, CI/CD, and production releases.
- Ability to balance scalability, reliability, maintainability, delivery speed, and cost.
- Strong communication skills and ability to work independently from problem definition to production.
- Healthcare or provider data experience.
- Elasticsearch or OpenSearch experience.
- Scala or Groovy for Spark.
- BigQuery and cross cloud workflows.
- Docker and Docker Compose.
- Terraform, CloudFormation, or AWS CDK.
- Data observability, lineage, and data contracts.
- Selenium, browser automation, and web data collection.
- Microsoft Graph, SharePoint, Slack, or Teams integrations.
- pytest, Ruff, mypy, or similar tooling.
- Experience with HIPAA or other compliance sensitive environments.
✨ Our intelligent job search engine discovered this job and republished it for your convenience.
Please be aware that the job information may be incorrect or incomplete. The job announcement remains the property of its original publisher. To view the original job and its full details, please visit the job's URL on the owner’s page.
Please clearly mention that you have heard of this job opportunity on https://ijob.am.

