Apply with hirly
Senior Data Engineer, Healthcare Data
Headwater Science · Research Triangle Park, North Carolina
Upload your resume to see how well you match this job — free, in seconds, no account needed.
Your resume is used only to score it against this job. If you don't create an account, it is deleted within 24 hours.
About Headwater Science Headwater Science (formerly NoviSci) is a data science and methods company specializing in principled, reproducible evidence generation for complex clinical and regulatory challenges. With deep expertise in comparative effectiveness, causal inference, healthcare utilization and expenditure research, and regulatory-grade analytical software, Headwater Science provides the methodological foundation that delivers reproducible analytic pipelines, novel epidemiologic and statistical methods, and regulatory-grade software validated to hold up under the most demanding scrutiny. The company works with life sciences organizations as a long-term scientific partner. Headwater Science is a Highlander Health company. Learn more at headwaterscience.com. The Role We are seeking a talented Senior Data Engineer to design, build, and maintain the data pipelines that power our real-world evidence research. Our studies depend on turning large, messy, vendored healthcare data into analytic-ready datasets that epidemiologists and statisticians can trust. You’ll help us construct these datasets alongside our statistical programming team, and you’ll also help shape the processes, systems, and infrastructure that produce regulatory-grade data. That includes collaborating with our software development team to define the tools that need to exist, the workflows those tools support, and the requirements they have to meet. It also includes serving as the technical point of contact for our administrative claims and EHR data vendors and becoming our in-house expert on each source’s structure, conventions, and limitations. This is a hands-on role. Much of your time will be spent inside real studies with tight deadlines, and the standards and tooling you build will grow out of that work rather than in isolation. Today the work runs on R, SQL, and AWS, and while AWS is here to stay, you'll have a voice in how the rest of the stack evolves to meet the needs of our studies. We’re a small team with broad scope, and as our data engineering work expands, there’s room for this role to grow with it. How We Work Inside a study, you’ll be one of several people working closely with the epidemiology team to ensure that the analytic-ready dataset faithfully follows the study protocol, and you’ll report back on what you learn about the data along the way. Those datasets are consumed by our statisticians, so you’ll also be in close contact with them around data deliveries: keeping them informed on timelines, documenting what’s in each delivery, and answering questions as the analysis takes shape. Across studies, you’ll work with the statistical programmers who use the systems and procedures you help shape. You’ll support them day to day as they build datasets on active projects, and you’ll also help refine the procedural requirements they follow: how data are accessed, how pipelines are structured and documented, and what a finished dataset has to satisfy before it’s released. The time you spend building datasets yourself is also how you’ll see where those systems and procedures need to improve What You’ll Do Build analytic-ready datasets
- Work on active studies building cohorts and datasets from administrative claims and EHR sources.
- Translate protocol specifications into concrete data logic.
- Prepare data deliveries for statisticians, with documentation on what’s included, how it was derived, and any known limitations. Strengthen the systems and standards
- Help evolve the path from raw vendor deliveries to analytic-ready datasets: where data land, how they're organized and versioned, and how each step is documented and verified.
- Contribute to the operating procedures that statistical programmers follow when building datasets.
- Help define the release criteria for datasets, including the quality checks they pass and the documentation that accompanies them.
- Inform the design of tools the software development team builds for data processing: the workflows they support, the interfaces they expose, and the requirements they must meet.
- Participate in the evaluation and adoption of tools and technologies for processing source data, including proposing changes where the current approaches fall short.
- Contribute to the design and administration of our AWS data infrastructure, including storage layout, access controls, and audit trails for protected health information. Own the source data
- Onboard new vendor datasets: profile the data, learn its conventions and gaps, and stand up its ingestion.
- Serve as the technical point of contact for data vendors on dictionaries, delivery mechanics, refresh schedules, and quality issues.
- Maintain the internal reference material that describes how each source represents information and where it falls short.
- Link datasets across sources as needed, coordinating with data vendors on the generation and management of privacy-preserving linkage keys such as Datavant tokens. What You’ll Bring
- 5+ years designing, building, and owning production data pipelines, including ingestion of large files, quality validation, reproducibility, and version control.
- Experience processing and transforming large datasets in relational databases or distributed query engines such as Spark, whether through SQL directly or APIs.
- Experience designing software, including writing requirements documents and specifying programming interfaces.
- Proficiency in R, or a willingness to learn it. Our current data processing pipelines are written in R.
- Experience building on AWS.
- A track record of turning ad hoc practices into documented, repeatable processes for a team of programmers or analysts.
- Experience evaluating and selecting data tools and architecture, including the tradeoffs involved in adopting or replacing a technology. Nice to Have
- Hands-on experience with administrative claims or EHR data, including their coding systems and enrollment logic. Familiarity with common data models such as OMOP is a plus.
- Exposure to modern data engineering technologies such as Snowflake, Spark, AWS Batch, S3, Arrow, Docker, and GitLab CI.
- A working knowledge of pipeline orchestration frameworks such as Airflow, targets, Nextflow, or Snakemake.
- Experience handling protected health information or other regulated data.
- Familiarity with software validation practices under frameworks such as GxP or 21 CFR Part 11. What We Offer
- Hybrid work — 3 days/week in office to collaborate with the team
- Comprehensive health, dental, and vision for you and your family
- 401(k) with company match
- Generous PTO and company holidays
- Paid parental leave If you are ready to be part of a team where your work truly matters- where your expertise is valued, your growth is supported, and your contributions help shape the future of healthcare- Headwater Science is the place for you. We’re building something meaningful together, and we’d love for you to be a part of it.