Emory logo

Data Scientist, Data & AI

Hiring from
United States
Work type
Remote
Posted
Is this job info correct?
Show job description

Overview

Be inspired. Be valued. Belong.

At Emory Healthcare we fuel your professional journey with better benefits, valuable resources, ongoing mentorship and leadership programs for all types of jobs, and a supportive environment that enables you to reach new heights in your career and be what you want to be. We provide:

  • Comprehensive health benefits that start day 1
  • Student Loan Repayment Assistance & Reimbursement Programs
  • Family-focused benefits
  • Wellness incentives

Ongoing mentorship, development, leadership programs...and more!

Remote position, however candidates must reside in one of the following states: Alabama, Arkansas, Florida, Georgia, Illinois, Louisiana, Michigan, New Hampshire, North Carolina, Ohio, Pennsylvania, South Carolina, Tennessee, Texas, Virginia, or Wisconsin.

Description

The Data Scientist turns Emory's clinical and operational data into measures, models and AI tools that leaders and clinicians use every week. You'll work directly with physician leaders, operations and finance on questions such as: are patients going home within the expected length of stay for their condition, and why not; how productive are our physicians against target; which patients belong in a disease cohort, and how are they doing; what does the text of a pathology or radiology report tell us that the structured data doesn't.
You'll build on the Emory Unified Data Platform (UDP): Epic Clarity and Caboodle data in Microsoft Fabric, using SQL, Python/PySpark and R, Power BI semantic models, and AI agents grounded in that data. The bar is high: our numbers have to match what leaders already trust, or explain exactly why they don't.

RESPONSIBILITIES:

  • Build trusted datasets and measures from Epic data (Clarity / Caboodle): define cohorts and metrics with clinical and business owners, write the SQL / PySpark, and validate every number against existing reports before it is used.
  • Develop and deploy models: statistical and machine-learning models for operational and clinical questions (length of stay vs GMLOS, readmission and transfer risk, throughput, productivity), from exploration through scheduled production pipelines in Fabric / Azure, with monitoring and documentation.
  • Apply NLP and LLMs to clinical text (pathology, radiology and clinical notes) to extract structured facts, with test sets, evaluation and human review before results are trusted.
  • Deliver to users: Power BI semantic models and reports, and conversational data agents that answer leaders' questions in plain English, with appropriate row-level security.
  • Protect patient data: minimum-necessary access, de-identification where possible, HIPAA, and research vs. quality-improvement boundaries (IRB) understood and respected.
  • Communicate: explain results, definitions and caveats to physicians and executives clearly; write short, plain-language documentation.
  • Raise the bar for the team: reusable code, peer review, versioned definitions, and mentoring of analysts.

PREFERRED QUALIFICATIONS:

  • Epic data (Clarity, Caboodle, Cogito) and its clinical and revenue-cycle concepts (encounters, DRG / GMLOS, wRVU, orders and results).
  • Microsoft Fabric / Azure (lakehouses, Spark notebooks, Power BI semantic models and DAX, Azure AI Foundry).
  • NLP / LLM work on pathology, radiology, clinical notes, including evaluation and human-in-the-loop review.
  • PHI handling, quality improvement vs research/IRB
  • Healthcare operations metrics (LOS, throughput, productivity) or oncology / quality-improvement analytics.
  • Semantic models, reports and AI data agenets, with security
  • Research data standards (OMOP, FHIR) and IRB processes.
  • Version control (Git), Agile delivery, MLOps.

Top skills being sought for this opening:

  1. SQL on large healthcare data, ideally Epic.
  2. Python or R.
  3. A validation mindset: be able to articulate discussion about a time your numbers didn't match an existing report.
  4. Communicating with physicians and executives.
  5. Clinical text NLP or LLMs, with evaluation.
  6. Fabric, Power BI or Azure (trainable).

MINIMUM QUALIFICATIONS:

  • Education: A bachelors degree, a master's degree is preferred
  • Experience: At least 5 years of experience either in healthcare or data analytics - At least 2 years of experience working for payor companies or healthcare systems - At least 2 years of experience in relational database - Proficient in SQL and at least 1 other programming languages (e.g., Python, R, SAS) - Experience in statistical modeling, machine learning, AI - Excellent presentation skills - Experience in strategic planning and workflow optimization - Mentor and coach the broader analytic team and service line operation leaders

Additional Details

Emory is an equal opportunity employer, and qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability, protected veteran status or other characteristics protected by state or federal law.

Emory Healthcare is committed to providing reasonable accommodations to qualified individuals with disabilities upon request. Please contact Emory Healthcare’s Human Resources at careers@emoryhealthcare.org. Please note that one week's advance notice is preferred.

Similar jobs

Apply for this job