The breakthrough AI production platform that allows anyone to create compelling commercials and spec spots in minutes.
Data/AI Scientist II
Location
United States
Posted
2 days ago
Salary
$95K - $164K / year
Seniority
Senior
Job Description
Data/AI Scientist II
Waymark
• Build and validate ML and AI models on claims and EHR data for care delivery use cases including risk stratification and care gap identification, under guidance of senior data scientists. • Support the development of LLM-based tools and AI applications for care team workflows. • Develop and maintain Python code for data processing, model development, and ML pipelines, following software engineering best practices including version control, code review, and testing. • Collaborate with engineering, product, and analytics teams to understand requirements and translate them to technical solutions. • Develop Healthcare Subject Matter Expertise: Show a strong interest in healthcare data structures, quantitative methods, and data science applications in healthcare.
Job Requirements
- A Master’s degree in Data Science, Computer Science, Statistics, or a related field.
- Python Proficiency: 3+ years of hands-on Python experience including academic and project work — demonstrated through a GitHub portfolio or equivalent.
- Solid foundations in statistical and machine learning methods, including classical ML and an understanding of modern AI and LLM applications.
- Strong SQL skills and comfort working with large, messy real-world datasets.
- Experience working with Git and collaborating in a team codebase.
- Strong communication skills and ability to work closely with cross-functional stakeholders.
- 1-2 years of industry data science experience (preferred).
- Exposure to healthcare data (claims or EHR) through coursework, research, or work experience (preferred).
Benefits
- Stock Options: Opportunity to invest in the company’s growth.
- Work-from-Home Stipend: A dedicated stipend for your first year to help set up your home office.
- Medical, Vision, and Dental Coverage: Comprehensive plans to keep you and your family healthy.
- Life Insurance: Basic life insurance to give you peace of mind.
- Paid Time Off: 20 vacation days, accrued over the year, plus 11 paid holidays.
- Parental Leave: 16 weeks of paid leave for birthing parents after six months of employment, and 8 weeks of bonding leave for non-birthing parents.
- Retirement Savings: Access to a 401(k) plan with a company contribution, subject to a vesting schedule.
- Commuter Benefits: Convenient options to support your commute needs.
- Professional Development Stipend: A dedicated stipend supports professional development and growth.
Related Guides
Related Categories
Related Job Pages
More Data Scientist Jobs
Data Lead – Advanced Primary Care Management
Cadence SolutionsProtecting your data. Simplifying your content.
• Own the analytical strategy for Cadence's digital primary care program end-to-end by defining what outcomes and KPIs matter, diagnosing ambiguous operational and clinical problems, and serving as the primary thought partner to clinical and operational leadership; exploring patient vitals, EHR data, and clinician-generated data to surface insights that close care gaps, drive better outcomes, and inform program strategy. • Define and communicate outcome frameworks for AI-powered primary care, translating clinical and operational performance into clear, credible narratives for external partners including health systems and payers, and serving as the analytical voice in those conversations. • Design and build core data models in dbt from the ground up, establishing canonical metric definitions, data collection standards, and the structural foundation that the broader team relies on; maintain the modern data stack (Snowflake, Fivetran, dbt) to keep infrastructure accurate, well-documented, and built to scale. • Deliver advanced analytics across the full stack from pipeline construction and data engineering to predictive modeling, staffing and capacity forecasting, and revenue modeling in partnership with operations and executive leadership. • Build and maintain reports and dashboards that monitor growth, engagement, clinical outcomes, and Cadence's demonstrated value to patients and health system partners with metrics that are trustworthy and tied to decisions. • Embed AI meaningfully into the analytics function by identifying high-value use cases, building reusable workflows and automation, and leading adoption of LLM-assisted tools that reduce manual overhead and expand what the team can produce.
Senior Data Scientist
Gilbane Building CompanyHeadquartered in Providence, Rhode Island, Gilbane, Inc. is a family of construction, facility management, and real estate development companies that includes Gilbane Building Comp
Role Description Gilbane is looking to add a highly skilled Senior Data Scientist to our Information Technology Team. You are a coach/leader who leads with an inclusive and empathetic mindset. You provide feedback and guidance to help others excel in their current or future roles. You determine priorities, delegate work, and effectively communicate progress. You establish measures to assess the impact, quality, and timeliness of results while praising successes and sharing lessons learned. You build high performing teams by attracting, engaging, developing, and retaining talented individuals through motivation and discipline to maximize impact on the organization and the individual. You leverage business insights by understanding industry trends, local market/economic conditions, and Gilbane’s business model to make critical decisions and create competitive advantage. You deploy a strategic mindset when considering solutions to long-term opportunities and risks that may develop in the future. Your core values match Gilbane’s: Integrity, Caring, Teamwork, Toughmindedness, Dedication to Excellence, Discipline, Loyalty, and Entrepreneurship. Use your deep scientific research methods, statistics, machine learning, causal inference, and AI skills to study our business and deploy analytical solutions that optimize business outcomes. - You are a scientist that is a master at working with data. - You’ll develop close stakeholder relationships, solve business problems, and educate stakeholders on analytics, AI / ML processes and capabilities. - Leverage the model development process to identify critical drivers of predicted outcomes (variable importance) and guide business to take action on these insights. - Partnering with business SMEs in the process to enrich model insights, improve model performance, and ensure user adoption. - Mine various data systems, clean and structure data, develop data pipelines, and identify areas for data quality improvements in the enterprise data environment. - Partner with Data Engineering and BI teams to put models and insights into production. - Develop causal inference / explanatory path models to help business understand drivers of key outcomes. - Work with stakeholders to ensure understanding of models and implement business process changes to improve outcomes based on model output. Qualifications - Bachelor’s degree - PhD desired - Master’s at least required in statistics, data science, economics, psychology (research / experimental) or other heavy quant focused social or other science (biometrics). - Need deep knowledge of statistical analysis, predictive modeling, machine learning, data mining, feature engineering, causal inference, scientific research methods. - Requires strong skills in SQL and Python or R. - Excellent communicator / consultative approach, PowerPoint and presentation skills required. Requirements - This position can be performed remotely or from any U.S. location where Gilbane has an office. - Salary to be determined based on factors such as geographic location, skills, education, and/or experience of the applicant, as well as the internal equity and alignment with the team. - The pay ranges from $141,000 - 220,000 plus benefits and retirement program. - Qualified applicants who are offered a position must pass a pre-employment substance abuse test. Benefits - Gilbane offers an excellent total compensation package which includes competitive health and welfare benefits and a generous profit-sharing/401k plan. - We invest in our employees’ education and have built Gilbane University into a top training organization in the construction industry.
• Own the roadmap and backlog for our reference data portfolio — payer, SDOH, mortality, formulary, and other reference datasets: which fields we carry, how they’re defined, how they’re sourced, and how they’re maintained over time • Write clear PRDs, user stories, and data specs that translate customer needs into requirements engineering and data teams can build from • Partner with develop to curate reference data across payer, SDOH, mortality, and formulary domains • Monitor attribute quality — coverage, fill rates, mapping accuracy, drift over time — and prioritize fixes based on customer impact • Track market and data-source changes that affect our reference datasets: payer M&A, plan rebrands, Medicare Advantage contract changes, PBM shifts, formulary updates, new SDOH data sources, and mortality data refresh cycles — and translate them into backlog updates • Partner with commercial, customer success, and analytics teams to understand how customers actually use our reference datasets and where definitions need to sharpen • Write and maintain data dictionaries, attribute definitions, and release notes so internal teams and customers can confidently use what we ship
• Translate customer goals — such as improving differential diagnosis, evaluating a clinical note summarizer, testing a RAG-based medical literature assistant, or creating preference data for patient-facing chatbots — into dataset specifications, taxonomies, rubrics, sampling plans, and acceptance criteria. • Make multimodal health AI a core focus: design training and evaluation datasets across clinical text, medical images, waveforms, structured EHR data, claims, trial data, medical literature, patient communications, payer policies, drug information, and other clinical artifacts, as well as use cases such as clinical reasoning, medical QA, note summarization, medical coding, patient communication, utilization management, and literature synthesis. • Design evaluations for retrieval-augmented and source-grounded health AI systems, including evidence citation, faithfulness, contraindication handling, guideline adherence, source freshness, and failure modes caused by incomplete, conflicting, or stale context. • Define sampling strategies, label schemas, inter-annotator agreement targets, adjudication workflows, SME review patterns, and quality thresholds in partnership with Language Data Scientists, clinicians, biomedical experts, and quality teams. • Build statistical and ML checks that make healthcare datasets trustworthy: stratified sampling across specialties and patient subgroups, bias and representation analysis, leakage detection, distribution shift checks, uncertainty estimates, reliability metrics, and subgroup performance analysis. • Partner with Applied Research Scientists and AI/ML Research Engineers to instrument datasets into evaluation and post-training pipelines, including rubric-grounded LLM-as-judge prompts, regression suites, model comparison workflows, experiment tracking, and model-improvement feedback loops. • Evaluate health AI behavior beyond surface accuracy: calibration, hallucination on safety-critical content, refusal appropriateness, robustness under ambiguity, equity across patient subgroups, and safe handoff in agentic or workflow-integrated systems. Reason concretely about clinical workflow fit: where outputs enter care delivery, what evidence a clinician or reviewer would need to trust them, when uncertainty must be surfaced, and how patient-facing, clinician-facing, payer, pharma, and operational use cases differ in risk. • Own data quality from source intake through delivery, including de-identified clinical text, medical literature, synthetic cases, structured records, client policies, and knowledge bases, with attention to PHI/PII handling, provenance, audit trails, versioning, and compliance documentation. • Stay current on the health AI landscape — regulatory developments such as FDA guidance on AI/ML-enabled medical devices and EU AI Act health provisions, benchmark releases such as MedQA, MedMCQA, and HealthBench, and emerging clinical evaluation methodology. • Support customer discovery and proposal work by scoping dataset programs, sizing annotation and SME review effort, identifying regulatory or data-access constraints, and explaining methodology choices to client clinical and ML leadership. • Contribute to Innodata internal IP: reusable health-domain taxonomies, evaluation rubrics, golden datasets, clinical review playbooks, dataset quality checks, and methodology templates.



