Job Closed
This listing is no longer active.
Focused patient recruitment.
Statistical Data Scientist
Location
United States
Posted
96 days ago
Salary
$180K - $200K / year
Seniority
Lead
Job Description
Statistical Data Scientist
Praxis
• Lead the design, development, and validation of R/Python code to automate generation of analytical datasets and TLFs. • Translate SAPs and metadata specifications (YAML/CSV) into executable and reproducible code. • Build and validate R packages and data science tools supporting both exploratory and confirmatory analyses. • Implement and validate statistical models (e.g., MMRM, ANCOVA, logistic regression). • Collaborate with IT to integrate data science workflows within Databricks and CI/CD pipelines. • Collaborate across programming, biostatistics, and data standards functions. • Conduct peer code reviews, unit testing, and automated validation. • Mentor and guide team members in best practices for programming, validation, and automation.
Job Requirements
- Bachelor’s or Master’s degree in Statistics, Biostatistics, Data Science, or a related field.
- 8+ years of statistical programming experience in the pharmaceutical/biotech industry including hands-on experience with R and/or Python.
- Proven experience preparing or supporting R-based regulatory submissions.
- Strong understanding of CDISC ADaM and SDTM data structures, and their use in analytical workflows.
- Experience developing and validating reusable R/Python libraries and functions.
- Proficiency with Git, Bitbucket, and CI/CD automation pipelines.
- Working knowledge of GxP and Part 11 compliance.
- Excellent documentation and validation practices.
- Collaborative and proactive mindset; able to operate independently in a small, agile team.
- Familiarity with YAML/JSON configuration and metadata-driven programming workflows (preferred).
- Prior experience migrating from SAS to R/Python (preferred).
- Knowledge of R validation frameworks (e.g., risk-based testing, reproducibility documentation) (preferred).
- Experience with exploratory analytics or visualization in R or Python within a regulated framework (preferred).
Benefits
- 99% of the premium paid for medical, dental and vision plans.
- Company-paid life insurance.
- AD&D and disability benefits.
- Voluntary plans to personalize your coverage.
- 401(k) matching up to 6% on eligible contributions.
- Long-term stock incentives and ESPP.
- Discretionary quarterly bonus.
- Flexible wellness benefit.
- Generous PTO and paid holidays.
- Company-wide shutdowns.
Related Guides
Related Categories
Related Job Pages
More Data Scientist Jobs
• Scope and co-develop production-level data science projects with our customers across different industries and use cases • Help users discover and master the Dataiku platform via user training, office hours, and ongoing consultative support • Provide strategic input to the customer and account teams that help make our customers successful. • Provide data science expertise both to customers and internally to Dataiku’s sales and marketing teams • Lead technical data science projects pre-sales scoping and design appealing proposals • Flag technical and non-technical account risks (onboarding issues, performance pitfalls, timeline slippage) • Develop custom Python-based “plugins” in collaboration with Solutions, R&D, and Product teams, to enhance Dataiku’s functionality • Lead Data Scientist engagements: You will coordinate agile sprints, prioritize tasks, estimate effort, do backlog grooming • Run demo booth/tech talk duties at company public events (e.g. Everyday AI) • Lead Junior Data Scientist technical interview • Contribute to 2 internal assets (internal best practice or external blog post/project on the public gallery) per year
Senior Data Scientist, Consultant
GuidehouseGuidehouse, a "next-generation consultancy" and a portfolio company of Veritas Capital, provides management, risk consulting, and technology services to help cl
• Help clients maximize the value of their data across the full analytics lifecycle, including data querying and wrangling, data visualization and business intelligence (BI), predictive analytics, machine learning, and artificial intelligence • Support clients in defining information strategy, data architecture, and data governance • Implement enterprise analytics and data management solutions that enable actionable insights, reduce cost and complexity, increase trust and data integrity, and improve operational effectiveness • Lead and coordinate internal teams to deliver high‑quality, client‑focused solutions • Engage directly with clients and communicate confidently about data products, analytics methodologies, and industry trends • Operate effectively in a high‑visibility, client‑facing role within a dynamic consulting environment
It's fun to work in a company where people truly BELIEVE in what they're doing! Fullsteam is a leading provider of vertical software and embedded payments technology dedicated to helping businesses flourish by providing their customers with seamless experiences. With a dynamic and growing team of over 1,900 employees, we are committed to driving innovation and delivering best-in-class software and payment solutions that empower small and medium-sized businesses across numerous industries. Our purpose is to help our customers grow their businesses and delight their customers. Join us and be a part of a forward-thinking company that values growth, excellence, and the success of our clients. Data Science, Machine Learning & AI Intern – PIC Business Systems Business Unit Overview: PIC Business Systems is the leading provider of web-based ERP software solutions for the window covering manufacturing industry, commanding over 50% market share with a perfect implementation track record spanning 35+ years. Our flagship product, e-PIC One Enterprise, serves industry leaders including Hunter Douglas, Budget Blinds (800+ franchises via HFC), and Springs Window Fashions. PIC was acquired by Fullsteam in January 2025 and operates within the Home ERP group, bringing enterprise-grade capabilities to a stable, profitable business generating approximately $10.8M in annual revenue. PIC is at the forefront of AI adoption within Fullsteam, having deployed Claude Code across all 14 engineers (achieving a documented 33:1 ROI) and built PICasso—a custom AI-powered Support Assistant with 40+ tools built on Strands Agents SDK and FastMCP, deployed on AWS ECS Fargate and leveraging AWS Bedrock for multi-model agent orchestration. The team of 25+ employees spans engineering, support, implementation, and operations, operating in a fully remote environment with a culture defined by four core values: Win Together, Embrace Change, Don’t Report the News, and Own the Outcome. Job Summary: The Data Science, Machine Learning & AI Intern will be PIC’s first dedicated data science and AI agent development role, responsible for developing predictive models, building analytics pipelines, creating and enhancing AI agents on AWS Bedrock, and embedding ML-driven insights into PIC’s products and operations. Reporting directly to the President, this intern will have access to rich multi-tenant ERP datasets spanning hundreds of window covering manufacturers and hands-on involvement with PICasso’s production AI agent infrastructure, providing a unique opportunity to work across the full spectrum of applied AI—from statistical modeling to autonomous agent development—in an industry-specific SaaS context. This role is ideal for a graduate student or advanced undergraduate in statistics, data science, machine learning, or a related quantitative field who wants hands-on experience building production ML systems and AI agents in a real business environment. The intern will contribute to PIC’s AI strategy while gaining exposure to enterprise SaaS operations, manufacturing domain expertise, and cloud infrastructure on AWS. Primary Responsibilities: - Build, maintain, and enhance AI agents using AWS Bedrock, including designing agent tool schemas, implementing post-condition validation, and optimizing agent performance across PICasso’s multi-agent architecture - Develop, test, and deploy new PICasso agent capabilities using Strands Agents SDK and FastMCP, contributing to the 40+ tool ecosystem that powers customer support operations - Design, train, and validate machine learning models for customer attrition prediction, leveraging ERP usage patterns, support ticket history, billing data, and engagement signals across PIC’s multi-tenant MySQL databases - Build supply and demand forecasting models for window covering manufacturing, incorporating seasonality, material pricing trends, and historical order data to improve production planning accuracy - Develop and maintain ETL data pipelines to extract, transform, and load data from Aurora MySQL multi-tenant databases into analytics-ready formats - Create interactive customer analytics dashboards and PicRite reports that surface actionable insights for account management, customer success, and executive leadership - Integrate ML-driven pricing validation logic to detect anomalies and optimize franchise pricing across large dealer networks (e.g., Budget Blinds) - Conduct exploratory data analysis to identify patterns, trends, and opportunities within PIC’s extensive ERP dataset - Participate in agent quality audits (review_low_scores) and contribute to reducing hallucination rates and improving agent accuracy - Document all models, agents, assumptions, data sources, and methodologies; present findings to engineering leadership and stakeholders - Collaborate with engineering, support, and implementation teams to understand business context and ensure model and agent relevance Skills & Competencies: - Strong foundation in statistics, probability, and machine learning algorithms (regression, classification, clustering, time series forecasting) - Proficiency in Python with data science libraries (pandas, NumPy, scikit-learn, TensorFlow/PyTorch) - Experience with SQL and relational databases; ability to write complex analytical queries against large datasets - Familiarity with LLMs, prompt engineering, and AI agent frameworks (experience with AWS Bedrock, LangChain, or similar a plus) - Familiarity with data visualization tools and libraries (Matplotlib, Seaborn, Plotly, or similar) - Understanding of ETL pipeline design and data engineering fundamentals - Strong analytical thinking with the ability to translate business problems into data science and AI solutions - Excellent written and verbal communication skills; ability to explain technical findings to non-technical stakeholders - Self-motivated with the ability to work independently in a remote environment - Intellectual curiosity and willingness to learn domain-specific knowledge in manufacturing and ERP systems Minimum Qualifications: - Currently pursuing or recently completed a Master’s degree in Statistics, Data Science, Computer Science, Machine Learning, Mathematics, or a related quantitative field - Coursework or project experience in machine learning, statistical modeling, or predictive analytics - Proficiency in Python and SQL - Experience with at least one ML framework (scikit-learn, TensorFlow, PyTorch, or similar) - Portfolio or academic projects demonstrating applied data science or AI work (Kaggle competitions, research projects, agent prototypes, or capstone work acceptable) Preferred Qualifications: - Experience with AWS services (Bedrock, SageMaker, Glue, Athena, ECS, or similar) - Hands-on experience building or working with LLM-based agents, tool-use patterns, or multi-agent architectures - Graduate-level coursework or research in machine learning, NLP, or time series analysis - Familiarity with MySQL/Aurora MySQL in production environments - Exposure to multi-tenant SaaS data architectures - Experience with version control (Git/GitHub) and collaborative development workflows - Experience with MCP (Model Context Protocol), FastMCP, or similar agent tooling frameworks - Knowledge of manufacturing, supply chain, or ERP domains Fullsteam supports an inclusive workplace that values diversity of thought, experience, and background. Fullsteam is an Equal Opportunity/Affirmative Action employer. All qualified applicants will receive consideration for employment without regard to race, religion, color, national origin, ancestry, age, physical or mental disability, sex, sexual orientation, gender identity/expression, pregnancy, veteran status, marital status, creed, status with regard to public assistance, genetic status or any other status protected by federal, state, or local law.
Location: Work from home (Pennsylvania) Shift: Days (United States of America) Scheduled Weekly Hours: 40 Worker Type: Regular Exemption Status: Yes Job Summary: The AI Data Scientist Team Lead (Manager, AI Platform Engineering) architects end-to-end AI solutions and leads the AI Platform team for Geisinger's AI Department. This is a hands-on technical leadership role, splitting time equally between solution architecture and engineering management (50% technical / 50% leadership). On the technical side, the Team Lead serves as the solution architect across the AI Platform portfolio: gathering requirements from clinical informaticists, data scientists, and business stakeholders; designing production-grade AI architectures spanning batch and real-time workloads; and making build-vs-buy calls for emerging AI capabilities. On the management side, the Team Lead runs the team's rituals, removes blockers, develops direct reports, and manages stakeholder expectations. The AI Platform team is an enabling team—not a delivery team—that builds the reusable capabilities, tooling, and infrastructure that let product teams deploy AI safely and quickly. The team consists of 8 engineers across 6 distinct roles (4 direct reports + 3 matrixed engineers from partner departments), currently supporting 10 platform capabilities serving 70 AI programs. The Team Lead owns the team's capability roadmap, capacity allocation, platform engineering standards, and architecture reviews, while translating organizational AI strategy into executable technical plans that deliver production-grade capabilities across the portfolio. Job Duties: What You Will Own: - Solution architecture across all platform capabilities (agentic AI systems, RAG pipelines, multi-model orchestration, real-time and batch ML infrastructure) - Requirements gathering and technical specification for AI programs across clinical and operational domains - Build-vs-buy and technology selection decisions for emerging AI capabilities, including generative AI, foundation models, and LLM applications - Platform engineering standards, architecture reviews, and governance compliance (HIPAA, AI risk management, responsible AI principles) - Team roadmap, capacity allocation, and intake triage for platform support requests - People management, career development, and performance evaluation for 4 direct reports (3 MLOps Engineers, 1 Full Stack Engineer) - Work direction, priorities, platform standards, and formal performance input for 3 matrixed engineers from partner departments (Sr. Platform Data Engineer, Sr. Software Engineer for Integration & Interfaces, Sr. Platform Engineer) What You Will Not Own: - Individual capability delivery (delegated to the team via RACI) - Product strategy or portfolio prioritization (owned by the AI Product Management function) - Discipline-specific technical standards (set department-wide by the MLOps and Data Science Technical Discipline Leads; set by home-department tech leads for matrixed engineers) - HR management or final performance evaluations for matrixed engineers (owned by their home departments) - Day-to-day Databricks workspace administration (owned by the Sr. Platform Data Engineer) Solution Architecture Responsibilities (50% Technical): - Design scalable AI architectures spanning batch and real-time workloads, ensuring solutions are production-grade, maintainable, and aligned with organizational priorities - Gather and refine requirements from clinical informaticists, data scientists, and business stakeholders; translate complex needs into actionable technical specifications - Architect agentic AI systems, RAG pipelines, and multi-model orchestration frameworks across clinical and operational domains - Serve as technical authority on end-to-end AI pipeline design across Databricks, cloud-native platforms, and Epic integration points - Drive build-vs-buy and technology selection decisions for emerging AI capabilities (generative AI, foundation models, LLM applications) - Ensure AI systems adhere to healthcare security standards (HIPAA), AI governance frameworks, and responsible AI principles - Partner with data architects and governance teams to enforce data quality, lineage, and access controls across AI data assets Engineering Management Responsibilities (50% Leadership): - Lead multiple concurrent AI projects; manage scope, timelines, and technical risk while removing obstacles for the team - Mentor and develop 4 direct-report engineers; provide technical leadership and formal performance input for 3 matrixed engineers - Establish platform engineering best practices, conduct architecture reviews, and foster engineering excellence across the full team - Align technical execution with strategic goals; contribute data-driven insights to inform organizational AI initiatives - Coordinate cross-functional collaboration between the AI Platform team and data scientists, software engineers, clinical informaticists, and business stakeholders - Champion scalable and governed AI practices across the organization - Run team rituals (daily standups, weekly planning, architecture office hours, biweekly demos, monthly capability health reviews, quarterly roadmap refresh) How the Role Operates: - Prioritization: The Team Lead owns the team's roadmap, balancing strategic alignment (capabilities that unblock the highest-value portfolio initiatives), breadth of impact (work that benefits the most programs wins over single-program requests), and operational urgency (production incidents, security issues, governance blockers jump the queue) - Intake: Product teams request platform support through a lightweight intake process the Team Lead manages; requests are triaged weekly—absorbed into the roadmap, handled as quick-turn asks, or redirected to self-serve documentation - Matrix management: For direct reports, owns the full management stack (roadmap, career development, performance, HR). For matrixed engineers, owns the work (roadmap, priorities, platform standards, architecture reviews) and provides formal input on performance reviews; the engineer's home department owns HR management and final evaluation - Escalation path: Engineer-level issues resolved directly between engineers; priority conflicts, scope disagreements, and technical decisions with broad impact come to the Team Lead; strategic trade-offs and cross-department conflicts escalate to the VP Work is typically performed in an office or remote environment. Accountable for satisfying all job specific obligations and complying with all organization policies and procedures. The specific statements in this profile are not intended to be all-inclusive. They represent typical elements considered necessary to successfully perform the job. *Relevant experience may be a combination of related work experience and degree obtained (Master's Degree = 2 years; PHD = 4 years ). Position Details: Key Technologies: - Databricks (Delta Lake, Unity Catalog, MLflow, Mosaic AI, Spark) - AWS (ECS/Fargate, Bedrock, S3, IAM), Terraform - Claude / Amazon Bedrock, LangChain, agentic AI frameworks - Epic APIs (FHIR, SDE) - Docker, CI/CD pipelines, MLOps tooling - Real-time streaming (Kafka, Spark Structured Streaming) Collaboration Points: - All AI Platform team roles: direct manager, solution reviewer, escalation point - Clinical informaticists and data scientists: requirements gathering and solution design - AI Product Management: roadmap alignment and portfolio prioritization - AI Department Technical Discipline Leads (MLOps, Data Science): alignment on discipline-specific standards applied to platform work - AI Governance: compliance with risk frameworks, responsible AI principles, and model risk management - Enterprise architecture and security: alignment of AI Platform infrastructure with organizational standards - Partner department managers (IT Platform, IT Software, CDIO Data Management): matrix coordination for matrixed engineers Required Skills & Qualifications: - 8+ years in data science, ML engineering, or AI solution architecture, with at least 3 years in a technical leadership or engineering management role - Demonstrated experience designing production ML/AI systems end-to-end: from data ingestion through model serving and monitoring - Strong fluency in Python and SQL; hands-on experience with Databricks (MLflow, Unity Catalog, Spark) and cloud-native ML infrastructure (AWS preferred) - Experience architecting agentic AI systems, LLM applications, or RAG pipelines in production settings - Proven ability to translate ambiguous business problems into technical specifications and actionable engineering plans - Track record of mentoring engineers across multiple specialties and managing concurrent technical projects - Familiarity with healthcare data standards (HL7/FHIR) and regulatory requirements (HIPAA) strongly preferred - Experience with Epic integration points (FHIR, SDE) a plus - MS or PhD in Computer Science, Data Science, or related quantitative field preferred; equivalent experience accepted Education: Bachelor's Degree- (Required) Experience: Minimum of 6 years-Relevant experience* (Required) Certification(s) and License(s): Skills: Analyzing, processing and building AI/ML solutions from Clinical and Operational data sources, such as Epic Clarity, HL7, DICOM, or ECG data, Clinical Databases, Communication, Critical Thinking, Data Analysis, Data Presentations, Group Collaboration, Leadership, Machine Learning Methods, Programming Languages, Structured Query Language (SQL) OUR PURPOSE & VALUES: Everything we do is about caring for our patients, our members, our students, our Geisinger family and our communities. - KINDNESS: We strive to treat everyone as we would hope to be treated ourselves. - EXCELLENCE: We treasure colleagues who humbly strive for excellence. - LEARNING: We share our knowledge with the best and brightest to better prepare the caregivers for tomorrow. - INNOVATION: We constantly seek new and better ways to care for our patients, our members, our community, and the nation. - SAFETY: We provide a safe environment for our patients and members and the Geisinger family. We offer healthcare benefits for full time and part time positions from day one, including vision, dental and domestic partners. Perhaps just as important, we encourage an atmosphere of collaboration, cooperation and collegiality. We know that a diverse workforce with unique experiences and backgrounds makes our team stronger. Our patients, members and community come from a wide variety of backgrounds, and it takes a diverse workforce to make better health easier for all. We are proud to be an affirmative action, equal opportunity employer and all qualified applicants will receive consideration for employment regardless to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or status as a protected veteran.




