Job Closed
This listing is no longer active.
We're the modern way of proving identity.
Data Science Lead
Location
United States
Posted
63 days ago
Salary
$179K - $200K / year
Seniority
Lead
Job Description
Data Science Lead
Prove
Role Description The Data Science Lead will serve as the strategic architect and research pioneer for the organization’s data ecosystem. This role is responsible for designing robust data architectures, leading research and development (R&D) for novel data sources, establishing rigorous analytical methodologies, and ensuring the seamless, scalable ingestion of high-quality data into downstream production solutions. Core Pillars of Responsibility - Data Architecture & Scalable Engineering - Blueprint Design: Design and oversee the evolution of scalable data architectures that support advanced analytics, machine learning (ML) modeling, and real-time processing. - R&D & Novel Data Source Evaluation - Exploratory Research: Scout, evaluate, and pressure-test new internal, external, and alternative data sources (e.g., synthetic data, IoT streams, third-party APIs) for predictive power and commercial viability. - Proof of Concepts (PoCs): Lead rapid prototyping and PoCs to validate new technologies, algorithms, and data structures before scaling them to production. - Vendor & Partner Assessment: Technical vetting of data vendors and partners to ensure data quality, density, and seamless integration capabilities. - Methodology & Analytical Rigor - Framework Standardization: Define and document the organization's gold-standard methodologies for statistical analysis, experimental design (A/B testing), and ML modeling. - Evaluation Metrics: Establish rigorous validation protocols and evaluation metrics (e.g., precision/recall, drift detection, bias/fairness audits) to ensure model and data integrity. - Continuous Improvement: Keep the organization at the cutting edge of data science by translating academic research and emerging industry trends into practical business methodologies. - Ingestion & Solution Integration - Productionalization Bridge: Serve as the critical bridge between R&D and Production, ensuring that complex analytical models and data sources are seamlessly ingested into core business products and solutions. - API & Interface Design: Oversee data delivery contracts between the DS ecosystem and downstream software applications to ensure the creation of clean, well-documented APIs. Key Deliverables (First 12 Months) - Data Source Playbook: A formalized framework for scoring, vetting, and onboarding new data assets. - Methodology Registry: A centralized repository of approved statistical models, evaluation metrics, and ingestion protocols to ensure team-wide consistency. - Feature Importance Registry & Feature Engineering Roadmap: A centralized repository connecting current data sources to their product value and impact of removal and/or possible substitutes to the roadmap of how Prove can leverage the signals in new and differentiated ways. - Architectural Roadmap: A 12 month to 3-year vision aligning data science infrastructure with corporate scaling goals. Qualifications - 6+ years in Data Science/Data Engineering, with at least 2 years in a technical leadership or architectural role. - Technical Stack: Python, R, SQL, Cloud Platforms (AWS/GCP/Azure), Big Data tech (Spark, Kafka), Orchestration (Airflow), and MLOps tools. - Expertise: Deep understanding of data modeling, schema design (SQL/NoSQL), statistical evaluation, and MLOps deployment patterns, especially in R&D functions that bridge research with production. - Soft Skills: Exceptional ability to translate complex technical architectures into strategic business value for non-technical stakeholders. Benefits - Competitive salaries & Bonus Plan (for eligible roles) and Equity Plan - Modern Health for financial, mental, and physical wellness - 401(k) Retirement Plan & Match (US Offices) and Local Country Pension (International Offices) - Unlimited Vacation and Flexible hours - Comprehensive medical benefits for you and your family ❤️ - Emotional & Physical Wellness – Access to wellness services (EAP & Prove Well-Being Reimbursement) - Bottomless snacks & beverages for certain office locations - Daily GrubHub stipend for lunch if coming into the office (US Offices) - A great place to work and connect with other talented Provers like yourself!
Related Guides
Related Categories
Related Job Pages
More Data Engineer Jobs
• Profissional "hands-on" focado na implementação da arquitetura proposta, garantindo que o agente compreenda linguagem natural, acesse os dados necessários e mantenha a coesão da conversa. • Implementação da Orquestração: Desenvolver a camada de orquestração do agente utilizando LangGraph, construindo os grafos de estado e fluxos da conversa. • Desenvolvimento de Ferramentas (Tool Calling): Implementar e integrar chamadas estruturadas às APIs (tool calling) e sistemas disponibilizados pelo cliente, garantindo mecânicas de retry e fallback. • RAG e Base de Conhecimento: Construir pipelines de RAG (Retrieval-Augmented Generation) para que o agente consulte a base de conhecimento N1 e recupere contextos de forma otimizada. • Gestão de Estado: Implementar a persistência de estado (checkpointing) para suportar conversas de múltiplos turnos (multiturno) de forma fluida. • Implantação no GCP: Realizar o deploy, integração contínua e configuração dos recursos necessários no Google Cloud Platform (ex: Cloud Run, Vector Databases no GCP, Cloud Storage). • Transferência de Conhecimento: Apoiar ativamente a transferência de conhecimento técnico diário para o time do cliente.
Head of Data Engineering, Real-World Data
NateraWe are a global leader in cell-free DNA (cfDNA) testing, dedicated to oncology, women’s health, and organ health.
Natera is seeking a leader to head Data Engineering and Platform for Real-World Data (RWD). This role will in partnership with Product drive the strategy and roadmap for the RWD platform and lead the development, operation, and delivery of data products that support clinical, research, and business priorities. The ideal candidate brings strong engineering leadership, deep technical expertise, and experience building reliable, scalable data platforms in healthcare or life sciences environments. This leader will oversee a team focused on data engineering, data platform development, and data delivery, and will partner closely with product, bioinformatics, analytics, and AI teams to build data solutions that are useful, reliable, and easy to consume across the organization.
• Design and maintain event-driven pipelines from backend systems (Ruby on Rails) to Snowflake. • Own event pipelines from backend app through Snowplow ingestion. • Establish event tracking standards, documentation, and training for engineering teams. • Support feature flagging and A/B testing infrastructure requirements. • Optimize existing data workflows for efficiency and performance. • Manage Snowflake administration, performance optimization, and cost management alongside the data team. • Build reverse ETL processes and marketing automation data flows using tools like Census. • Collaborate to optimize Snowplow Analytics for event tracking and schema management. • Partner with analytics engineers on dbt deployment, testing, and production workflows. • Act as the technical bridge between engineering and analytics for pipeline questions. • Provide technical guidance to product and engineering teams on event implementation. • Lead training sessions and office hours on event tracking and troubleshooting. • Participate in architecture discussions and cross-team technical decision-making.
• Lead the design and development of gold-tier DBT models and BigQuery resources • Act as the senior technical authority on the team • Serve as the primary liaison between local and global engineering teams • Design, build, and optimize scalable ELT pipelines using Airflow, DBT, and BigQuery • Establish robust data quality frameworks and observability practices • Mentor data and analytics engineers




