Empowering companies to find the specialized inputs they need to build the future of sustainable & branded products
Chemical Data Scientist
Location
California
Posted
2 days ago
Salary
0
Seniority
Senior
Job Description
Chemical Data Scientist
Valdera
• Design and build pipelines to collect supplier data and chemical product information (specifications, CAS numbers, certifications, SDS/regulatory documents, NAICS classification of manufacturing plants) from supplier sites, distributor catalogs, trade databases, and other public and semi-structured sources • Develop and maintain web scrapers and automated ETL workflows to keep supplier and product data current at scale • Clean, normalize, and reconcile inconsistent supplier data into structured, standardized formats suitable for internal tools and analytics • Apply chemical domain knowledge to validate and enrich data — resolving product names, CAS numbers, synonyms, and specifications across suppliers • Evaluate and improve matching and classification models to map suppliers and products to buyer requirements, and to identify overlapping or equivalent chemical offerings • Partner with Supplier Management and Engineering to define data quality standards, identify gaps in supplier coverage, and prioritize new data sources. • Own pipeline health and data quality, and drive the KPIs that measure overall data coverage
Job Requirements
- 5+ years of experience in a data science, data engineering, or applied data role, ideally with exposure to messy, real-world or industrial datasets.
- Working knowledge of chemistry or chemical industry data — comfort with CAS numbers, chemical properties, SDS documents, NAICS classification, and supplier certifications
- Strong Python skills, with experience building web scrapers and data pipelines
- Experience with data cleaning and normalization at scale, and a good eye for spotting inconsistencies in unstructured data
- Familiarity with building or applying matching, deduplication, or classification models (traditional ML or LLM-based approaches)
- Hands-on experience using AI tools and LLMs to accelerate data extraction, enrichment, or engineering workflows
- Startup mindset with a strong sense of ownership — comfortable working independently in a fast-moving, remote environment with ambiguous, evolving priorities
Benefits
- Valdera offers generous benefits to employees. You will be provided a more detailed breakdown of your options prior to joining Valdera.
Related Guides
Related Categories
Related Job Pages
More Data Scientist Jobs
Data Scientist I
HelioCampusTransforming institutional effectiveness for the most forward thinking institutions
Role Description Bringing experience analyzing student-level data at a Higher Ed institution or EdTech company in the areas of admissions, enrollment, financial aid, student success, institutional effectiveness, or finance and budget, you will step into a key role within HelioCampus’ Client Experience division. As a member of the DS-OPS (Data Science Operations, Products, and Services) team, you will partner directly with university leadership to turn "business questions" into "data questions," delivering analyses and predictive models that support institutional decision-making. In this role, you will work independently and collaboratively with engineers and clients to: - Design datasets - Perform statistical analysis - Build dashboards - Deploy machine learning models Managing your projects autonomously, you will use your strong communication skills to translate complex analytical results into clear, actionable insights for a variety of higher education stakeholders. Qualifications - Experience analyzing student-level data at a Higher Ed institution or EdTech company in the areas of admissions, enrollment, financial aid, student success, institutional effectiveness, or finance/budget. - 3+ years of experience delivering analytical insights to higher ed stakeholders, including building and evaluating machine learning models (python and scikit-learn experience required). - Analytical dataset design and feature engineering skills (SQL experience required). - Ability to conduct, interpret, and explain statistical analyses. - Experience using interactive data visualization and business intelligence tools (Tableau and/or Power BI) to design and publish interactive reports and dashboards, enabling data exploration and communication of analysis results. - Excellent communication and collaboration skills, and experience working with both business users and technical development teams, as well as presenting findings to decision-makers. - Experience working with student data from PeopleSoft, Banner, Colleague, Workday, Slate, Salesforce/TargetX, or other higher ed-specific data systems. - Ability to work effectively and independently in a remote role on an Eastern time zone business hour schedule, managing multiple priorities and meeting deliverable deadlines. - Understanding of model transparency & explainability concepts, and ethical issues in data science. - Familiarity with production data pipeline and model deployment and management. - Familiarity with relational database and data warehouse concepts. - Familiarity with a variety of machine learning methodologies, forecasting techniques, and generative AI (LLMs). - A tool-agnostic approach to data science: excitement for adopting new tools and techniques, while having solid fundamentals that underpin quick learning and high quality work delivery. Requirements - Advanced degree in an analytical/quantitative field (Comparable depth of experience can be substituted for quantitative field of study or advanced degree). - Experience working remotely with distributed teams and clients. - Experience with version control (git), software development processes, object oriented programming, and machine learning framework development. - MLOps experience deploying and maintaining production pipelines and models. - Experience applying a wide range of supervised and unsupervised machine learning and forecasting techniques using a variety of python packages. - Experience working with any of the following tools: jupyter notebooks, pandas, Docker, linux server/command line, EC2, S3, Amazon Redshift, airflow, SHAP, MLFlow, prophet, VS Code. - Experience using AI tools/services to increase delivery efficiency without sacrificing quality (for example - having built enough depth of coding experience to critically evaluate AI-generated code and avoid introducing issues that lead to future rework). - Experience with LLM Evaluation methods and frameworks. Benefits - Flexible work schedule - 18 PTO days on day one with additional PTO per year - Fifteen (15) paid holidays - Medical/dental/vision coverage - Parental leave - 401K with match - Company events - Referral bonus - Professional development opportunities - Home office perks - Two office locations Company Description - Commitment to Higher Education: We believe that serving higher ed makes the world a better place. This drives our desire to deliver high-quality products and services to help higher education institutions serve their students and communities. - Flexibility to Fit Your Life: HelioCampus employees enjoy the autonomy, freedom, and trust of a remote-first working environment. We embrace an “on it” mentality both in and out of the workplace, and employees are supported in pursuing personal and professional passions. - A Community of Bright Minds with Big Hearts: HelioCampus hires curious and talented practitioners who bring a diversity of thought and genuine care for one another and our clients to work every day. Our team members are more than a job description and are encouraged to show up as their full and authentic selves.
Lead Data Scientist – Growth & Experimentation
FullscriptDispense your way | Currently hiring across North America!
• Own experimentation (the heart of the role). Partner with Product and Commercial teams to design, run, and read growth experiments — from framing the hypothesis and sizing the test to calling the result and pushing the decision. Bring rigor (power analysis, causal inference, guardrail metrics) to a business moving fast. • Build the reporting backbone. Own the core dashboards and self-serve tooling the segment runs on, evolving an AI-augmented reporting stack that already automates much of the routine work. Your job is to make performance easy to understand for execs. • Generate insight, not just analysis. Break down revenue and margin performance into drivers - volume, mix, pricing, cohort behavior - and proactively tell us where growth is coming from, where it's leaking, and what to do next. • Keep the data foundation healthy. Manage the partnership with our central Data Engineering team to make sure the segment's core data models and pipelines support the analytics. • Sharpen the forecast. Support Strategic Finance with driver-based revenue forecasting analytics, connecting operational signals to financial outcomes.
Data Scientist Intern
CAREER PANACEAWe help Graduates & Skilled migrants get their Professional Job faster via our proven PROFESSIONAL INTERNSHIP PROGRAM.
• Collect, clean, and preprocess structured and unstructured data from multiple sources. • Perform exploratory data analysis (EDA) to identify trends, patterns, and anomalies. • Develop and test predictive models using statistical and machine learning techniques. • Assist in building data pipelines for data collection, transformation, and validation. • Create dashboards and visualisations using Power BI, Tableau, or Python libraries to communicate insights. • Write SQL queries to extract, manipulate, and analyse large datasets. • Support feature engineering and model optimisation to improve prediction accuracy. • Evaluate model performance using appropriate metrics and validation techniques. • Collaborate with cross-functional teams to understand business requirements and translate them into data-driven solutions. • Prepare reports and presentations summarising analytical findings and recommendations. • Document methodologies, code, and model assumptions to ensure reproducibility. • Assist in automating repetitive data processing and reporting tasks using Python or R. • Conduct research on emerging AI, machine learning, and data science techniques. • Ensure data quality, integrity, and compliance with organisational data governance standards. • Participate in team meetings, code reviews, and knowledge-sharing sessions.
Data Programs Manager
PAR TechnologyPAR Technology is a leading provider of systems, software, and service solutions to the retail and restaurant industries. The company is working to redefine the
• Own retailer onboarding from kickoff through go-live. • Manage weekly scan data submission health across the full retailer base. • Serve as the primary point of contact for all retailer and CPG partner communications. • Publish the weekly health dashboard for program performance. • Coordinate with Scan Data Engineering to ensure technical setup and test submissions are completed. • Monitor each submission cycle and communicate directly with retailers on action items.




