C the Signs is a cancer prediction system that identifies patients at risk of cancer at the earliest, most curable stage
AI Data Engineer
Location
New Hampshire + 4 moreAll locations: New Hampshire | New Jersey | New York | Massachusetts | Rhode Island
Posted
98 days ago
Salary
0
Seniority
Senior
Job Description
AI Data Engineer
C the Signs
• Collaborate with data scientists and machine learning engineers to understand data requirements for LLM and machine learning model fine-tuning. • Design, build, and maintain scalable data pipelines to ingest, process, and store massive and diverse healthcare datasets. • Implement robust data validation and monitoring to ensure the integrity, accuracy, and consistency of all training datasets. • Implement robust data cleaning, validation, and transformation processes to ensure data quality and integrity. • Develop and optimize data structures and schemas for efficient access and utilization by LLMs and machine learning models. • Work with the team to identify and acquire new data sources, ensuring compliance with relevant healthcare regulations (e.g., HIPAA). • Monitor data pipeline performance, troubleshoot issues, and implement optimizations to improve efficiency and reliability. • Document data engineering processes, data models, and data dictionaries. • Stay up-to-date with the latest advancements in data engineering, big data technologies, and machine learning.
Job Requirements
- Required
- Bachelor's degree in Computer Science, Engineering, or a related field.
- Proven experience as a Data Engineer, with a focus on big data technologies.
- Strong proficiency in programming languages such as Python, Scala, or Java.
- Extensive experience with data warehousing, ETL processes, and data modeling.
- Experience with major cloud providers (e.g., AWS, GCP, Azure) and their data storage and processing services.
- Hands-on experience with big data frameworks like Apache Spark for distributed processing.
- Excellent problem-solving skills and the ability to work independently and as part of a team.
- Strong communication and interpersonal skills.
- Preferred
- Master's degree in a related field.
- Experience with healthcare data and a good understanding of healthcare data standards (e.g., FHIR, HL7).
- Familiarity with machine learning concepts and LLM fine-tuning processes.
- Experience with data orchestration tools (e.g., Apache Airflow).
- Work Authorization:
- Must be a US Citizen, Green Card holder, or currently in the US have valid H1B visa
Benefits
- Why Join Us?**
- Joining **C the Signs** is not just about building AI; it’s about shaping the future of healthcare. If you are a technical leader with an unshakable belief in the power of AI to save lives and the ability to make it happen at scale, this is your opportunity to create a tangible, global impact.
- Benefits:**
- Competitive salary and benefits package.
- Flexible working arrangements (remote or hybrid options available).
- The opportunity to work on life-changing AI technology that directly impacts patient outcomes.
- Join a team that combines cutting-edge innovation with a mission to save lives and improve health equity.
- Continuous learning opportunities with access to the latest tools and advancements in AI and healthcare.
Related Guides
Related Categories
Related Job Pages
More Data Engineer Jobs
• Diseño e implementación de pipelines de datos: Construir y optimizar procesos ETL/ELT en entorno GCP, utilizando BigQuery y herramientas de orquestación, garantizando flujos eficientes, escalables y fiables. • Migración y transformación de datos: Adaptar y trasladar lógica existente desde entornos como Talend/MySQL hacia SQL nativo en BigQuery, mejorando rendimiento y mantenibilidad. • Arquitectura cloud de datos: Diseñar soluciones de procesamiento y almacenamiento en la nube (GCP principalmente, valorable AWS/Azure), asegurando escalabilidad, seguridad y disponibilidad. • Calidad e integridad del dato: Aplicar buenas prácticas de validación, limpieza y monitorización para garantizar datos precisos, consistentes y confiables. • Innovación tecnológica: Explorar e incorporar nuevas herramientas, frameworks y metodologías que optimicen la ingesta, transformación y análisis de datos. • Documentación y buenas prácticas: Documentar pipelines, modelos y procesos para asegurar su comprensión, mantenimiento y reutilización dentro del equipo. • Integración de sistemas: Conectar distintas fuentes de datos (internas y externas) mediante APIs y otros mecanismos, garantizando flujos eficientes e interoperables. • Colaboración con equipos de analítica y BI: Facilitar datos estructurados y de calidad para su explotación en dashboards, análisis y modelos avanzados.
Title: Data Engineer II Location: Remote - Ohio Job Description: Based in St. Louis, Core & Main is a leader in advancing reliable infrastructure™ with local service, nationwide®. As a specialty distributor with a focus on water, wastewater, storm drainage and fire protection products and related services, Core & Main provides solutions to municipalities, private water companies and professional contractors across municipal, non-residential and residential end markets, nationwide. With over 370 locations across the U.S., the company provides its customers local expertise backed by a national supply chain. Core & Main’s 5,700 associates are committed to helping their communities thrive with safe and reliable infrastructure. Job Summary Design, build, test, and maintain reliable data pipelines and data solutions that support analytics, reporting, and operational use cases. This role owns end‑to‑end data pipelines or data flows within defined domains and works independently on moderately complex data engineering tasks. Partner closely with senior engineers, architects, and stakeholders to deliver high‑quality, secure, and well‑documented data solutions that meet business and technical requirements. Major Tasks, Responsibilities and Key Accountabilities - Design, develop, and maintain production‑ready data pipelines, data transformations, and data models. - Collaborate with business and technical stakeholders to understand data requirements and translate them into scalable solutions. - Implement and maintain data warehouse, data lake, or lakehouse solutions aligned with architectural standards. - Perform unit testing and support QA, regression, and user acceptance testing for data solutions. - Troubleshoot, debug, and resolve data pipeline failures, performance issues, and data quality defects. - Contribute to technical documentation, including design artifacts, data mappings, and operational runbooks. - Participate in peer code reviews and apply feedback to improve code quality and maintainability. - Support enhancements, patches, and upgrades to existing data platforms and tooling. Preferred Qualifications - Bachelor’s degree in Computer Science, Information Technology, or related field. - 4 years of hands-on development experience in SQL and/or Python for data warehouse management, data integration, and data lake management. - Deep working knowledge in SQL development using T-SQL code to design, implement, and optimize complex database objects, such as tables, views, stored procedures, indexes, and functions. - Experience working with Azure data architecture, including a solid understanding of tools for building data pipelines on cloud-based data platforms, such as Delta Lakehouse Medallion architecture and data warehousing solutions. - Exposure to modern Spark-based data platforms like Databricks or Microsoft Fabric for data engineering tasks, including leveraging their capabilities for scalable data processing, analytics, and machine learning workflows in a cloud-based environment. - Understanding of ELT vs ETL and how to build efficient data pipelines with modern Change Data Capture processes. - Hands-on experience with CI/CD pipelines in Azure DevOps and understanding of Agile development methodologies. - Familiarity with common data mapping and transformation techniques for Dynamics 365 Data Entities and Data Management Framework for the Finance and Operations modules. - Familiarity with Power BI and its integration with Microsoft Fabric for end-to-end analytics. - Strong communication skills with the ability to translate complex technical concepts into business-friendly language. Career Level Dimensions Typical Training/Experience - Typically requires BS/BA in a related discipline. Generally, 3-5 years of experience in a related field; certification is required in some areas OR MS/MA and generally 2+ years of experience in a related field. Problem Complexity - Applies established problem‑solving skills to moderately complex situations. Identifies root causes for common and recurring data issues and escalates more complex or ambiguous problems appropriately. Troubleshoots and resolves issues within defined data pipelines, systems, or domains, using documented patterns and guidance from senior team members. Autonomy - Performs assignments with moderate independence, operating within established practices and architectural standards. - Determines appropriate approaches to solutions for well‑defined problems. - Receives regular technical guidance on complex problems, design decisions, or unfamiliar technologies. Collaboration - Works closely with other engineers, analysts, and business partners to deliver reliable, high‑quality data solutions. - Actively participates in knowledge sharing, documentation, and peer reviews. - May provide informal guidance to less experienced engineers but does not have formal leadership or people management responsibilities. Core & Main is an Equal Employment Opportunity employer. Employment at Core & Main is based solely on a person’s merit and qualifications directly related to professional competence. Core & Main does not discriminate against any employee or applicant on the basis of race, creed, color, religion, national origin, nationality, ancestry, age, disability, veteran status, pregnancy or related condition (including breastfeeding), affectional or sexual orientation, gender identity or expression, marital status, status with regard to public assistance, citizenship, or any other basis protected by law. None of the questions in this application are intended to elicit information regarding any protected characteristics, nor imply any limitation, illegal preferences or discrimination based upon non-job-related information or protected characteristics.
Data Engineer
G2i Inc.G2i is a hiring platform run by engineers that match you with pre-vetted React and React Native engineers.
• Own the data stack end-to-end: ingestion → transformation → modeling → serving → monitoring • Build and maintain ETL/ELT pipelines from APIs, webhooks, and operational systems • Design resilient data models that handle evolving and imperfect source systems • Implement monitoring and alerting for data quality, freshness, and pipeline failures • Ensure high reliability and observability across the data layer • Lead improvements, migrations, and infrastructure decisions • Collaborate with engineering leadership on architecture, including modern data access patterns
Data Engineer I
eSimplicityAn engineering firm that delivers high-quality Healthcare IT, Cybersecurity, and Telecommunication solutions.
• Develop production-grade ETL workflows using Python and Microsoft-based frameworks to ingest, transform, and validate large-scale structured and unstructured data. • Implement schema enforcement, data validation, and quality checks to maintain integrity across diverse sources. • Optimize pipelines for performance, scalability, and fault tolerance using open-source and cloud-native patterns. • Manage Azure-based data solutions, including Data Lake Storage, Azure SQL, and cloud storage access from Python services. • Deploy workflow orchestration using Azure Data Factory or Foundry for scheduling, monitoring, and automation. • Ensure secure integration of APIs and services within the Microsoft ecosystem for seamless data exchange. • Build Python-based data services leveraging libraries such as Pandas, Pytorch, and other open-source frameworks for high-performance processing. • Implement logging, monitoring, and performance tuning for robust operational reliability. • Develop API endpoints and microservices to enable interoperability with analytics and ML platforms. • Work closely with data scientists, analysts, and cloud architects to deliver clean, reliable data for predictive modeling and real-time dashboards. • Apply data governance best practices, ensuring compliance, reproducibility, and auditability across workflows. • Contribute to Agile team processes, driving iterative improvements and shared problem-solving. • Work in Agile teams; drive iterative delivery, joint problem-solving, and continuous improvement. • Engage closely with project managers, technical leads, client representatives, and cross-functional teams to provide timely updates, resolve issues, and ensure alignment with business goals. • Translate technical specifications into code and design documents.




