Fusemachines is a leading AI strategy, talent, and education services provider. Founded by Sameer Maskey Ph.D., Adjunct Associate Professor at Columbia University, Fusemachines has a core mission of democratizing AI. With a presence in 4 countries (Nepal, the United States, Canada, and the Dominican Republic) and more than 450 full-time employees, Fusemachines brings global AI expertise to transform companies worldwide. Founded in 2013, Fusemachines is a global provider of enterprise AI products and services, on a mission to democratize AI. Leveraging proprietary AI Studio and AI Engines, the company helps drive the clients’ AI Enterprise Transformation, regardless of where they are in their Digital AI journeys. With offices in North America, Asia, and Latin America, Fusemachines provides a suite of enterprise AI offerings and specialty services that allow organizations of any size to implement and scale AI. Fusemachines serves companies in industries such as retail, manufacturing, and government. Fusemachines continues to actively pursue the mission of democratizing AI for the masses by providing high-quality AI education in underserved communities and helping organizations achieve their full potential with AI.
Senior Data Engineer
Location
EST (UTC-5)
Posted
4 days ago
Salary
0
Seniority
Senior
Job Description
Senior Data Engineer
Fusemachines
Role Description This is a full-time, high-impact position for a Senior Data Engineer with expertise in Databricks, dbt, and Apache Airflow to support a critical CRM data architecture migration for a key client in the Life Sciences industry. - Join an urgent initiative to backfill key engineering capabilities and maintain momentum during an ongoing CRM system transition. - Migrate enterprise customer data from Veeva CRM to Salesforce Life Sciences Cloud, integrated with an underlying AWS S3 cloud environment and Databricks data warehouse. - Build out, configure, and redirect data ingestion pipelines out of Life Sciences Cloud into the data warehouse. - Implement dbt models and Airflow orchestrations to ensure complete data accuracy. Candidates must be able to operate strictly on US East Coast business hours (location is flexible across North America, LATAM, or remote with full Eastern Time overlap). Qualifications - 5+ years of hands-on data engineering experience with deep expertise in AWS, Databricks, dbt, and Apache Airflow. - Strong programming proficiency in Python/PySpark and Advanced SQL (complex joins, analytical window functions). - Hands-on expertise in Databricks platform architecture, Lakehouse implementation, Delta Lake, Unity Catalog, and cluster performance tuning. - Proven track record of architecting and executing migrations. - Demonstrated experience scaling platform performance. - Proven experience building scalable transformations pipelines using dbt for data transformation, testing, and documentation. - Solid background orchestrating complex workflow DAGs with Apache Airflow. - Experience working with AWS cloud infrastructure, specifically AWS S3 as an underlying data lake storage layer. - Hands-on experience developing integrations and data ingestion pipelines for CRM platforms, specifically Salesforce, Salesforce Life Sciences Cloud, and/or Veeva CRM. - Understanding of data structures, customer master data, and analytics workflows within the Life Sciences. - Deep understanding of SDLC/Agile and DevOps for CI/CD and artifact management. - Knowledge of AWS and Databricks security best practices and compliance standards. - Certifications Preferred: Databricks Certified Data Engineer Associate/Professional, Databricks Spark Developer, and major cloud certifications in AWS. Requirements - Ability to maintain 100% full working time overlap with US East Coast business hours (ET). - Flexible location (US, Canada, LATAM, or remote ET). Responsibilities - Architect, build, and deploy data integration pipelines connecting Salesforce Life Sciences Cloud to the client’s Databricks warehouse environment. - Execute pipeline modifications to transition legacy data feeds from Veeva CRM to Salesforce Life Sciences Cloud, updating warehouse models accordingly. - Write clean, modular dbt transformation models and organize end-to-end DAG execution using Apache Airflow. - Manage Delta tables and optimize Databricks clusters and AWS S3 storage for high performance and cost efficiency. - Implement data quality testing, schemas, and verification rules in dbt and Python to guarantee accurate data delivery. - Build and enforce proactive monitoring frameworks. - Work closely with project leads, solution architects, and technical stakeholders during US East Coast hours to ensure rapid iteration and goal completion. Equal Opportunity Employer Race, Color, Religion, Sex, Sexual Orientation, Gender Identity, National Origin, Age, Genetic Information, Disability, Protected Veteran Status, or any other legally protected group status.
Related Guides
Related Categories
Related Job Pages
More Data Engineer Jobs
Role Description We're seeking an experienced Data Engineer/ Sr Data Engineer / Lead Data Engineer with expertise in data engineering across major data platforms. The ideal candidate will have a strong background in Python, SQL, ETL, and data modeling, with experience in tools like Teradata, Informatica, Hadoop, Spark, PySpark, ADF, Snowflake, and Big Data. Cloud knowledge (AWS, Azure, or GCP) is a plus. The role requires a willingness to transition and upskill into Databricks & AI/ML projects. - Design, develop, and maintain large-scale data systems - Develop and implement ETL processes using various tools and technologies - Collaborate with cross-functional teams to design and implement data models - Work with big data tools like Hadoop, Spark, PySpark, and Kafka - Develop scalable and efficient data pipelines - Troubleshoot data-related issues and optimize data systems - Transition and upskill into Databricks & AI/ML projects Qualifications - Relevant years of experience in data engineering - Strong proficiency in Python, SQL, ETL, and data modeling - Experience with one or more of the following: - Teradata - Informatica - Hadoop - Spark - PySpark - ADF - Snowflake - Big Data - Scala - Kafka - Cloud knowledge (AWS, Azure, or GCP) is a plus - Willingness to learn and adapt to new technologies, specifically Databricks & AI/ML Requirements - Experience with Databricks - Knowledge of AI/ML concepts and tools - Certification in relevant technologies Benefits - Competitive salary and benefits - Opportunity to work on cutting-edge projects - Collaborative and dynamic work environment - Professional growth and development opportunities - Remote work opportunities & flexible hours
Data Engineer
Encora DigitalEncora, a leader in digital engineering, drives innovation by crafting cutting-edge, cloud-first, data-first, and AI-first solutions that redefine industries. S
Role Description Location: Brazil Job Mode: Full-time Work Mode: Work from home Essential Skills: - Desire to work at high level with stakeholders to devise, understand and communicate clearly requirements and architecture design for data platforms in an efficient manner. - Large Experience with architecture, governance, security, design, business mapping and understanding, performance and tuning for data lake, data warehouse and other data storage systems and transformation. - Experience with talking with customers and stakeholders and extracting business and technical requirements in high and low level. - Strong knowledge of Data frameworks: Hadoop (YARN, HDFS), Hive, Spark, Kafka, Pentaho, Airflow, AWS data tools. - Experience manipulating data using SQL, NoSQL and unstructured data sources, including metadata. - Strong knowledge of ETL frameworks and workflow processes (data ingestion, clean up and preparation). Highly Desirable Skills: - Experience with Python and Java for system administration. - Large experience with data modeling and data design. - System administration experience on Linux platforms. - Financial and/or banking marketing knowledge is a plus. - Experience with AWS platform. Additional Skills: - Machine learning algorithms and workflow. - Pipeline and workflow orchestration (Oozie, Luigi). - Experience with Kubernetes. Company Description Encora is the preferred digital engineering and modernization partner of some of the world’s leading enterprises and digital native companies. With over 9,000 experts in 47+ offices and innovation labs worldwide, Encora’s technology practices include: - Product Engineering & Development - Cloud Services - Quality Engineering - DevSecOps - Data & Analytics - Digital Experience - Cybersecurity - AI & LLM Engineering At Encora, we hire professionals based solely on their skills and qualifications, and do not discriminate based on age, disability, religion, gender, sexual orientation, socioeconomic status, or nationality.
Principal Data Operations Engineer
State of ColoradoThe State of Colorado is located in the Rocky Mountain region of the western United States. It entered the 100-year-old Union in 1876, earning the nickname "Cen
Role Description This position is term limited with an anticipated end date of approximately two (2) years from the date of hire. This position is eligible for State employee benefits and may be extended as the situation warrants. We are looking for a team player who is passionate about data and finding ways to improve solutions for a diverse customer base. As our new Principal Data Operations Engineer you will be responsible for: - Implementing designs and standards provided by the Data and Integrations Architect. - Providing platform and application administration, configuration management, end user management, security and application level patching and upgrades, as well as source code control, deployment and release management. - Providing operational support ranging from minor bug fixes to major enhancements. - Participating in planning for application replacement and modernization. - Collaborating across departments, establishing configurations and tools for efficient data and integration service consumption. - Driving continuous improvement and managing vendor interactions to achieve organizational goals. Some of the day-to-day opportunities include: - Consulting with Data and Integrations Architects, Principal Developers, and other Data Operations and OIT team members to maintain and enhance existing platforms. - Performing platform administration and support for applications within Data Operations. - Working with Data Architects, Data Engineers, and Integration Developers on data ingestion, transformation, and presentation tasks. - Establishing automation of manual processes, including code deployment and environment provisioning. - Acting as Tier-2 escalation point for on-call/break-fix efforts. - Working with SecOps resources to ensure network security policy is established consistently. - Collaborating with Business Analysts, Customers, Project Managers, and others to assist in the creation of estimates and timelines. - Performing coding or configuration management in accordance with standards and best practices. - Coordinating update releases and other system changes. - Organizing, building, and validating all segments of the code and configurations related to a specific build through CI/CD pipelines. - Ensuring application maintenance and configuration activities are consistent with established service portfolio policies. - Identifying and recommending changes to application and platform policies to improve service quality. Qualifications - A minimum of seven (7) years of experience as a data engineer, DevOps Engineer, or similar software engineering role. - A minimum of one year (1) of experience designing, building, implementing, and maintaining data and system integrations. - Experience with MS Azure DevOps CI/CD, Terraform, and Python. Requirements - Additional appropriate education will substitute for the required experience on a year-for-year basis. - Training or Certification related to the work assigned to the position will be assigned credit towards substitution for experience and/or education. - If the minimum qualifications include a degree requirement, additional appropriate paid or unpaid experience will substitute for the required education on a year-for-year basis. Benefits - Eligible for State employee benefits. - Support for a healthy balance of work and personal time.
Clinical Data Engineer
Oregon Health & Science UniversityWe are Oregon's only public academic health center. In addition to caring for patients, we lead groundbreaking research. We also train the next generation of health care professionals. As Portland's largest employer, we give you opportunities to learn and advance in a system of hospitals and clinics across Oregon and Southwest Washington. All are welcome. OHSU welcomes people of all ages, ethnicities, genders, national origins, religions and sexual orientations. We are striving to build an anti-racist, multicultural institution and encourage people with diverse backgrounds to apply. To request reasonable accommodation, contact askhr@ohsu.edu.
Role Description The Clinical Data Engineer sits on the clinical data team within the broader Clinical Business Intelligence unit, alongside analysts, engineers, administrators, and architects. This position builds new clinical data warehousing solutions, data transformations, and data integration assets, and supports the changes, enhancements, and maintenance of existing assets in support of OHSU clinical data initiatives. You will work closely with cross-functional teams — clinical and operational stakeholders, data architects, and IT specialists — to develop robust data pipelines, implement data quality controls, and deliver trustworthy data that supports clinical decision-making. Development happens primarily in the Epic Caboodle Console, Microsoft SQL Server tools, and Microsoft Fabric Data Engineering tools, with Azure DevOps for version control, code management, and deployment. Duties may extend to other cloud data engineering tools, such as Apache Airflow, as needed. - Design and develop ETL pipelines that extract, transform, and load clinical data from a variety of sources into structures suitable for analysis. - Implement data quality controls that validate, monitor, and maintain the accuracy, completeness, and consistency of clinical data across the pipeline lifecycle. - Contribute to the growth of the OHSU Caboodle Data Warehouse by designing, developing, testing, and implementing custom clinical data models. - Partner with BI architects, developers, analysts, and customers (practice managers, data scientists, quality analysts) to build and publish data models, ETL processes, Lakehouses, Warehouses, Notebooks, and metadata using the Epic Caboodle Console and third-party ETL tools. - Develop data feeds using SSIS or similar tools, ensuring appropriate security review and transport consistent with information privacy and security requirements, business associate agreements, and data use agreements. - Document warehouse content in the Caboodle Console and the Analytics Marketplace so users can determine what data exists, how it is defined, and how it traces back to Epic Clarity. - Troubleshoot ETL failures, data anomalies, and warehouse issues surfaced by automated monitoring, other developers, and end users; resolve or escalate through established processes and communicate status to affected groups. - Recommend improvements to ETL processes, tool sets, data models, and monitoring techniques that increase reliability and efficiency. - Deploy warehouse content using approved Azure DevOps systems and processes, and follow approved SDLC practices throughout. - Manage assigned projects by building timelines, identifying risks and milestones, and reporting status. - Respond to and track issues in Jira Service Desk, gathering information from customers and triaging to resolution. Qualifications - Bachelor’s degree in computer science, a related field, or a clinical field and six years of work-related experience in the information technology field or a combination of clinical or operational healthcare environments; - OR Associate’s degree in computer science, a related field, or a clinical field and seven years of work-related experience in the information technology field or a combination of clinical or operational healthcare environments; - OR Eight years work related experience in the information technology field or a combination of clinical or operational healthcare environments; - OR Equivalent combination of education and experience where one year of experience will be substituted for an Associate’s degree and two years of experience will be substituted for a Bachelor’s degree. Requirements - Minimum of two (2) years of experience as an Application Engineer or Developer (or equivalent classification) developing data warehouse objects and data integration ETL solutions. - Minimum of three (3) years SQL Server Experience, including SSIS and T-SQL, coding, performance tuning, and system optimization. - Minimum of five (5) years with Microsoft SQL Server T-SQL. - Minimum of two (2) years of experience in a medallion architecture data warehouse environment. - One year of experience with Microsoft Fabric using OneLake and Data Engineering tools. - One year of experience with Python or PySpark. Skills and Abilities - Knowledge of data warehousing architecture and dimensional modeling concepts. - Knowledge of data validation and testing methodologies for ETL processes. - Familiarity with data governance and cataloging practices. - Proven communication, analytical, and problem-solving skills. - Ability to manage competing priorities and communicate progress on an ongoing basis with excellent attention to detail. - Ability to accurately document system technical artifacts at a level of detail sufficient for ongoing production support. Certifications - Epic Clarity Data Model Certifications and Epic Caboodle Developer Certification within 6 months of hire. - Microsoft DP-700 Fabric Data Engineer certification within 9 months of hire. Preferred Qualifications - Experience with the Epic Clarity and Caboodle data models. - Experience planning and managing small projects. - Experience with HIPAA and PHI compliance. - Microsoft DP-700 Fabric Data Engineer Certification. Benefits - Healthcare for full-time employees covered 100% and 88% for dependents. - $50K of term life insurance provided at no cost to the employee. - Two separate above market pension plans to choose from. - Vacation - up to 200 hours per year dependent on length of service. - Sick Leave - up to 96 hours per year. - 9 paid holidays per year. - Substantial Tri-Met and C-Tran discounts. - Employee Assistance Program. - Childcare service discounts. - Tuition reimbursement. - Employee discounts to local and national businesses. Why apply to OHSU? We are Oregon's only public academic health center. In addition to caring for patients, we lead groundbreaking research. We also train the next generation of health care professionals. As Portland's largest employer, we give you opportunities to learn and advance in a system of hospitals and clinics across Oregon and Southwest Washington. All are welcome. OHSU welcomes people of all ages, ethnicities, genders, national origins, religions and sexual orientations. We are striving to build an anti-racist, multicultural institution and encourage people with diverse backgrounds to apply. To request reasonable accommodation, contact askhr@ohsu.edu.
