Job Closed
This listing is no longer active.
Accelerate innovation for a healthier world.
Data Engineer
Location
United States
Posted
17 days ago
Salary
$69.3K - $173.2K / year
Seniority
Mid Level
Job Description
Data Engineer
IQVIA
Role Description The Data Engineer (Data Operations) is responsible for building, maintaining, and optimizing downstream data pipelines and analytics-ready datasets that power business intelligence and reporting. This role focuses on data troubleshooting, transformation, and delivery, ensuring that upstream data is validated and converted into reliable, curated data sources for BI and analytics teams. - Work hands-on with Snowflake, SQL, and Python to develop data models, resolve data issues, and support production data operations. - Enable the BI team to efficiently consume high-quality data through their visualization tools. Essential Functions: - Build and maintain curated data layers and data sources. - Develop and optimize SQL- and Python-based transformations in Snowflake. - Troubleshoot data quality issues originating from upstream pipelines, including identifying root causes and coordinating fixes where needed. - Validate incoming datasets and ensure accuracy, completeness, and consistency before they are surfaced to downstream users. - Design and manage analytics-ready data models and tables to support reporting and dashboards. - Serve as a key point of contact for reporting data layers ensuring data sources meet usability and performance needs. - Monitor data pipelines and datasets for failures, anomalies, and performance issues, and proactively resolve them. - Support production operations, including incident management, backlog prioritization, and SLA adherence. - Partner with upstream engineering teams to escalate and resolve ingestion-related issues, while not owning ingestion directly. - Improve processes for data validation, monitoring, and operational efficiency. Qualifications - Bachelor’s degree. - 3+ years of experience in data engineering, data operations, or backend data development. - Strong hands-on experience with: - Snowflake (data warehouse design, optimization, and management). - SQL (advanced querying, performance tuning, and data transformation). - Python (data processing, scripting, and automation). - Proven experience building and supporting production-grade data pipelines and workflows. - Strong experience with ETL/ELT tools, orchestration frameworks (e.g., Airflow, Atomic), or similar pipeline management tools. - Experience troubleshooting data issues, pipeline failures, and data inconsistencies, with strong root cause analysis skills. - Solid understanding of data modeling, data warehousing concepts, and data lifecycle management. - Familiarity with CI/CD, version control (Git), and production deployment practices. - Strong problem-solving skills and ability to work independently in fast-paced environments. - Excellent communication skills with the ability to explain data issues and solutions to both technical and non-technical stakeholders. - Experience working with healthcare, pharmaceutical, or life sciences data strongly preferred. Benefits - The potential base pay range for this role, when annualized, is $69,300.00 - $173,200.00. - The actual base pay offered may vary based on a number of factors including job-related qualifications such as knowledge, skills, education, and experience; location; and/or schedule (full or part-time). - Dependent on the position offered, incentive plans, bonuses, and/or other forms of compensation may be offered, in addition to a range of health and welfare and/or other benefits.
Related Guides
Related Categories
Related Job Pages
More Data Engineer Jobs
• Owning the design and development of robust dbt Core models that transform raw data into trusted, analytics‑ready datasets in Snowflake • Architecting scalable, high‑performance data models that support enterprise reporting, analytics, and AI use cases • Translating complex business and analytical requirements into efficient, well‑structured ELT solutions through close collaboration with BI, analytics, and business stakeholders • Embedding best practices in data quality, testing, documentation, and lineage to ensure transparency, reliability, and trust in our data ecosystem • Leveraging Python to support automation, data validation, orchestration, and performance monitoring across ELT pipelines • Monitoring, tuning, and optimizing Snowflake query performance and cost efficiency • Leading technical design discussions and contributing hands‑on to critical data initiatives • Serving as a technical lead and mentor, guiding other engineers and elevating standards across the full data transformation lifecycle • Providing thought leadership on modern data transformation patterns, tooling, and architecture to help shape enterprise data strategy • Supporting data governance and metadata enrichment initiatives in alignment with broader enterprise data goals
Data Governance – Platform Manager
LawnStarterLawnStarter is a lawn care application designed to provide users with high-quality landscaping services “at the click of a button” along with exceptional cu
• You'll be the first person at LawnStarter dedicated to data governance - the owner of whether our data can be trusted. • That means the quality and freshness of our source data, pipelines, and reports; the definitions behind our metrics; the standards behind our Segment event tracking; the health of our Lightdash workspace; the data feeding our machine learning models; and the security of the data itself. • This is a hands-on role. You'll work solo at first, with the Analytics team around you but nobody under you - building automation, writing checks, fixing what's broken, and putting processes in place that scale past you. If the scope grows the way we expect, this becomes the foundation of a team you'd build. • Data quality and freshness - automated monitoring across source data, pipelines, and reports; catching upstream schema and source changes before they break anything downstream; running incidents to resolution when they happen. • Data lineage and impact analysis - a living map from production source to warehouse model to dashboard, and the process that uses it: when a production change is proposed, its downstream impact on pipelines, metrics, and reports gets assessed before it ships, not discovered after. The end-state is data contracts with engineering, so breaking changes get caught in their workflow, not ours. • Lightdash - administration, workspace structure, permissions, and the rollout itself. Your job is to give the company self-serve autonomy while keeping the workspace tidy enough that people can find and trust what's there. Enablement is part of the deal - people follow standards they've been taught - and so is keeping queries fast and warehouse costs sane. • The semantic layer - we just shipped it for our most critical metrics: one governed definition per metric, in code. You'll extend definition and mapping to the rest and guard the layer against uncontrolled growth as it scales. • Event tracking governance - our governed Segment event catalog: reviewing new events against its standards, keeping it matched to what production actually sends, and evolving the guardrails (naming, property dictionary, drift detection) as tracking grows. • AI data readiness - AI agents query our warehouse every day through Brain, our internal AI toolkit. You'll govern what data AI tools can access and keep the warehouse AI-legible: documented, consistent, and safe for an agent to query and get the right answer. • Data security and privacy - access controls, PII handling and retention under US state privacy laws, and periodic reviews of who - and which AI tools - can see what. • The governance system itself - the documentation, ownership models, and review loops that keep all of the above running without heroics.
AWS Snowflake Data Architect, 12+ Years of Experience
3Pillar GlobalBuilding digital businesses, together.
• Design and implement scalable data platforms using Snowflake, Databricks, Delta Lake, and cloud technologies. • Build batch and real-time data pipelines using PySpark, Kafka, and Spark Structured Streaming. • Develop AI-ready data architectures supporting analytics, ML, LLMs, and RAG applications. • Design semantic models, data governance, metadata, and data lineage solutions. • Implement vector databases, embedding pipelines, and retrieval solutions for AI applications. • Build and manage ML/LLMOps pipelines, model deployment, monitoring, and CI/CD. • Ensure data security, RBAC, compliance, and governance across the platform. • Mentor engineering teams and define architecture best practices.
• Assist in developing and maintaining basic data ingestion and transformation pipelines (ETL/ELT) using PySpark and SQL. • Help monitor data pipelines and implement basic checks to ensure data reliability for internal consumers (such as Data Scientists). • Learn and assist in automating pipeline testing and deployment processes. • Work alongside data scientists and software engineers to understand and support integrated data flows. • Assist in documenting data schemas, pipeline architectures, and metadata cataloging.




