Job Description
Principal Data Engineer
Gather AI
• Architect a greenfield, multi-layer data warehouse (raw, refined, serving) that separates analytical workloads from production OLTP traffic. • Deliver a governed, self-service data-access layer for internal consumers first (Product, CSM, Deployment/Operations, and Leadership) as Phase 1, ahead of customer-facing conversational analytics. • Build a semantic and metrics layer so every metric, such as "scan accuracy by site," is defined once in code and stays identical across every dashboard and product, making self-service safe from metric drift. • Own the quality bar: 99%+ availability SLA with freshness guarantees, 100% traceability, zero cross-tenant leakage, 99.5%+ pipeline success, and no data loss. • Design tenant isolation, per-tenant cost attribution, and schema and row-level RBAC to scale toward hundreds of tenants (300+ target), not today's fleet size. • Own data-ingestion correctness at the boundary with the integration/backend team, covering data contracts, schema validation, and pipeline quality, so WMS data lands in the right place, shape, and time across WMS versions. • Stand up a data catalog and lineage layer (Purview as the Azure-native fit, DataHub as the open-source alternative) so every consumer can find data, see ownership, and trace lineage when a metric looks wrong. • Prove the foundation end to end on Gather's drone product, then generalize it so each new product extends the model instead of rebuilding it • Act as the connective tissue between product and ML (3DCC, damage detection). Link structured records to unstructured drone imagery and video with full traceability, and stand up the data-infra readiness for feature stores and annotation pipelines on one trusted foundation.
Job Requirements
- 10+ years in data engineering, with 3+ years architecting data platforms for data products, analytics, or AI-driven products.
- Proven experience building a greenfield data warehouse and leading an OLTP to OLAP transition, not just maintaining an existing one.
- Deep expertise designing multi-layer transformation architectures and reusable frameworks that scale across multiple product areas.
- Expert SQL and dbt, hands-on ELT and orchestration, and large-scale or streaming data experience.
- Production experience on a major cloud (Azure preferred, AWS or GCP acceptable), plus infrastructure as code and CI/CD.
- Track record with data quality, security, governance, and multi-tenancy in production environments.
- Data transformation and modeling that turns raw multi-source data into refined, serving-ready datasets (raw to refined to serving).
- Pipeline orchestration and workflow automation for scheduling, dependency management, and reliable execution across data flows.
- Large-scale and distributed processing of high-volume batch data.
- Real-time and streaming ingestion that captures and processes event data as it arrives.
- Semantic and metrics-layer design that defines business metrics once and serves them consistently to every consumer.
- Serving-layer optimization for fast, low-latency consumption through wide and flattened tables and pre-computed metrics.
- Cloud data engineering and infrastructure automation that provisions, deploys, and operates the platform reproducibly (cloud-native, infrastructure as code, CI/CD).
- Data quality, observability, and lineage that ensure trust, freshness, and end-to-end traceability.
- Security, governance, and multi-tenancy including tenant isolation, access control, and resiliency.
- Multimodal data integration that links structured records to unstructured image and video (drone captures) with traceability.
Related Guides
Related Categories
Related Job Pages
More Data Engineer Jobs
• Define company data assets (data model) • Define and design data integrations, data quality frameworks and design and evaluate open source/vendor tools for data lineage • Work closely with Dropbox business units and engineering teams to develop strategy for long term Data Platform architecture to be efficient, reliable and scalable • Conceptualize and own the data architecture for multiple large-scale projects, while evaluating design and operational cost-benefit tradeoffs within systems • Collaborate with engineers, product managers, and data scientists to understand data needs, representing key data insights in a meaningful way • Design, build, and launch collections of sophisticated data models and visualizations that support multiple use cases across different products or domains • Optimize pipelines, dashboards, frameworks, and systems to facilitate easier development of data artifacts
Senior Software Engineer, Data
NavaBuilding simple, effective government services. Want to contribute? We're hiring!
• Lead a team of 2-3 engineers • Working with fellow Nava engineers to design, review, and build well-crafted software • Working in an agile manner to efficiently ship new features that meet user needs • Design and develop scalable data ingestion and processing pipelines • Implement large-scale data ecosystems within cloud-based platforms that include data management and data governance of structured and unstructured data • Implement data validation and quality checks to ensure accuracy and consistency • Work with cross-functional project teams to gather business requirements and translate to detailed technical specifications • Work with Government partners to assist and develop data engineering applications and pipelines that will enable data services and processing capabilities • Participate in software design and code reviews • Maintain security and privacy standards in all aspects of the data pipeline • Taking part in hiring activities (e.g., submitting referrals, conducting interviews, and attending interview debriefs), as needed
Senior Data Engineer – Contractor
Cars & BidsCars & Bids has daily auctions of cool enthusiast vehicles from the 1980s to the 2020s.
• Support entire data stack from ingestion through warehousing and transformation • Build and maintain bespoke syncs for operational tools • Investigate data quality issues • Conduct team syncs for alignment
Role Description You will work day-to-day alongside our Senior Data Engineer, Full Stack Developer and AI Automation Developer in a small, practical and collaborative team. The role will involve: - Building and maintaining data pipelines. - Supporting integrations between our core systems. - Cleaning and transforming clinic data. - Improving reporting workflows. - Helping migrate legacy Microsoft 365 / Power Automate processes onto a more scalable modern data stack. This is not a narrow “reporting only” role. You will be close to the real operational problems of a growing veterinary group and will be expected to take ownership of useful, practical projects that make a visible difference across the business. What you’ll do: - Build and maintain data pipelines and integrations between our core systems. - Extract, clean and transform data from multiple sources, including veterinary practice management platforms. - Support reporting, analytics and operational dashboards for teams across the group. - Help maintain existing Microsoft 365 automations and support their migration onto a more robust pipeline stack. - Work with clinic, finance, operations and leadership data to make it more reliable, usable and connected. - Support AI-driven automation projects by preparing, structuring and validating the data that powers them. - Help document, simplify and de-risk our data infrastructure. - Troubleshoot data issues, integration failures and reporting inconsistencies. - Work closely with the Senior Data Engineer to improve how we build, monitor and maintain data processes. Qualifications - Strong SQL, including joins, aggregations, window functions and the ability to read and improve queries written by others. - Working Python, including pandas, REST APIs, CSV and Excel handling, and basic error handling. - A good understanding of data pipelines, including scheduling, task dependencies, incremental versus full loads, and how to safely re-run failed jobs. - Git fundamentals, including branching, commits and merge requests. - Confidence working with Microsoft 365 tools such as Excel, SharePoint and Outlook. - Ability to open, understand and troubleshoot existing Power Automate flows. - A practical, problem-solving mindset and the confidence to ask questions, document your work and improve existing processes. - Interest in AI-driven automation and how clean, structured data can safely power internal AI tools. Requirements - Experience with any of the following would be useful, but we do not expect you to know everything: - Google BigQuery - Apache Airflow - dbt - Python 3 - Firebase, including Cloud Functions, Firestore and hosting - GitLab and CI/CD - Docker - Google Cloud Platform - AWS - Power BI - Power Apps - Power Automate Benefits - Salary: up to 40k - Home-based or hybrid working, depending on location and business needs. - Real ownership and variety from day one. - The chance to work on practical data and automation projects that directly affect 50+ clinics. - A small, collaborative team where your work is visible and your ideas are heard. - The opportunity to help shape the future data and AI infrastructure of a growing veterinary group.



