Job Closed
This listing is no longer active.
Veriff is an industry leader in online identity verification, helping businesses achieve greater levels of trust.
Senior Data Engineer
Location
Estonia
Posted
136 days ago
Salary
0
Seniority
Senior
Job Description
Senior Data Engineer
Veriff
• Owning and evolving our data lake and data warehouse infrastructure using technologies such as Spark, Apache Iceberg, S3, Trino/Athena, and Redshift. • Designing and maintaining platform-level data transformation pipelines in Python and SQL — focused on schema evolution, partitioning, compaction, and deduplication. • Implementing optimized storage formats (Parquet, Avro, ORC), partitioning strategies, and indexing to improve query performance and reduce platform costs. • Driving data governance initiatives — PII detection and classification, access control policies, data cataloging, lineage tracking, and data quality frameworks. • Ensuring the availability, reliability, and cost efficiency of the data platform, including observability, monitoring, and alerting for pipeline and query engine health. • Collaborating with ML, analytics, product, and engineering teams to define data contracts, maintain schema consistency, and provide clean, well-governed datasets. • Contributing to disaster recovery strategy and multi-region reliability of the data platform.
Job Requirements
- Strong experience with Python, SQL, and Apache Spark / PySpark for large-scale data processing.
- Deep knowledge of modern analytics platform architecture — object stores, columnar and row-based data formats (Parquet, Avro, ORC), orchestration tools, analytical query engines, schema registries, and data catalogs.
- Experience with data governance and data management at scale — PII handling, data cataloging, schema management, access control, and data quality frameworks.
- Experience designing and operating data lake and data warehouse infrastructure.
- Solid understanding of storage optimization — partitioning, compaction, and compression trade-offs.
- Experience building observability, monitoring, and alerting for data platforms.
- Strong problem-solving skills and comfort working with ambiguity — defining problems before solving them.
- A collaborative mindset — this role serves ML, analytics, and product teams as internal customers.
- Experience with Infrastructure as Code (IaC) and Terraform.
- Familiarity with containerization — Docker and Kubernetes.
- Experience with CI/CD pipelines for data platform deployments.
- Knowledge of data lake table formats beyond Iceberg (Delta Lake, Hudi).
- Familiarity with data catalog and metadata management tools (e.g., DataHub, Amundsen, AWS Glue Catalog).
- Understanding of data privacy regulations (GDPR) in the context of data engineering.
- Experience building streaming data pipelines.
- Experience with the AWS data stack.
Benefits
- Flexibility to work from home
- Stock options that ensure your share in our success
- Extra recharge days on top of your annual vacation
- Comprehensive relocation support to Estonia or Spain
- Extensive medical, dental, and vision insurance to ensure you’re feeling great physically and mentally
- Learning and Development & Health and Sports budget that you are free to tailor to your own needs
- Four weeks of fully paid sabbatical leave after reaching your 5th work anniversary
Related Guides
Related Categories
Related Job Pages
More Data Engineer Jobs
• Support the Data Solutions- Operations team in a variety of functions, such as: • Support development and testing of configuration-driven ETL jobs in AWS Glue. • Assist in standardizing job patterns, schema contracts, and reusable utilities. • Help refactor legacy or hard-coded logic into config-driven structures. • Document pipeline dependencies, data flows, and operational runbooks. • Contribute to data quality checks, validation rules, and monitoring improvements. • Assist with troubleshooting job failures and performance optimization efforts. • Support CI/CD, deployment validation, and environment configuration consistency. • Perform additional duties as assigned or requested.
• Support master data and analytics initiatives by assisting with taxonomy development, data standardization, and enterprise data modeling. • Conduct taxonomy and classification research to support master and reference data modeling. • Assist in documenting business definitions, hierarchies, and data relationships. • Support data modeling and schema updates within the Nexxus project. • Contribute to data quality analysis and remediation efforts. • Help prepare datasets and transformations supporting master data workflows. • Collaborate with Data Solutions and business stakeholders to validate taxonomy structures. • Perform additional duties as assigned or requested.
Cloud Data Architect
CitizantWorking with Federal agencies on IT and business transformation. Follow us for info on Agile/DevOps & Data Strategies.
• The Cloud Data Architect serves as the senior technical authority responsible for defining, governing, and evolving enterprise architecture to support mission and business objectives • Leads the development of architecture strategies, standards, and technical roadmaps • Ensures alignment with enterprise architecture frameworks and federal best practices • Provides architectural oversight across engineering, product development, infrastructure, data platforms, and digital services • Develops and maintains enterprise architecture framework, reference architectures, and technology standards • Establishes and leads architecture governance processes • Maintains enterprise architecture artifacts • Defines and maintains reference architectures for cloud environments, SaaS platforms, enterprise applications, and data services • Guides engineering and development teams in designing integrated, scalable, and resilient solutions • Develops and maintains multi-year technology roadmaps aligned with organizational goals
Data Engineering Lead
MediaRadar, Inc.Sales enablement platforms customized for media, and ad tech companies that help you close more deals.
This description is a summary of our understanding of the job description. Click on 'Apply' button to find out more. Role Description We are seeking a visionary and hands-on Data Engineering Lead to spearhead the design, development, and optimization of our next-generation data platform. As a Lead, you will balance technical excellence with people leadership, ensuring our data architecture is scalable, resilient, and perfectly aligned with our business goals while collaborating cross-functionally to support analytics, reporting, and operational data needs. The ideal candidate should be a PySpark expert who thrives in the Azure ecosystem and has a deep appreciation for clean, modular code and robust ETL patterns. This is an exciting opportunity to work along with a great team of data engineers, demanding technologies, and an engaging work environment to help shape our data engineering best practices. Key Responsibilities - Technical Leadership: Design and supervise the implementation of comprehensive data pipelines utilizing Azure Databricks and PySpark. - Team Mentorship: Direct a team of data engineers, performing code reviews, offering technical expertise, and cultivating a culture of ongoing learning. - Data Modeling & Optimization: Develop high-performance schemas in PostgreSQL and refine complex SQL queries for large datasets. - ETL Strategy: Establish and apply optimal practices for data ingestion, transformation, and storage (Delta Lake/Lakehouse patterns). - Strategic Collaboration: Collaborate closely with Data Analysts, Architects, and Product Managers to convert business requirements into technical specifications. - Process Improvement: Promote the implementation of CI/CD, unit testing, and automated monitoring to achieve 99.9% data reliability. - Ensure data quality, governance, and compliance through validation, documentation, and secure practices. - Continuously improve data systems for enhanced performance, reliability, and scalability. - Effectively engage within an agile, cross-functional project team. Mandatory Skillset - Azure Databricks: - Expert-level experience managing workspaces, clusters, and job scheduling. - Solid understanding of data lakehouse architectures and Delta Lake. - Proven experience in Performance Tuning, Spark Optimization and Cost Reduction. - PySpark: Advanced proficiency in Spark DataFrame APIs and Spark SQL for large-scale data processing involving various data formats. - SQL Mastery: Exceptional ability to write, tune, and troubleshoot complex queries. - PostgreSQL: Hands-on experience with relational database design, indexing, and performance optimization. - ETL/ELT Frameworks: Proven track record of building scalable data pipelines from scratch. Desired Skills - Workflow Orchestration: Experience with Apache Airflow for managing complex task dependencies. - Containerization: Familiarity with Azure Kubernetes Service (AKS) for deploying containerized data services. - Infrastructure as Code (IaC): Knowledge of Terraform or Bicep for managing Azure resources. Qualifications - 10+ years of experience in Data Engineering or Software Engineering. - 3+ years as a formal technical Lead managing an agile team and implementing E2E solutions. - Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field. - Strong communication skills with the ability to explain complex technical concepts to non-technical stakeholders. - Strong problem-solving skills and attention to detail. Benefits - You won't just be "maintaining" pipelines; you'll be the primary architect of a data ecosystem that powers real-time decision-making across the entire organization.



