Job Closed
This listing is no longer active.
Senior Data Engineer
Location
United States
Posted
88 days ago
Salary
$143.1K - $174.6K / year
Seniority
Senior
Job Description
Senior Data Engineer
Spokeo
• Develop, optimize, and improve our data systems, including ETL pipelines, storage, and entity resolution • Build infrastructure and data automation pipelines to ingest, process, and load data from various sources • Automate and integrate new components into the data pipeline • Collaborate with stakeholders and data science teams to develop data products, including entity resolution and best selection, to efficiently execute product vision and strategy in alignment with organizational goals and priorities • Create unit and stress-test components to monitor technical performance and ensure that identified issues are resolved • Develop data analysis tools to provide data insights and capture key metrics • Research solutions and maintain technical documentation • Follow best practices for data governance, quality, cleansing, and other ETL-related activities
Job Requirements
- 7+ years of development experience in data engineering within a production environment (internships and academic settings excluded)
- Proven experience working with large datasets exceeding 100M+ records or multiple terabytes
- 2+ years of development experience in highly scalable, distributed systems and cluster architectures using AWS and utilizing EMR
- 5+ years of hands-on programming experience with Python
- 5+ years of professional experience working in big data ecosystems; Spark is required; PySpark is preferable
- 3+ years of experience with SQL, schema design, and dimensional data modeling
- 2+ years of professional experience working with dataflow orchestration tools, such as Airflow
- 2+ years of experience with non-relational databases (e.g., DynamoDB, Elasticsearch, etc.)
- A bachelor’s degree in Computer Science, Information Systems, Mathematics, or a related field is required
Benefits
- 100% medical/dental/vision coverage
- Unlimited employee PTO
- Bonus program
- Equity plans
- 401(k)
Related Guides
Related Categories
Related Job Pages
More Data Engineer Jobs
Role Description The Corporate Workspace Technology area is currently hiring for the role of Senior Data Engineer for the modernization of tooling enablement function. As a Senior Data Engineer, you will be responsible for design, modeling, development and management of data warehouse objects in Snowflake data store utilizing effective data pipelines such as Talend, Informatica and APIs. - Designing, implementing, managing scalable data solutions using Snowflake environment for optimized data storage and processing. - Migrate existing data domains/flows from relational data store to cloud data store (Snowflake). - Identify and optimize new/existing data workflows. - Identify and implement data integrity practices. - Integrate data governance and data science tools with Snowflake ecosystem as per practice. - Support the development of data models and ETL processes to ensure high quality data ingestion in cloud data store. - Collaborate with team members to design and implement effective data workflows and transformations. - Assist in the maintenance and optimization of Snowflake environments to improve performance and reduce costs. - Contribute to proof of concept, documentation and best practices for data management and governance within the Snowflake ecosystem. - Participate in code reviews and provide constructive feedback to improve team deliverables quality. - Design and develop data ingestion pipeline using Talend/Informatica using industry best practices. - Writing efficient SQL and Python scripts for large dataset analysis and building end to end automation process on a set schedule. - Design, implement data distribution layer using Snowflake REST API. Qualifications - Self-driven, dedicated individual who works well in a team and thinks and acts strategically. Requirements - Bachelor’s Degree, in lieu of a degree, demonstrating in addition to the minimum years of experience required for the role, three years of specialized training and/or progressively responsible work experience in technology for each missing year of college is required. Benefits - Competitive compensation. - Professional development opportunities. - Flexible work arrangements.
• Konzeption und Umsetzung moderner Datenpipelines: Du entwickelst und optimierst Batch- und Streaming-Datenpipelines und sorgst dafür, dass Daten zuverlässig, performant und skalierbar verarbeitet werden. • Aufbau leistungsfähiger ETL-/ELT-Prozesse: Mit Python und Databricks konzipierst, implementierst und betreibst Du robuste Datenintegrationsprozesse – von der Rohdatenaufnahme bis zur Bereitstellung für Analytics. • Entwicklung zukunftssicherer Datenmodelle: Du entwirfst, verwaltest und optimierst Datenmodelle für analytische Anwendungen und nachgelagerte Systeme und stellst deren Konsistenz und Wartbarkeit sicher. • Qualität und Stabilität der Datenplattform: Du überwachst und verbesserst kontinuierlich Datenqualität, Performance und Stabilität über die gesamte Plattform hinweg und etablierst geeignete Monitoring- und Testing-Strategien. • Enge Zusammenarbeit mit Team und Fachbereichen: Du arbeitest eng mit dem Entwicklungsteam zusammen und stimmst Dich mit den Fachbereichen zu datenbezogenen Anforderungen ab, um fachlich und technisch optimale Lösungen zu schaffen.
Role Description As a Senior Engineer on the Data Platform team, you’ll help shape the backbone of the client's data ecosystem. The reliability and scalability of the systems you build will directly influence how quickly and confidently teams across the company can innovate using data. - Design and build scalable data pipelines, batch jobs, and streaming processes that power analytics and machine learning. - Implement and optimize data storage systems (warehouses, lakes, analytical databases) with a focus on reliability, lineage, and governance. - Develop tools and frameworks to support data orchestration, ETL workflows, and big data processing. - Collaborate with engineers, data scientists, and product managers to align platform capabilities with business needs. - Ensure data systems meet requirements for quality, scalability, and cost efficiency in cloud environments. - Contribute to technical discussions, code reviews, and system design sessions. - Take ownership of features and projects, driving them from concept to production. Qualifications - Strong background in software engineering, data engineering and distributed systems, ideally at scale (terabytes+). - Hands-on experience with modern data warehouses, lakes, and compute/storage technologies (e.g., Snowflake, Pinot, TiDB, Spark, Airflow, Kafka, Trino/Presto). - Experience designing and building robust ETL pipelines and real-time streaming processes. - Solid understanding of data governance, lineage, and cost optimization in cloud environments. - Ability to write clean, maintainable, and performant code in production systems. - Strong problem-solving skills and ability to collaborate across teams. Benefits - Competitive salary + equity - RRSP matching - 3 weeks vacation + 3 personal care days + 1 Culture & Belief day + birthdays off - Access to a comprehensive mental health care platform - Full benefits from day one of employment - Work from home reimbursements - Optional global WeWork membership for those who want a change from their home office - Robust training and onboarding program - Coverage and support of personal development initiatives (conferences, courses, etc) - Access to industry-related programmatic courses and certifications to support continuous learning - Mentorship opportunities with industry leaders - An awesome parental leave policy - A friendly, welcoming, and supportive culture - Our social and team events!
• Consegue propor o melhor tipo de arquitetura para criação de pipeline de dados com base em um problema pré definido (ex: arquitetura lambda, arquitetura kappa, ETL, Scrapping, etc); • Tem conhecimento avançado de cloud computing e consegue defender as melhores peças de infraestrutura para um problema proposto (ex: Kubernetes, VMs, ECS, etc); • Consegue identificar e propor otimizações para problemas de performance em alguma etapa da arquitetura de dados (ex.: clusters de Spark, clusters de Kafka, bancos de dados, etc); • Capacidade de atuar em incidentes complexos, implementar monitoramentos e propor soluções que resolvam o problema na causa raiz.



