Job Closed
This listing is no longer active.
Kayzen powers the world's best mobile marketing teams to take programmatic in-house.
Senior Data Engineer/Data Ops
Location
India
Posted
118 days ago
Salary
0
Seniority
Senior
Job Description
Senior Data Engineer/Data Ops
Kayzen
• Create innovative solutions for handling peta-bytes of data with billions of rows & joins • Program and maintain our data pipelines that fuels our on-premise/cloud data warehouse used to generate and serve our models • Maintain and improve our fleet of data servers (the software), making sure they are reliable and able to process our billions of logs and data points • Develop and productionize data pipelines for our ML models in both bare-metal and the cloud environment • Make suggestions and lead projects to improve our data processing capabilities • Contribute to the team enabling us to be always better
Job Requirements
- Minimum 5+ yrs of professional experience in creating and maintaining big data pipelines
- Identifying data related process improvements
- Maintaining Kubernetes, Hadoop and Spark infrastructure
- Bachelor's/Master’s degree in a quantitative field of Mathematics, Physics, Computer Science, Machine learning Engineering, Business Analytics, Information Management or related field
- Knowledge of relevant programming languages (Python, Java, etc)
- Expert in SQL & NoSQL and big data processing pipelines (we use Python, Spark, Airflow)
- Proven experience managing data infrastructure that can store and process Petabytes of data (we use Hadoop, Spark)
- Kubernetes wizard
- Proven affinity with data
- Strong analytical and problem-solving skills
- Ability to translate business requirements into data solutions
- Excellent stakeholder management skills
- Previous experience with Clickhouse is a plus
- Previous experience with ad-tech is a plus
- Experience with Real time big data processing is a plus
- General understanding of Machine Learning techniques (Neural Networks, Random Forest, etc.) and ML frameworks (Mlflow, PyTorch, Tensorflow, etc.) is a plus
Benefits
- Exceptional career growth and learning opportunity
- A unique opportunity to be part of an experienced team of industry experts and entrepreneurs who bring massive change to the Adtech market
- Direct, day-to-day work experience with the management
- A fun, driven, and multinational team located across Germany, India, Argentina, Ukraine, Turkey, the UK and soon more countries
- A flexible work-from-home arrangement
- A 500-dollar home-office setup budget
- A 1000-dollar annual learning and development budget
Related Guides
Related Categories
Related Job Pages
More Data Engineer Jobs
• Partner with the engineering team to lead the knowledge transfer of "Atlas" (built in Python), taking full architectural ownership and ensuring its continued health and progression. • Using Apache Airflow to manage and optimize the daily ingestion of analytics from multiple disparate sources (social media, podcast platforms, etc.) to ensure a clean, reliable, and "healthy" data stream. • Act as a "data detective" to identify what information we are missing and prioritize new data collection that aligns with our financial stability and growth levers. • Build and maintain intuitive dashboards (e.g., Power BI or custom builds) that allow non-technical peers to answer basic everyday questions without needing manual engineering support. • Transform raw data into structured, "customer-ready" packages to support our heavy push into data licensing. • Utilize your knowledge of the media analytics landscape to guide the organization on whether to build custom internal tools or leverage existing 3rd-party social media analytics solutions.
Principal Data Engineer
WaymarkThe breakthrough AI production platform that allows anyone to create compelling commercials and spec spots in minutes.
• Architect production-grade data pipelines that integrate clinical data through multiple channels—direct EHR connections (e.g., Epic, Cerner, Athenahealth), health information exchanges (HIEs), health alliance networks, and third-party integration vendors—via FHIR R4 APIs, HL7v2 feeds, CCDA documents, and bulk data exports, while enforcing healthcare standards and clinical terminologies (ICD-10, SNOMED CT, LOINC, RxNorm), targeting sub-hour latency from clinical event to actionable insight and ≥99.9% pipeline reliability. • Lead end-to-end integration efforts with health plan partners, HIEs, health alliance networks, provider organizations, and data integration vendors—owning partner technical evaluation, connectivity, data mapping, validation, and production rollout. • Build and optimize cloud-native data infrastructure (AWS) and ETL/ELT workflows using modern orchestration tools (e.g., Step Functions, dbt) with robust data quality monitoring and lineage tracking. • Develop reusable backend frameworks, libraries, and internal tooling that accelerate onboarding of new data sources—whether direct EHR feeds, HIE connections, or vendor integrations—with the goal of reducing new-partner integration time by 50% or more while improving developer productivity and reducing operational toil across engineering. • Partner with data science and data analytics teams to build and operationalize the data foundations for predictive risk scores, care gap identification, clinical alerting, and patient outreach prioritization. • Collaborate with product and clinical teams to translate care delivery requirements into scalable, production-grade technical solutions and influence. • Ensure all data systems comply with HIPAA, HITECH, and applicable privacy regulations, implementing access controls, audit logging, encryption, and de-identification processes for PHI. • Provide technical leadership and mentorship to engineers across Waymark, raising the bar on healthcare data best practices.
Senior AWS Redshift Data Engineer
Techmango Technology Services Private LimitedData Engineering | Generative AI | Microsoft Dynamics 365 | AI &ML | Application Modernization | Business Intelligence
• Design and maintain scalable, efficient, and well-partitioned schemas in MSSQL, Redshift, and Snowflake. • Architect and optimize complex queries, stored procedures, indexing strategies, and partitioning for large datasets. • Build, monitor, and maintain data pipelines that ensure timely and accurate delivery of data to internal and external consumers. • Own and enforce data refresh SLAs, ensuring availability, consistency, and reliability across production and reporting environments. • Collaborate with software engineers, analysts, and DevOps teams to ensure data models and queries align with product and reporting requirements. • Proactively identify and remediate performance bottlenecks, slow queries, and data inconsistencies. • Implement and manage database change workflows using schema migration/versioning tools. • Define and promote best practices for data access, security, compliance, and observability.
• Own the end-to-end migration workflow for new customers. • Design mapping strategies to transform competitor data into our unified schema. • Build and maintain migration tooling (scripts, internal apps, reusable templates). • Collaborate with Client Success and Product teams to resolve data discrepancies and edge cases. • Optimize and automate recurring migration tasks to drive scalability. • Diagnose and troubleshoot migration issues quickly and systematically. • Continuously refine best practices for schema evolution, error handling, and data validation.




