Job Closed
This listing is no longer active.
Effortless Payments, Endless Possibilities
Data Engineer – Mid-level
Location
Brazil
Posted
58 days ago
Salary
0
Seniority
Senior
Job Description
Data Engineer – Mid-level
Bemobi
• Operate and evolve the Data Lake: work with data in its Raw, Processed and Refined zones, including deduplication processes, cataloging and storage optimization (Parquet, Iceberg). • Operate and monitor streaming pipelines via Kafka: create topics, connectors, ACLs, credentials and monitor consumer lag for real-time streaming pipelines. • Contribute to the evolution of the Data Platform API: create and maintain modules of our platform that abstract ingestion processes for users and developers (which may be via files, CDC or data streaming), data processing on Big Data platforms (Spark / Redshift) and pipeline orchestration (Airflow). • Investigate and resolve incidents: interpret failures in pipelines, dataset loads, data duplication and cluster errors, and assist in incident resolution. • Use and contribute to Infrastructure as Code (IaC): use Terraform and operate AWS resources under team guidance. • Participate in modernization and integration initiatives with AI tools: use LLMs to build agents and create tools that integrate AI into our systems (MCP Servers). • Documentation and monitoring maintenance: contribute to technical documentation of our tools and architecture. Maintain our observability tools.
Job Requirements
- Solid experience (3+ years) in Data Engineering or related areas.
- Proficiency in Python for developing pipelines, automation scripts and integrations.
- Basic understanding of software architecture.
- Hands-on experience with advanced SQL.
- Experience with Apache Airflow.
- Experience with AWS services: S3, Redshift, EMR (Spark).
- Knowledge of Apache Kafka: topics, producers/consumers, connectors (Debezium, S3 Sink).
- Familiarity with Git and CI/CD workflows (Bitbucket Pipelines or similar).
- Knowledge of Data Lake architectures (Lakehouse, Medallion Architecture).
- Good communication skills and ability to work autonomously within an agile team.
- Nice to have / Differentials:
- Experience with Apache Spark (PySpark, SparkSQL).
- Experience with Terraform or another Infrastructure as Code tool.
- Knowledge of C# / .NET.
- Familiarity with Debezium for Change Data Capture (CDC).
- Experience with modern table formats (Apache Iceberg, Hudi, Delta).
- Knowledge of Grafana for monitoring and operational dashboards.
- Experience with OpsGenie/JSM for incident and alert management.
- Familiarity with Redshift.
- Technical English for reading documentation and communicating with LATAM teams.
Benefits
- Bradesco National Network Health Plan - extended to dependents with no per-dependent discount;
- Optional Bradesco Dental Plan;
- Flexible meal/food allowance (VR/VA) - maintained during vacation;
- Profit-sharing (PLR);
- Wellhub;
- Birthday day off;
- Home office allowance;
- Commuter allowance (VT) as needed - legal deductions apply;
- Life insurance;
- Free access to all our products - AppsClub, Discount Club, TrueCaller, BTFit and Busuu;
- Access to internal training via digital platforms;
- Internal recognition program among employees - Bemobucks.
Related Guides
Related Categories
Related Job Pages
More Data Engineer Jobs
Senior Data Scraping Engineer
MindriftApply → Pass qualification(s) → Join a project → Complete tasks → Get paid. Project time expectations: Tasks are estimated to require around 10–20 hours per week during active phases, based on project requirements; This is an estimate, not a guaranteed workload, and applies only while the project is active. Note: Rates vary based on expertise, skills assessment, location, project needs, and other factors. Higher rates may be offered to highly specialized experts. Lower rates may apply during onboarding or non-core project phases. Payment details are shared per project.
Role Description Mindrift is looking for highly skilled Python Data Scraping Engineers to join the Tendem project and drive specialized data scraping workflows within our hybrid AI + human system. In this role, as an AI Pilot – that’s how we refer to this role at Mindrift – you’ll collaborate with Tendem Agents that handle repetitive tasks, while you provide critical thinking, domain expertise, and quality control to deliver accurate and actionable results. This part-time remote opportunity is ideal for technical professionals with hands-on experience in web scraping, data extraction and processing. Key Responsibilities - Own end-to-end data extraction workflows across complex websites, ensuring complete coverage, accuracy, and reliable delivery of structured datasets. - Leverage internal tools (Apify, OpenRouter) alongside custom workflows to accelerate data collection, validation, and task execution while meeting defined requirements. - Ensure reliable extraction from dynamic and interactive web sources, adapting approaches as needed to handle JavaScript-rendered content and changing site behavior. - Enforce data quality standards through validation checks, cross-source consistency controls, adherence to formatting specifications, and systematic verification prior to delivery. - Scale scraping operations for large datasets using efficient batching or parallelization, monitor failures, and maintain stability against minor site structure changes. Qualifications - At least 3 years of relevant experience in data engineering, web scraping, automation, or software development (required). - Bachelor's or Master’s Degree in Engineering, Applied Mathematics, Computer Science, or related technical fields is a plus. - Strong experience in Python web scraping (BeautifulSoup, Selenium or similar), including dynamic content (JS, AJAX, infinite scroll) and APIs via proxies. - Proven ability to extract data from complex structures (hierarchies, archived pages, inconsistent HTML). - Solid background in data cleaning, normalization, and validation, delivering structured datasets (CSV, JSON, Google Sheets). - Hands-on experience with LLMs and AI frameworks to enhance automation and problem-solving. - Strong attention to detail and commitment to data accuracy. - Self-directed work ethic with ability to troubleshoot independently. - A link to GitHub is a plus. - English proficiency: Upper-intermediate (B2) or above (required). Compensation On this project, contributors can earn up to $37 per hour equivalent, depending on their level and pace of contribution. Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements. Benefits - Work fully remote on your own schedule with just a laptop and stable internet connection. - Gain hands-on experience in a unique hybrid environment where human expertise and AI agents collaborate seamlessly — a distinctive skill set in a rapidly growing field. - Participate in performance-based bonus programs that reward high-quality work and consistent delivery.
Principal Data Engineer – Contract
NasstarFrom cloud optimisation and application modernisation to connectivity and collaboration, we are Nasstar.
• Act as a trusted technical advisor across client engagements. • Build strong relationships with senior stakeholders, delivery leads and engineering teams. • Translate complex technical concepts into clear business value and delivery outcomes. • Support clients through technical decision-making, platform strategy and roadmap planning. • Lead the design and implementation of enterprise-scale Databricks solutions and modern data platforms. • Provide architectural leadership across cloud-native data engineering initiatives. • Define engineering standards, best practices and scalable delivery approaches. • Drive platform optimisation, resilience and performance improvements. • Lead technical delivery across multiple client engagements within a consultancy environment. • Work closely with cross-functional teams including Data Engineers, Architects, Product Owners and client stakeholders. • Support project planning, technical governance and delivery execution. • Ensure high-quality, scalable and maintainable engineering solutions. • Mentor and guide engineering teams across modern data engineering practices. • Lead technical problem-solving and incident resolution where required. • Drive adoption of engineering best practices, automation and DevOps principles.
Data Engineer II
InStride HealthEvidence-based anxiety and OCD treatment for kids, teens, and young adults (ages 7-22).
• Design, develop, and maintain robust, scalable ETL/ELT data pipelines using Python, SQL, and data processing frameworks including dbt, Matillion, and AWS services. • Implement data quality checks, monitoring, and alerting across all data pipelines to ensure data integrity and reliability. • Optimize existing pipelines for performance, cost-efficiency, and error handling. • Contribute to the design and maintenance of InStride’s data warehouse and data lake solutions, including schema design, data modeling, and indexing strategies in Amazon Redshift. • Ensure data security, HIPAA compliance, and proper handling of protected health information (PHI) within all data infrastructure. • Work closely with data analysts, data scientists, and business intelligence engineers to understand their data requirements and deliver reliable, high-quality data access. • Troubleshoot and resolve complex production data issues with urgency and root-cause rigor. • Develop and maintain clear documentation for data models, pipelines, and data sources to enable self-service analytics. • Participate in code reviews and contribute to technical discussions, bringing a constructive and detail-oriented perspective. • Stay current on emerging data technologies and tools, bringing relevant insights to the team.
Data Engineer Azure Data Factory
Talent Connect100% remoto desde cualquier país de Latam. Trabajo con equipos internacionales (principalmente USA). Horario full-time.
Role Description Integrarte a un equipo de datos para diseñar, desarrollar y operar pipelines ETL/ELT en Azure, garantizando la ingesta, transformación y disponibilidad de datos para consumidores de negocio y modelos analíticos. El rol implica responsabilidad técnica en la implementación de soluciones con Azure Data Factory, Azure Databricks y ecosistemas de almacenamiento (Blob Storage, Data Lake) aplicando buenas prácticas de calidad, rendimiento y gobernanza de datos. Responsibilities - Diseñar y desarrollar pipelines de ingestión y transformación de datos utilizando Azure Data Factory y Azure Databricks. - Implementar procesos ETL/ELT con Apache Spark y PySpark para procesamiento por lotes y en tiempo casi real. - Modelar y optimizar consultas SQL y estructuras de datos para consumo analítico y reporting. - Gestionar y organizar datos en Azure Blob Storage y Azure Data Lake (Gen2), definiendo particionamiento, formatos y políticas de retención. - Construir pipelines reproducibles, parametrizables y observables; instrumentar logging, métricas y alertas. - Colaborar con equipos de ciencia de datos, BI y producto para definir contratos de datos, SLAs y requisitos de calidad. - Automatizar despliegues y pruebas de pipelines; aplicar control de versiones y prácticas CI/CD para artefactos de datos. - Garantizar cumplimiento de políticas de seguridad y privacidad en el manejo de datos sensibles. - Documentar arquitecturas, pipelines y runbooks; transferir conocimiento al equipo y participar en on-call rotativo cuando aplique. Qualifications - Experiencia comprobable como Data Engineer (mín. 3 años) trabajando con Azure Data Factory y Azure Databricks. - Dominio de Apache Spark y PySpark para procesamiento distribuido y optimización de jobs. - Sólidos conocimientos de SQL: modelado, tuning y optimización de consultas. - Experiencia en diseño e implementación de pipelines ETL/ELT y en integración de datos. - Conocimiento práctico de Azure Blob Storage y Azure Data Lake (Gen2), incluyendo formatos como Parquet/Delta Lake. - Familiaridad con herramientas de orquestación, monitorización y CI/CD aplicadas a datos. - Inglés técnico: nivel intermedio (lectura de documentación y comunicación con equipos técnicos). - Residir en Argentina. Bonus Points - Experiencia con Delta Lake, optimizaciones de storage y compaction. - Conocimientos en seguridad y gobernanza de datos: RBAC, Data Catalog, políticas de enmascaramiento. - Familiaridad con Python avanzado y librerías de ingeniería de datos (pandas, numpy, airflow o equivalentes). - Experiencia con performance tuning de Spark y troubleshooting de jobs en Databricks. - Certificaciones Azure relacionadas con datos (DP-203, Azure Data Engineer Associate) o Databricks. - Experiencia en entornos cloud a gran escala y en trabajo ágil con equipos multidisciplinarios. Selection Process - Recruiting screen - Entrevista técnica - Entrevista con el cliente - Oferta Benefits - Envía tu solicitud y únete a nosotros para impulsar soluciones de datos robustas y escalables. - Para postular: - Envía tu CV a través de nuestro portal. - Completa tu perfil en talentconnect.ai. - Participa en una entrevista inicial utilizando inteligencia artificial. - Al completar estos pasos estarás más cerca de sumarte a nuestro equipo.


