Job Closed

This listing is no longer active.

Mindrift logo
Mindrift

Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid. Project time expectations: Tasks are estimated to require around 10–20 hours per week during active phases, based on project requirements; This is an estimate, not a guaranteed workload, and applies only while the project is active. Note: Rates vary based on expertise, skills assessment, location, project needs, and other factors. Higher rates may be offered to highly specialized experts. Lower rates may apply during onboarding or non-core project phases. Payment details are shared per project.

Senior Data Scraping Engineer

Location

United States

Posted

59 days ago

Salary

$37 / hour

Seniority

Senior

Job Description

Senior Data Scraping Engineer

Mindrift

Role Description Mindrift is looking for highly skilled Python Data Scraping Engineers to join the Tendem project and drive specialized data scraping workflows within our hybrid AI + human system. In this role, as an AI Pilot – that’s how we refer to this role at Mindrift – you’ll collaborate with Tendem Agents that handle repetitive tasks, while you provide critical thinking, domain expertise, and quality control to deliver accurate and actionable results. This part-time remote opportunity is ideal for technical professionals with hands-on experience in web scraping, data extraction and processing. Key Responsibilities - Own end-to-end data extraction workflows across complex websites, ensuring complete coverage, accuracy, and reliable delivery of structured datasets. - Leverage internal tools (Apify, OpenRouter) alongside custom workflows to accelerate data collection, validation, and task execution while meeting defined requirements. - Ensure reliable extraction from dynamic and interactive web sources, adapting approaches as needed to handle JavaScript-rendered content and changing site behavior. - Enforce data quality standards through validation checks, cross-source consistency controls, adherence to formatting specifications, and systematic verification prior to delivery. - Scale scraping operations for large datasets using efficient batching or parallelization, monitor failures, and maintain stability against minor site structure changes. Qualifications - At least 3 years of relevant experience in data engineering, web scraping, automation, or software development (required). - Bachelor's or Master’s Degree in Engineering, Applied Mathematics, Computer Science, or related technical fields is a plus. - Strong experience in Python web scraping (BeautifulSoup, Selenium or similar), including dynamic content (JS, AJAX, infinite scroll) and APIs via proxies. - Proven ability to extract data from complex structures (hierarchies, archived pages, inconsistent HTML). - Solid background in data cleaning, normalization, and validation, delivering structured datasets (CSV, JSON, Google Sheets). - Hands-on experience with LLMs and AI frameworks to enhance automation and problem-solving. - Strong attention to detail and commitment to data accuracy. - Self-directed work ethic with ability to troubleshoot independently. - A link to GitHub is a plus. - English proficiency: Upper-intermediate (B2) or above (required). Compensation On this project, contributors can earn up to $37 per hour equivalent, depending on their level and pace of contribution. Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements. Benefits - Work fully remote on your own schedule with just a laptop and stable internet connection. - Gain hands-on experience in a unique hybrid environment where human expertise and AI agents collaborate seamlessly — a distinctive skill set in a rapidly growing field. - Participate in performance-based bonus programs that reward high-quality work and consistent delivery.

Related Categories

Related Job Pages

More Data Engineer Jobs

Nasstar logo

Principal Data Engineer – Contract

Nasstar

From cloud optimisation and application modernisation to connectivity and collaboration, we are Nasstar.

Data Engineer59 days ago
ContractRemoteTeam 1,001-5,000H1B No Sponsor

• Act as a trusted technical advisor across client engagements. • Build strong relationships with senior stakeholders, delivery leads and engineering teams. • Translate complex technical concepts into clear business value and delivery outcomes. • Support clients through technical decision-making, platform strategy and roadmap planning. • Lead the design and implementation of enterprise-scale Databricks solutions and modern data platforms. • Provide architectural leadership across cloud-native data engineering initiatives. • Define engineering standards, best practices and scalable delivery approaches. • Drive platform optimisation, resilience and performance improvements. • Lead technical delivery across multiple client engagements within a consultancy environment. • Work closely with cross-functional teams including Data Engineers, Architects, Product Owners and client stakeholders. • Support project planning, technical governance and delivery execution. • Ensure high-quality, scalable and maintainable engineering solutions. • Mentor and guide engineering teams across modern data engineering practices. • Lead technical problem-solving and incident resolution where required. • Drive adoption of engineering best practices, automation and DevOps principles.

United Kingdom
InStride Health logo

Data Engineer II

InStride Health

Evidence-based anxiety and OCD treatment for kids, teens, and young adults (ages 7-22).

Data Engineer59 days ago
Full TimeRemoteTeam 51-200H1B No Sponsor

• Design, develop, and maintain robust, scalable ETL/ELT data pipelines using Python, SQL, and data processing frameworks including dbt, Matillion, and AWS services. • Implement data quality checks, monitoring, and alerting across all data pipelines to ensure data integrity and reliability. • Optimize existing pipelines for performance, cost-efficiency, and error handling. • Contribute to the design and maintenance of InStride’s data warehouse and data lake solutions, including schema design, data modeling, and indexing strategies in Amazon Redshift. • Ensure data security, HIPAA compliance, and proper handling of protected health information (PHI) within all data infrastructure. • Work closely with data analysts, data scientists, and business intelligence engineers to understand their data requirements and deliver reliable, high-quality data access. • Troubleshoot and resolve complex production data issues with urgency and root-cause rigor. • Develop and maintain clear documentation for data models, pipelines, and data sources to enable self-service analytics. • Participate in code reviews and contribute to technical discussions, bringing a constructive and detail-oriented perspective. • Stay current on emerging data technologies and tools, bringing relevant insights to the team.

United States
$110K - $125K / year
Talent Connect logo

Data Engineer Azure Data Factory

Talent Connect

100% remoto desde cualquier país de Latam. Trabajo con equipos internacionales (principalmente USA). Horario full-time.

Data Engineer59 days ago

Role Description Integrarte a un equipo de datos para diseñar, desarrollar y operar pipelines ETL/ELT en Azure, garantizando la ingesta, transformación y disponibilidad de datos para consumidores de negocio y modelos analíticos. El rol implica responsabilidad técnica en la implementación de soluciones con Azure Data Factory, Azure Databricks y ecosistemas de almacenamiento (Blob Storage, Data Lake) aplicando buenas prácticas de calidad, rendimiento y gobernanza de datos. Responsibilities - Diseñar y desarrollar pipelines de ingestión y transformación de datos utilizando Azure Data Factory y Azure Databricks. - Implementar procesos ETL/ELT con Apache Spark y PySpark para procesamiento por lotes y en tiempo casi real. - Modelar y optimizar consultas SQL y estructuras de datos para consumo analítico y reporting. - Gestionar y organizar datos en Azure Blob Storage y Azure Data Lake (Gen2), definiendo particionamiento, formatos y políticas de retención. - Construir pipelines reproducibles, parametrizables y observables; instrumentar logging, métricas y alertas. - Colaborar con equipos de ciencia de datos, BI y producto para definir contratos de datos, SLAs y requisitos de calidad. - Automatizar despliegues y pruebas de pipelines; aplicar control de versiones y prácticas CI/CD para artefactos de datos. - Garantizar cumplimiento de políticas de seguridad y privacidad en el manejo de datos sensibles. - Documentar arquitecturas, pipelines y runbooks; transferir conocimiento al equipo y participar en on-call rotativo cuando aplique. Qualifications - Experiencia comprobable como Data Engineer (mín. 3 años) trabajando con Azure Data Factory y Azure Databricks. - Dominio de Apache Spark y PySpark para procesamiento distribuido y optimización de jobs. - Sólidos conocimientos de SQL: modelado, tuning y optimización de consultas. - Experiencia en diseño e implementación de pipelines ETL/ELT y en integración de datos. - Conocimiento práctico de Azure Blob Storage y Azure Data Lake (Gen2), incluyendo formatos como Parquet/Delta Lake. - Familiaridad con herramientas de orquestación, monitorización y CI/CD aplicadas a datos. - Inglés técnico: nivel intermedio (lectura de documentación y comunicación con equipos técnicos). - Residir en Argentina. Bonus Points - Experiencia con Delta Lake, optimizaciones de storage y compaction. - Conocimientos en seguridad y gobernanza de datos: RBAC, Data Catalog, políticas de enmascaramiento. - Familiaridad con Python avanzado y librerías de ingeniería de datos (pandas, numpy, airflow o equivalentes). - Experiencia con performance tuning de Spark y troubleshooting de jobs en Databricks. - Certificaciones Azure relacionadas con datos (DP-203, Azure Data Engineer Associate) o Databricks. - Experiencia en entornos cloud a gran escala y en trabajo ágil con equipos multidisciplinarios. Selection Process - Recruiting screen - Entrevista técnica - Entrevista con el cliente - Oferta Benefits - Envía tu solicitud y únete a nosotros para impulsar soluciones de datos robustas y escalables. - Para postular: - Envía tu CV a través de nuestro portal. - Completa tu perfil en talentconnect.ai. - Participa en una entrevista inicial utilizando inteligencia artificial. - Al completar estos pasos estarás más cerca de sumarte a nuestro equipo.

Argentina
Job Closed
Blend360 logo

Data Engineer Analyst

Blend360

Optimizing business performance through people, data, tech & analytics

Data Engineer59 days ago
Full TimeRemoteTeam 501-1,000H1B Sponsor

Role Description We are looking for a Data Engineer Analyst to contribute to the design and implementation of scalable, domain-oriented data solutions for a high-impact enterprise initiative. You will work within a specific data domain — such as Ingestion, Customer Data, Enrichment, or Marketing Signals — ensuring high-quality, reliable data pipelines that power analytics and business-critical workflows. You may work 100% remotely from anywhere in Mexico! - Contribute to the design and development of scalable data pipelines on AWS (Glue, S3, Athena, Step Functions). - Build and optimize batch and streaming ingestion pipelines, including CDC-based architectures. - Support the design and implementation of data models aligned with business and analytics needs. - Ensure data quality, reliability, and performance across pipelines and datasets. - Collaborate with cross-functional teams to align technical solutions with business requirements. - Apply best practices in data engineering, including modular design, reusability, and governance. - Leverage AI-assisted development tools to improve productivity and code quality. - Proactively identify risks, bottlenecks, and optimization opportunities across data workflows. Qualifications - Bachelor's degree in Computer Science, Data Engineering, Software Engineering, or equivalent field. - Hands-on experience with the AWS ecosystem, including Glue, S3, Athena, and Step Functions (required). - Solid proficiency in Python and PySpark for data pipeline development (required). - Experience with Snowflake or similar cloud data warehouse platforms (plus). - Understanding of data pipeline architecture and design patterns. - Exposure to CDC (Change Data Capture) patterns and streaming ingestion. - Experience with data modeling for analytics and reporting use cases. - Familiarity with AI-assisted development tools (plus). - Strong problem-solving skills and ability to work effectively in collaborative, agile environments. - English: Advanced (required for effective communication with global teams). Requirements - 3+ years of experience in Data Engineering or related disciplines, with hands-on work building and maintaining data pipelines on cloud platforms such as AWS. Benefits - Learning Opportunities: - Certifications in AWS (we are AWS Partners), Databricks, and Snowflake. - Access to AI learning paths to stay up to date with the latest technologies. - Study plans, courses, and additional certifications tailored to your role. - Access to Udemy Business, offering thousands of courses to boost your technical and soft skills. - English lessons to support your professional communication. - Travel opportunities to attend industry conferences and meet clients. - Mentoring and Development: - Career development plans and mentorship programs to help shape your path. - Celebrations & Support: - Special day rewards to celebrate birthdays, work anniversaries, and other personal milestones. - Company-provided equipment. - Flexible working options to help you strike the right balance. - Other benefits may vary according to your location in LATAM. For detailed information regarding the benefits applicable to your specific location, please consult with one of our recruiters.

Mexico
Job Closed