Join our team at Welo Data and embark on a journey of growth and innovation.
Data Trainer
Location
Worldwide
Posted
3 days ago
Salary
$24 / hour
Seniority
Mid Level
Job Description
Data Trainer
Welo Data
Role Description We are looking for detail-oriented Trainers to support an AI data annotation project. In this role, you will review pre-seeded questions paired with images and provide accurate "golden" answers based on what you observe in the image. - Review a pre-seeded question along with an accompanying image (e.g., "What is the title of the Excel file based on what you see in the image?") - Carefully examine image content to identify the relevant details needed to answer the question - Provide a clear, accurate "golden” answer based solely on the visual information provided - Flag any images that are unclear, corrupted, or insufficient to answer the question - Maintain consistency and quality across a batch of tasks - Follow project-specific guidelines and rubrics as provided by the project team Qualifications - Native/Strong proficiency at Somali - Fluent English proficiency (reading/writing) - Bachelor's degree required - Strong attention to detail and accuracy - Ability to follow written instructions precisely and consistently - Reliable access to a computer and internet connection - Prior experience with data annotation, data labeling, or quality review is a plus Requirements - Freelance, remote work type - 10 hours per week schedule - Compensation: $24/hr - Duration: Long-term
Related Guides
Related Categories
Related Job Pages
More Data Engineer Jobs
Data Scientist Architect
Navtech, Inc.NAVTECH INC 1600 Golf Road. Suite 1200, Rolling Meadows, IL 60008 Ph: (224) 348-1340 Email: alex@navtechusa.com Website: www.navtechusa.com E-Verified Company
Role Description Designs and develops scalable solutions using AI tools and machine-learning models. - Performs research and testing to develop machine learning algorithms and predictive models. - Utilizes big data computation and storage tools to create prototypes and datasets. - Conducts model training and evaluation. - Integrates, tests, tunes, and monitors solutions. - Proficient with multiple AI tools such as Python, Java, or R and machine learning frameworks like Spark, TensorFlow, or scikit-learn. Requires a master's degree in computer science, mathematics, engineering or equivalent. Typically reports to a manager or head of a unit/department. - P05-Expert: Works autonomously. Goals are generally communicated in "solution" or project goal terms. - May provide a leadership role for the work group through knowledge in the area of specialization. - Works on advanced, complex technical projects or business issues requiring state of the art technical or industry knowledge. - Typically requires 10+ years of related experience. Qualifications - Master's degree in computer science, mathematics, engineering or equivalent. - 10+ years of related experience. Requirements - Proficiency in AI tools such as Python, Java, or R. - Experience with machine learning frameworks like Spark, TensorFlow, or scikit-learn. - Ability to work autonomously and lead a work group. - Expertise in advanced technical projects or business issues.
Data Engineer
PairesPaires is where founders come to raise capital. We pair them with the right investors from a large, engaged global investor network, then our agents run the warm outreach and manage the relationships that turn into meetings.
Role Description We are hiring our first Data Engineer to own the database our agents and outreach are built on. Paires is where founders come to raise capital. We pair them with the right investors from a large, engaged global investor network, then run the warm outreach that turns into meetings. It is a two-sided platform, live with paying clients, profitable and self-funded, built by a small, senior, flat team that ships fast. The role involves: - Managing a comprehensive database of companies, investors, funding rounds, and relevant news. - Designing, scaling, and maintaining the database to ensure it serves as the single source of truth. - Ensuring the database is not merely a reporting or analytics warehouse but a live product memory. What you will own - The database itself: Postgres and Supabase with hybrid search, schema design, modeling, scaling, and performance. - Data quality end to end: validation gates for vendor and third-party data, deduplication, entity resolution, provenance, monitoring. - The communications layer: raw emails and call transcripts stored, linked to the right people and companies, and searchable. - Ingestion and enrichment pipelines: funding rounds, market news, and contact and company research at scale. - The knowledge graph: companies, investors, funding rounds, and news as entities and relationships. - The unified data layer: one clean spine that every campaign, agent, and product feature reads from. Qualifications - Experience owning a database of companies, people, deals, or communications. - Strong skills in SQL and Python with real pipeline work experience. - Ability to catch bad data before it impacts the business. - Experience thinking in schemas and contracts, designing for future queries. - Experience modeling entities and relationships at scale. - Ability to move fast with AI tooling and own outcomes. Requirements - No specific title required; experience in RevOps or growth roles is acceptable. - Hands-on experience with AI tools is preferred. Benefits - Fully remote and async work environment. - Flexible hours with some overlap with US Eastern time. - Meetings batched on Mondays and Thursdays, allowing for deep work. - Access to the best AI tooling, including Claude Code, Cursor, and top models. - Collaboration with the GTM lead and founding engineers. How to apply Hit apply, which takes you to our short application form. We read every application.
Role Description Forge requires a Mid Data Engineer to support legacy-to-modern data transformation in a secure AWS environment for a DoW customer. The role will develop batch and event-driven pipelines, automate data quality and testing, integrate with application services, and provide observable, recoverable, high-quality data flows across mission and external interfaces. Key Responsibilities - Build secure Python and AWS ETL/ELT pipelines for ingestion, transformation, reconciliation, and delivery. - Develop and evolve relational data models, schemas, indexes, constraints, views, and access patterns for MariaDB, PostgreSQL, or comparable platforms. - Develop data workflows and interfaces using Python on AWS Lambda and PySpark for event-driven, batch, and distributed transformation workloads. - Create automated data-quality checks for accuracy, completeness, consistency, timeliness, uniqueness, and business-rule conformance. - Implement source-to-target mapping, lineage, auditability, restartability, exception handling, and controlled replay. - Develop parity tests that compare legacy and modern processing outcomes and document the disposition of intentional differences. - Tune SQL and pipeline performance for high-volume batch and near-real-time workloads while protecting transactional integrity. - Implement monitoring, logging, alerting, and operational dashboards for pipeline health, latency, failures, and data quality. - Automate CI/CD, version-controlled data changes, deployments, rollback, and operational recovery controls. - Collaborate with architects, mission SMEs, Appian developers, testers, security personnel, and interface partners. Qualifications - Ability to think strategically, act tactically, and demonstrate strong analytical and critical-thinking skills. - Build strong cross-group working relationships and demonstrate exceptional organizational skills and attention to detail. - Thrive and succeed in an entrepreneurial environment and not be hindered by ambiguity or competing priorities. - Self-managing candidates who enjoy working collaboratively in a fast-paced environment and with dynamic teams. Requirements - U.S. Citizen (Authorization to Work in the U.S. will not suffice); previous professional experience supporting the U.S. Federal Government, either as a federal employee or contractor, is required. - 4+ years of professional experience in data engineering, database development, or data-platform delivery. - Bachelor's degree in Computer Science, Information Systems, Data Engineering, or equivalent, OR 4 additional years of relevant professional experience in lieu of a degree. - Advanced Python software-engineering skills and experience building AWS Lambda functions and PySpark data-transformation pipelines. - Experience building, testing, and operating production ETL/ELT pipelines with automated data-quality controls. - Experience with data modeling, schema migration, source-to-target mapping, lineage, reconciliation, and data-quality automation. - Experience integrating data platforms with REST APIs, application services, file exchanges, and event-driven interfaces. - Experience with Git, CI/CD, automated testing, logging, monitoring, performance tuning, and production support. - Active CompTIA Security+ or equivalent DoW-approved baseline cybersecurity certification, or ability to obtain within the first 30 days of starting. - Active Tier 2 background investigation or higher, completed or favorably adjudicated within the previous 18 months. Highly Desired Qualifications - Experience using Palantir Foundry for data integration, transformation, lineage, governance, and operational workflows. - Active Secret security clearance preferred. - Previous professional experience supporting a DoW organization, mission, or customer is strongly preferred; experience in modernizing COBOL flat files or legacy relational data into a modern relational architecture. - Experience integrating with Appian, Python microservices, financial transactions, logistics workflows, or high-volume external interfaces. - Experience serving as a technical team lead, mentoring junior engineers, or assisting teammates across delivery tasks. - Experience with BI, analytics, archival, records-retention, or NARA-aligned data lifecycle requirements. Benefits - Complete Flextime - 401k With Employer Matching - Healthcare, Including Medical, Dental, and Vision - Health Savings Account (HSA) And Pre-Tax Premium Options - Supplementary healthcare and family support - Extended Short-Term Disability and Long-Term Disability - Healthcare Insurance Deductible Paydown - Health and Wellness Programs - Tuition Reimbursement, Student Loan Repayment, and Education & Training Stipends - Cell Phone / Internet Stipends - College Saving Plans with Employer Contributions - Alternative Work Locations and Tele-Commuting - Employee Referral Awards - Retention, Signing & Performance Bonuses - Commuter Benefits - Paid Sabbatical
Senior Data Engineer, Microsoft Fabric Engineer
Weekday (YC W21)We are a Y-Combinator-backed startup building your AI-powered Recruiter Agent
• Design, develop, and maintain scalable data pipelines using Microsoft Azure Fabric, Databricks, and Azure data services. • Build and optimize robust ETL/ELT processes for structured, semi-structured, and unstructured data across enterprise environments. • Collaborate with business stakeholders to understand data requirements and translate them into scalable technical solutions. • Develop and maintain reliable data integration workflows that ensure data accuracy, consistency, and availability. • Optimize data processing pipelines for performance, scalability, cost efficiency, and operational reliability. • Leverage AI-assisted development tools and modern engineering practices to accelerate solution delivery and improve code quality. • Monitor, troubleshoot, and resolve production data pipeline issues while proactively identifying opportunities for automation and optimization. • Work closely with architects, analysts, developers, and cross-functional teams throughout the project lifecycle to deliver high-quality data solutions. • Implement best practices for data engineering, governance, documentation, testing, and operational support. • Take ownership of project deliverables by ensuring quality, meeting timelines, communicating risks proactively, and continuously improving engineering processes.
