Mindrift logo
Mindrift

Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid. Project time expectations: Tasks are estimated to require around 10–20 hours per week during active phases, based on project requirements; This is an estimate, not a guaranteed workload, and applies only while the project is active. Note: Rates vary based on expertise, skills assessment, location, project needs, and other factors. Higher rates may be offered to highly specialized experts. Lower rates may apply during onboarding or non-core project phases. Payment details are shared per project.

Senior Python Engineer - AI Coding Agent Evaluation

Location

Worldwide

Posted

4 days ago

Salary

$200 / hour

Seniority

Senior

No structured requirement data.

Job Description

Senior Python Engineer - AI Coding Agent Evaluation

Mindrift

Role Description Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. What this opportunity involves: - Building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. - Creating challenging tasks and evaluation criteria within realistic simulated environments: - Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history. - Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent. - Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient. - Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust. What this is NOT: - Not data labeling. - Not prompt engineering. - Not writing code from scratch - the agent writes most of the code; you guide and evaluate. Qualifications - 8+ years in software development. - Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis. - Experience writing tests (functional, integration). - English proficiency - B2+. Requirements - Deep understanding of where models fail and what scenarios reveal the difference between a good and a bad solution. - Ability to create tasks that genuinely challenge the best models. - Writing tests that accept all correct solutions and reject incorrect ones. Benefits - Compensation: Up to $200/hr equivalent, depending on level and pace. - Tasks are estimated at ~30 hours each; you set your own schedule. How it works - Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid. - Tasks for this project are estimated to take 30 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. - Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.

Related Job Pages

More AI Engineer Jobs

Celonis logo

Senior Software Engineer - AI and Task Mining

Celonis

Celonis GmbH, founded in 2011, provides AI-enhanced, enterprise-ready process-mining technology and automated process-discovery solutions to help companies visu

AI Engineer4 days ago

Title: Senior Software Engineer - AI & Task Mining Location: Munich, Germany Job Description: Celonis is the global leader in Process Intelligence and the pioneer of Process Mining technology. As one of the world’s fastest-growing enterprise SaaS companies, we are changemakers pushing the boundaries of what’s possible. We invest heavily in advanced AI capabilities—specifically our Process Intelligence Graph—to turn data insights into immediate business action. We believe there is a massive opportunity to unlock global productivity and sustainability by placing intelligence at the core of every business process. Join our mission to make processes work for people, companies, and the planet. The Team Our team builds Celonis' end-to-end Task Mining solution. Task Mining is the technology that lets businesses capture user interaction data (desktop) so they can analyze how teams actually get work done — and how they can do it even better. We own the full stack behind it: the desktop client, backend services, data-processing pipeline, and Studio frontend applications. The Role We're looking for a Senior Software Engineer to help shape the AI and algorithmic core of Task Mining. Our mission is broad and growing: turning raw signals about how people work into structured, actionable insight. You'll design the algorithms and AI techniques behind these capabilities and ship them as robust, production-grade services to our customers. This is a role for someone who is energized by AI and algorithms first, and equally comfortable building the scalable services and data pipelines that deliver those capabilities reliably at enterprise scale. You'll work primarily in Python, with SQL across our data pipeline, and touch Java on the wider backend where it matters. You'll set technical direction, raise the engineering bar, and mentor others along the way. What you'll do - Design and build AI- and algorithm-powered features from scratch — from research and prototyping through to production rollout — across a growing set of problems. - Advance our core insight engine: evolve techniques that combine LLMs, machine learning, and rule-based/algorithmic methods to extract structure from complex, noisy interaction data — keeping results accurate, explainable, and cost-efficient. - Integrate new and diverse data sources into the product, designing the abstractions that let us plug them in cleanly. - Develop and deploy production-ready Python services — clean async APIs, containerized and integrated into a multi-tenant SaaS platform — to serve our AI and data pipelines. - Build and extend robust, scalable data pipelines for ingesting, processing, and transforming large datasets, and design the data models behind them. - Build strong evaluation into everything: guardrails, hallucination and schema-adherence checks, and latency/cost/quality metrics, so we can trust and continuously improve model and algorithm outputs. - Own end-to-end delivery: lead design, implementation, build, and shipping to customers. - Provide technical leadership and mentorship — drive design discussions, code reviews, and technical planning to keep standards high and knowledge shared. The qualifications you need - 5+ years of practical experience in a Computer Science / Data Science related field, or a PhD in Data Science/AI/ML with 2+ years of practical experience. - Hands-on experience applying AI/ML to real problems — including LLMs and Retrieval-Augmented Generation (RAG): writing effective prompts, using vector search, and evaluating outputs for hallucination, schema adherence, and performance (latency/cost). - A strong algorithmic foundation and the judgment to choose between AI-based and classical approaches for a given problem. - Experience designing, building, and deploying robust, scalable, production-ready Python services and APIs. - Strong understanding of Python's async I/O model and why it matters for high-performance, I/O-bound applications. - Solid grasp of ETL, data warehouses/lakes, data modeling, and schema design. - Experience with containerization (Docker) and CI/CD (e.g. GitHub Actions). - Strong communication and collaboration skills (English is a must). What Celonis can offer you: - Pioneer Innovation: Work with the global leader in Process Mining and the Process Intelligence Graph to shape the future of AI-driven business operations. - Ownership from Day 1: Every full-time "Celonaut" is an owner, receiving Restricted Stock Units (RSUs) and merit-based refresh grants. - Unrivaled Family Support: Benefit from our inclusive parental leave policy—24 weeks of fully paid leave for primary carers and 12 weeks for supporting carers, available from your first day of employment. - Work-Life Integration: Enjoy Unlimited PTO (in applicable regions) and generous PTO globally, as well as a flexible hybrid work model that balances remote focus with vibrant office collaboration. - Continuous Growth: Elevate your skills through our 70-20-10 learning framework, mentorship programs, and access to a dedicated learning platform. - Holistic Well-being: Prioritize your health with subsidized Wellhub memberships, mental health counseling, and dedicated "Wellness Weeks" that prioritize work/life balance. - Drive Sustainability: Participate in annual Impact Days, where you receive paid time off to volunteer for community and environmental causes with your local office, or virtually. - Global Inclusion & Belonging: Find community through our Inclusion Think Tank and participate in our annual Inclusion Days, ensuring every voice is heard and valued. - Value-Driven Impact: Join a mission-led organization where our core values—Live for Customer Value, The Best Team Wins, We Own It, and Earth Is Our Future—drive every decision. About Us: Celonis makes processes work — for people, companies, and the planet. Powered by process mining and AI, the Celonis Process Intelligence Platform integrates process data and business context to create a living digital twin of business operations. We enable thousands of companies worldwide to understand how their business actually runs and, together with their partners, build intelligent solutions that transform and continuously improve the way they operate — unlocking billions in value. Celonis is headquartered in Munich, Germany, and New York City, USA, with more than 20 offices worldwide.

Germany
Travoom logo

Principal AI Search, Conversation Architect

Travoom

Travoom is the marketplace for bucket list travel experiences.

AI Engineer4 days ago
Full TimeRemoteTeam 11-50H1B No Sponsor

• The Principal AI Search & Conversation Architect will be responsible for designing and leading the architecture of a next-generation conversational AI platform that replaces traditional search with intelligent dialogue. • This role owns the systems that understand user intent, determine what information is needed to satisfy a request, retrieve data from multiple real-time providers, reason across competing options, and generate conversational responses that lead users from discovery to transaction. • The successful candidate will architect the AI orchestration layer that coordinates multiple large language models, specialized agents, retrieval systems, recommendation engines, and external APIs into a single seamless conversational experience. • Rather than returning lists of search results, the platform will assemble and rank information from ticketing providers, hotels, airlines, restaurants, transportation services, mapping platforms, weather services, and other third-party data sources, presenting recommendations through natural conversation. • This role requires deep experience designing AI-native systems capable of planning multi-step tasks, maintaining conversational context, coordinating tool execution, ranking competing results, and continuously adapting recommendations based on user feedback throughout a conversation. • The Principal Architect will define the technical vision, establish architectural standards, evaluate emerging AI technologies, mentor senior engineers, and make long-term decisions regarding LLM strategy, retrieval architecture, agent design, memory management, semantic search, vector databases, knowledge graphs, and AI evaluation frameworks. • The position requires extensive experience building large-scale distributed systems, production AI platforms, retrieval-augmented generation (RAG), multi-agent orchestration, function calling, Model Context Protocol (MCP), recommendation systems, API orchestration, prompt architecture, conversation state management, personalization, and low-latency cloud infrastructure. • The Principal AI Search & Conversation Architect will work closely with the company’s other Principal Architects responsible for backend infrastructure, payments, blockchain, messaging, mini-programs, and platform services to ensure the conversational platform integrates seamlessly across the entire ecosystem while maintaining exceptional scalability, reliability, performance, and user experience.

Texas
IFS logo

AI Adoption Lead

IFS

Be your best when it really matters. At the #MomentOfService

AI Engineer4 days ago
Full TimeRemoteTeam 5,001-10,000Since 1983H1B Sponsor

• Lead the overall post-sale value realization strategy for Nexus Black. • Build, scale, and manage the AI Adoption Engineering organization. • Set the vision, define operating models, and influence GTM and product strategy. • Own portfolio-wide KPIs around adoption, time-to-first-value, retention, and expansion. • Drive cross-team coordination between Product, Engineering, and Support.

Illinois
$150K - $175K / year
SureSwift Capital logo

Senior AI-Augmented Full Stack Developer

SureSwift Capital

We are a remote-first, people-first company creating dream exits for bootstrapped founders.

AI Engineer4 days ago
Full TimeRemoteTeam 51-200Since 2015H1B No Sponsor

• Own feature development end-to-end, from UI design through backend implementation, QA, and deployment. • Partner with the Portfolio Leader and business teams to prioritize and deliver features for existing applications. • Build new features, fix bugs, and optimize performance across diverse tech stacks. • Set and uphold a high bar for how things look and feel to the user. • Develop repeatable approaches and best practices for AI-assisted development. • Own AI as a core part of the development workflow to accelerate delivery and reduce friction. • Apply and champion AI-powered code generation and agentic tools (e.g., Claude Code, Cursor, Copilot) throughout the build process. • Take ownership of understanding and extending legacy production systems while maintaining stability and performance. • Drive modernization of existing applications without disrupting customers. • Establish robust test coverage, regression checks, and safe deployment workflows.

Canada
$140K - $190K / year