Senior AI Backend Engineer - Agent Evaluation & Quality
Location
Saudi Arabia
Posted
2 days ago
Salary
0
Seniority
Senior
No structured requirement data.
Job Description
Senior AI Backend Engineer - Agent Evaluation & Quality
Salla
Role Description We run production multi-agent systems that handle real work for a large base of users. As those systems grow, our biggest constraint is confidence: we need to know how well the agents perform, catch regressions before they ship, and keep quality steady as we release. This role owns that. You'll build the evaluation systems behind our agents - the judges, test harnesses, and simulators that tell us whether an agent is working and where it's failing. The goal is to let us ship agents faster because we can trust what the evaluation tells us. Evaluation is the focus, but it won't be the boundary. Because you'll understand the agents' failure modes better than anyone there will also be opportunities to contribute to agent development itself, building and improving the agents alongside the systems that evaluate them. Responsibilities - Own the evaluation stack. Design and build LLM-as-judge systems, calibrate them against human labels, and make agent quality measurable per-agent and per-failure-mode. - Make the release gate real. Build per-PR eval harnesses and regression detection wired into CI, so quality is enforced automatically, not by manual passes. - Build user simulators to generate test coverage and adversarial cases before real users hit them. - Turn production signal into improvement - pipe real failures back into evaluation sets so the system compounds over time. - Partner with product to turn "what good looks like" into concrete, measurable criteria. - Grow into agent development - contribute to building and hardening the agents themselves, starting with the components you know most deeply from evaluating them. Qualifications - Strong software engineering fundamentals. Production Python or Typescript (or similar), clean API and system design, testing, CI/CD. You write code others build on - evaluation infrastructure is real engineering. - Hands-on LLM/agent experience. You've built with LLMs - agents, RAG, tool/function calling, orchestration frameworks (LangGraph, LangChain, or equivalent) - and understand how they behave and break. - A measurement mindset. You reason about metrics, calibration, and experiments; you want to quantify whether something works, not just ship it. - Production experience. You've run LLM systems in production and dealt with reliability, latency, cost, and observability. - 5+ years software engineering, with recent hands-on LLM/agent work. Nice to have - Direct experience evaluating LLM/agent systems - offline/online eval, LLM-as-judge, systematic regression testing. - Observability tooling (Arize, LangSmith, or similar). - Arabic language / NLP experience. - E-commerce or merchant-facing product experience.
Related Guides
Related Job Pages
More AI Engineer Jobs
Title: AI Engineer (Chicago) Location: IL-Lincolnshire Job Description: Job Description We are seeking an AI Engineer to join our Technology & AI Organization. This role sits at the intersection of AI-assisted application development and full-stack engineering. You will take AI-generated front-end prototypes and transform them into production-grade, fully integrated applications-connecting user interfaces to backend services, APIs, databases, and enterprise systems. You will work across the full software delivery lifecycle-from vibe-coded prototypes built in tools like Lovable and Claude, through GitHub-managed source control, CI/CD via Harness, and deployment to Vercel, GCP, and Azure. This is a hands-on engineering role for someone who thrives at turning rapid AI-generated concepts into reliable, secure, scalable enterprise software. Key Responsibilities: Full-Stack Integration of AI-Generated Applications - Take front-end applications generated through AI coding tools (Lovable, Claude Code, Cursor) and extend them into full-stack solutions with backend logic, API integrations, and database connectivity. - Wire up front-end React/Next.js applications to backend services including Supabase (PostgreSQL), Snowflake, Salesforce APIs, and internal microservices. - Implement authentication and authorization flows using Microsoft Entra ID (Azure AD) SSO, OAuth 2.0, and RBAC patterns. - Build and maintain RESTful and GraphQL API layers to connect UI components with enterprise data sources. AI-Assisted Development Workflow - Operate within Camping World's "Vibe Code" development workflow: Lovable → GitHub → Claude Code → Wiring → Harness → Azure/Vercel. - Leverage AI coding assistants (Claude, GitHub Copilot, Gemini Code Assist) to accelerate development while applying engineering judgment to AI-generated output. - Review, refactor, and harden AI-generated code for production readiness-addressing security vulnerabilities, performance concerns, and maintainability. - Collaborate with Agent Developers and AI Solution Architects to integrate AI agent capabilities into applications. DevOps & Deployment - Manage CI/CD pipelines in Harness for automated build, test, and deployment workflows. - Deploy applications to Vercel (front-end/serverless), GCP (Kubernetes/Cloud Run), and Azure as required. - Implement observability with Dynatrace, including application performance monitoring, log analytics, and alerting. - Follow software supply chain security practices including container image scanning (Chainguard), secrets management (CyberArk), and dependency auditing. Database & Data Integration - Design and optimize PostgreSQL schemas (with pgvector for AI/embedding workloads) in Supabase. - Build data integration pipelines connecting Snowflake data warehouse to application layers. - Work with Snowflake MCP connectors to enable AI-driven data access within applications. - Ensure data integrity, query performance, and proper indexing across application databases. API Management & Security - Configure and manage APIs through Kong/Konnect API gateway including rate limiting, authentication, and traffic management. - Implement application security best practices including input validation, prompt injection prevention, and secure coding standards. - Participate in code reviews with a focus on security and quality of AI-generated code. Pay Rate: $70-75 We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/. Skills and Requirements - 4+ years of professional software development experience with at least 2 years in full-stack roles. - Strong proficiency in JavaScript/TypeScript, React, and Next.js for front-end development. - Experience with Node.js, Python, or Go for backend/API development. - Working knowledge of PostgreSQL and at least one cloud data platform (Snowflake, BigQuery, or Redshift). - Experience with CI/CD pipelines and modern DevOps practices (GitHub Actions, Harness, Jenkins, or similar). - Familiarity with cloud platforms: GCP, Azure, or AWS (multi-cloud experience preferred). - Experience implementing OAuth 2.0 / SSO authentication flows (Microsoft Entra ID / Azure AD preferred). - Demonstrated ability to read, review, and improve code generated by AI tools. - Strong understanding of RESTful API design, versioning, and documentation. - Comfortable working in Git-based collaborative workflows with branching strategies and pull request reviews. - Hybrid/remote work flexibility. 3 Days in Office is a requirement. - Experience with AI-assisted development tools such as Claude Code, Lovable, Cursor, or GitHub Copilot. - Familiarity with Supabase (PostgreSQL + Auth + Edge Functions) as a backend platform. - Experience with Vercel for front-end deployment and serverless functions. - Knowledge of vector databases and pgvector for AI/embedding workloads. - Experience with Kong or similar API gateway platforms. - Exposure to Kubernetes, Docker, and container orchestration on GCP or Azure. - Understanding of AI agent architectures and how applications interact with LLM-based services. - Experience with Dynatrace, Datadog, or similar APM/observability platforms. - Familiarity with software supply chain security tools and practices (SBOM, image signing, Chainguard). - Background in retail, automotive, or RV dealership technology systems is a plus.
Applied AI Engineer
PairesPaires is where founders come to raise capital. We pair them with the right investors from a large, engaged global investor network, then our agents run the warm outreach and manage the relationships that turn into meetings.
Role Description We are hiring an Applied AI Engineer to own the conversation layer of an AI-first fundraising platform: the agents that handle every reply, carry conversations through to booked meetings, and look after every relationship across our warm outreach, our investor network, and investor relations. Paires is where founders come to raise capital. We pair them with the right investors from a large, engaged global investor network, then our agents run the warm outreach and manage the relationships that turn into meetings. It is a two-sided platform, live with paying clients, profitable and self-funded, built by a small, senior, flat team that ships fast. The role involves: - Owning how Paires converses: every reply handled well and in time, booking, nurture, and the coaching experiences we are building on the same data. - Starting as our reply and messaging system and growing into a core part of the product. - Building broader applied-AI features across the platform. - Engaging in real conversations with real people outside the team, day after day. What you will own - Agents that read inbound messages and draft the right reply, fast, with a human in the loop. - The messaging layer across our warm outreach, our warm investor network, and investor relations, not just one inbox. - Conversational agents end to end: replies, booking, nurture, and support that reads human. - The eval spine that gates quality: golden sets, judges, and the guardrails that hold as the conversations grow. - Broader applied-AI features: classification, extraction, routing, summarization. Qualifications - Ship conversational LLM systems to production with real users: reply handling, support, booking, or nurture agents. Inbound work counts fully here. - Have built evals yourself: golden sets from scratch, judge criteria, ship or rollback decisions made on the numbers. - Have integrated LLMs with email, CRM, or messaging systems. - Think like an operator, and know the business goal behind the message. - Are an engineer first. We run roughly 80/20 engineering to research. - Move fast with AI tooling and own outcomes. Requirements - Our stack: Python, Supabase, Pydantic AI, Claude Agent SDK, AWS. If you have shipped on any of it, lead with that. - Bonus: email and CRM integrations, RevOps exposure, sales or IR experience. Benefits - Fully remote and async. Your day overlaps with US Eastern time for a few hours - not full US hours. - Meetings batch on Mondays and Thursdays, the rest is deep work. - The best AI tooling, paid (Claude Code, Cursor, top models). How to apply Hit apply, which takes you to our short application form. We read every application.
Founding AI Engineer
PairesPaires is where founders come to raise capital. We pair them with the right investors from a large, engaged global investor network, then our agents run the warm outreach and manage the relationships that turn into meetings.
Role Description We are hiring our Founding AI Engineer to own the agent layer of an AI-first fundraising platform: the conversational agents that run our warm outreach and investor relationships, and the matching engine behind them. Paires is where founders come to raise capital. We pair them with the right investors from a large, engaged global investor network, then our agents run the warm outreach and manage the relationships that turn into meetings. It is a two-sided platform, a product for founders and a living network on the investor side, not one-way matching. We are live with paying clients, profitable and self-funded, and we run as a small, senior, flat team that ships fast. The role involves: - Owning our core: the agent system that runs the product. - Building, scaling, and running the agent layer and matching engine. - Shipping to production end to end, with the judgment to know what to build next as we scale. What you will own: - The agent layer: orchestrating work to specialist agents that research investors, draft outreach, and carry investor conversations end to end. - The matching engine: embeddings, ranking, scoring, and the feedback loop that sharpens matches over time. - The backend and data: Python, Postgres, Supabase, AWS, clean pipelines. - The agent layer built on the Claude Agent SDK and Pydantic AI. Qualifications - Have been one of the first engineers at a quick-growing company, running the show. - At least 3 years of shipping production software, taking a full app from zero to production and keeping it running end to end. - Experience shipping LLM systems to production with real examples. - Built retrieval, ranking, or matching with embeddings in production. - Strong Python skills and clear reasoning about data and systems. - Designed evals, caring about the quality of output. - Owned outcomes end to end and moved fast with AI tooling. Requirements - Write the code yourself: this is a builder seat, not a management seat. - Bonus: TypeScript in production, FastAPI, vector databases, fintech or fundraising. Benefits - Fully remote and async work environment. - Day overlaps with US Eastern time for a few hours. - Meetings batch on Mondays and Thursdays; the rest is deep work. - The best AI tooling, paid (Claude Code, Cursor, top models). - Real ownership of the core product, with equity potential for the right person. Company Description Paires is where founders come to raise capital. We pair them with the right investors from a large, engaged global investor network, then our agents run the warm outreach and manage the relationships that turn into meetings.
AI/ML Engineer
JR Software Solutions, Inc.JRSS is an IT consulting service provider for Government and Fortune 500 clients. As a certified WOSB, SDB, and HUBZone small business, we specialize in Federal Civilian Data, Cloud, and AI/ML enterprise-level solution architecture, design, implementation, and professional staffing. Recognized on the Inc. 5000 list of fastest-growing private companies, we're proud to foster a collaborative culture where innovation and client impact come first.
Role Description JRSS is seeking an AI/ML Engineer. True innovation happens where machine learning meets cloud technology and real-world impact. As an AI/ML Engineer, you'll join a collaborative team of technologists, data scientists, and stakeholders to tackle meaningful challenges using ML, Generative AI, and modern tools. You'll contribute to building and scaling intelligent systems — from core ML models to chatbot and Retrieval-Augmented Generation (RAG) applications. With strong skills in Python, SQL, and cloud platforms like Azure, you'll help deliver practical, forward-looking solutions in dynamic environments. Bring your technical expertise, curiosity, and customer-first mindset to help shape the future of intelligent systems across industries. Core Responsibilities - Design, develop, and deploy scalable machine learning models and AI-driven solutions to address complex business and operational challenges. - Build and enhance Generative AI, LLM, and Retrieval-Augmented Generation (RAG) applications, including chatbot and conversational AI capabilities. - Develop and optimize data pipelines, feature engineering workflows, and large-scale data processing solutions using Python, SQL, and Spark. - Implement and support MLOps practices, including model training, deployment, monitoring, and lifecycle management using tools such as MLflow and Azure cloud services. - Collaborate with cross-functional teams and stakeholders to deliver customer-focused AI/ML solutions while staying current on emerging technologies and industry best practices. Qualifications - 3+ years of experience designing, developing, and deploying machine learning models. - 3+ years of experience with Generative AI, LLMs, or RAG applications. - 4+ years of hands-on experience with Python for ML and data engineering. - Experience with SQL for data manipulation and feature engineering. - Experience with big data tools such as Apache Spark. - Experience with Databricks and MLOps tools like MLflow. - Experience with cloud, preferably in Azure. - Ability to exhibit strong communication and customer-facing skills. - Ability to thrive both independently and in cross-functional teams. - Ability to problem solve and stay current with emerging ML trends. Nice If You Have - Experience with full-stack development or deploying end-to-end ML applications. - Experience with chatbot development or conversational AI. - Experience fine-tuning large models. - Experience with deploying ML solutions using MLOps pipelines. - Knowledge of Agile workflows and tools like JIRA. Benefits - Join a fast-growing, award-winning team that values professional development, collaboration, and innovation. - At JRSS, our employees are our greatest asset — and your work will directly support high-impact federal and enterprise missions. Company Description JRSS is an IT consulting service provider for Government and Fortune 500 clients. As a certified WOSB, SDB, and HUBZone small business, we specialize in Federal Civilian Data, Cloud, and AI/ML enterprise-level solution architecture, design, implementation, and professional staffing. Recognized on the Inc. 5000 list of fastest-growing private companies, we're proud to foster a collaborative culture where innovation and client impact come first.
