Accelerating Product Development Through Science-Based AI
Lead Machine Learning Engineer
Location
Mexico
Posted
2 days ago
Salary
0
Seniority
Lead
No structured requirement data.
Job Description
Lead Machine Learning Engineer
NobleAI
Role Description As a Lead Machine Learning Engineer specializing in conversational and agentic systems at NobleAI, you will be responsible for architecting, building, and deploying intelligent features to our VIP platform. This role is ideal for individuals passionate about the cutting edge of LLMs and eager to build AI systems that can reason, plan, and act. - Design domain specific AI systems and chatbots capable of complex dialogue management and workflow execution via tools, API calls and multi-step tasks based on user goals, multi-agent orchestration. - Collaborate with scientists to assess, fine-tune, and deploy LLMs on domain specific data to support accuracy measurement for use cases. - Build and maintain Retrieval-Augmented-Generation (RAG) systems, Reinforcement Learning frameworks, guardrail and assessment mechanisms for end to end lifecycle for customized models. - Collaborate with product and software engineers to integrate the features into our platform. - Establish prompt engineering and data management best practices for transparency and governance. - Establish best practices for monitoring and evaluation of data and models across the model lifecycle (development, testing, and production). - Keep a pulse on the latest advancements in NLP, LLMs, and agentic AI research, and act as a subject matter expert on architecture decisions on platform and use cases. Qualifications - MSc (preferred) or BSc in Computer Science, Artificial Intelligence, or a related quantitative field. - 5+ years of hands-on experience building and deploying AI/ML systems, with a strong focus on Natural Language Processing (NLP). - Proven experience designing and shipping chatbots, virtual assistants, or agentic systems using modern LLM-based architectures. - Strong programming proficiency in Python (5+ years) and deep experience with core ML/NLP libraries such as PyTorch, TensorFlow. - Hands-on experience with LLM agent frameworks for building complex, tool-using applications, multi-agent orchestration, or establishing MCP services for platform capabilities. - Demonstrated experience with Retrieval-Augmented Generation (RAG), including the use of vector databases like Pinecone, Weaviate, or ChromaDB. - Familiarity with techniques for fine-tuning LLMs (e.g., LoRA/QLoRA) and experience working with open-source models (e.g., Llama, Mistral) or major model APIs (e.g., OpenAI, Anthropic). - 5+ years of experience with cloud platforms (Azure preferred) and familiarity with deploying AI models as scalable microservices using Docker and Kubernetes (KFP, KServe). - Solid software engineering fundamentals, including version control (Git), automated testing, and CI/CD principles. - Excellent communication skills with the ability to articulate complex technical ideas to both technical and non-technical stakeholders. Benefits - Benefits coverage including medical, dental, vision, disability, and life insurance. - Retirement fund employer contribution. - Generous Paid Time Off & Holidays. - Stock options. - Performance-based bonus. - Salary range depending on experience.
Related Guides
Related Job Pages
More Machine Learning Engineer Jobs
Machine Learning Engineer
AbbVieA biopharmaceutical company based in Chicago, Illinois, AbbVie makes and markets advanced therapies and medicines to treat serious illnesses and medical conditi
• Own small to medium components of machine learning systems from technical design through implementation and delivery • Translate technical requirements into high-quality, maintainable code and deliver workstreams according to plan • Build and maintain data pipelines and feature engineering workflows to support machine learning and AI solutions • Design, train, evaluate, and refine machine learning models with minimal supervision, applying sound statistical and engineering practices • Implement ML solutions that can be deployed into production environments as microservices, APIs, batch jobs, or streaming components • Support production monitoring efforts by helping define and implement metrics for model performance, data drift, anomalies, and retraining triggers • Collaborate with Data Engineers, Software Engineers, Data Scientists, Product partners, and business stakeholders to deliver project objectives • Understand system design, data models, and technical artifacts well enough to contribute to implementation decisions and tradeoffs • Follow governance, documentation, coding, and source control standards consistently • Demonstrate flexibility and proactively support teammates with day-to-day responsibilities as needed • Clearly document and communicate work progress, technical decisions, and outcomes to technical and non-technical audiences
Role Description Insurance isn’t the first industry most engineers think of when they imagine cutting-edge AI work. That’s exactly why this role is interesting. CFC’s Data & AI unit is building production agentic systems that automate complex underwriting decisions - not chatbots, not copilots bolted onto legacy workflows, but autonomous multi-step AI agents that reason over unstructured data, assess risk, and drive real business outcomes. We’re using frameworks like LangChain to orchestrate LLM-driven services that sit at the heart of how the business operates. The problems are genuinely hard: ambiguous inputs, high-stakes decisions, and the kind of domain complexity that makes for satisfying engineering. We’re also fundamentally rethinking how we deliver software. We’re moving toward an agentic-first development model - using AI agents not just in what we build for the business, but in how we build it. The goal is to multiply engineering delivery by an order of magnitude: not by cutting corners or generating throwaway code, but by designing robust systems and processes that let agents handle well-defined work while engineers focus on architecture, design, and the problems that actually require human judgment. This is a deliberate, engineering-led approach. We care about code quality, testability, and maintainability - the agent-generated code meets the same standards as everything else. We expect that a successful candidate will be able to bring their expertise to help guide and refine our agentic development process as it matures - a meaningful opportunity to influence how we work as well as what we ship. This is a Senior Machine Learning / AI Engineer role with genuine technical leadership scope. You’ll shape the architecture of AI-driven production microservices, own system design decisions, and work at the intersection of traditional software engineering and applied AI. You’ll collaborate closely with engineers, data scientists, and product managers in a team that’s small enough for your decisions to matter and ambitious enough for the work to stay interesting. Qualifications - Significant experience building production-grade AI agents or LLM-powered services. - Strong Python development experience, ideally with 6+ years of professional software engineering experience. - A track record of writing clean, maintainable and high-quality production code. - Experience supporting business-critical systems in live production environments. - Strong understanding of asynchronous programming, Docker, containerised deployments and modern service architectures. - Experience designing distributed systems, asynchronous microservices and event-driven architectures. - Confidence leading system design discussions and making pragmatic architectural trade-offs. - Strong cloud experience, ideally within Azure. - Experience deploying, monitoring and maintaining ML or LLM models in production. - Familiarity with MLOps, LLMOps, evals, monitoring and lifecycle management. - Hands-on experience orchestrating LLM workflows using frameworks such as LangChain. - An interest in how agentic software development can improve engineering delivery without compromising code quality. - Strong communication skills and the ability to collaborate effectively in remote or asynchronous environments. - An ownership mindset, with the ability to work independently and contribute effectively to shared codebases. Requirements - Design, develop, and maintain business-critical AI agent services. - Build tailored agent workflows and services from business requirements, using LangChain and related frameworks with reliable patterns for LLM-driven decision-making in production. - Integrate tests and validation to improve AI agents through evals and monitoring. - Work with software engineers and architects to lead system design and architectural decisions. - Translate technical specifications into clean, testable, and scalable production code. - Work closely with cross-functional teams - engineers, data scientists, product managers - to deliver features on time and to a high standard. - Write unit and integration tests to maintain reliability and service correctness. - Monitor, troubleshoot, and continuously improve production services. - Produce clear, structured documentation for systems, architecture, and processes. - Mentor junior engineers through code reviews, best-practice guidance, and knowledge sharing. Benefits - Love what you do: We show up each day ready to take on the world. Our passion and intensity set us apart and makes the difference to our colleagues, customers, brokers and carriers. - Challenge everything: We’re never afraid to question the way that things are done and we constantly challenge ourselves and others to makes things better. - Have fun, be good: Insurance is a serious business, but we don’t take ourselves too seriously. We make it fun to work at CFC, we welcome all viewpoints, and we treat everyone how we would expect to be treated.
MLOps / LLMOps Engineer
Rootshell Enterprise Technologies, Inc.Rootshell Enterprise Technologies Inc. is a recognized provider of professional IT Consulting services in the US.
Role Description Operationalizing Large Language Models requires specialized expertise beyond traditional MLOps practices. LLMs present unique operational challenges including significantly larger computational requirements, complex data pipelines, specialized infrastructure needs, and unique performance optimization requirements. This specialized role ensures GenAI solutions can scale effectively from proof-of-concept to enterprise-wide deployment in a utility environment. - Ensures GenAI solutions move successfully from prototype to production with proper operational support - Establishes specialized monitoring for model performance, inference latency, and data quality - Enables efficient scaling of LLM solutions across multiple business units - Creates high-performance deployment architectures that balance speed, cost, and reliability - Develops operational data pipelines to continuously improve model performance with new utility-specific data Qualifications - DevOps + ML: Expertise in Kubernetes, Docker, CI/CD tools, and MLflow or similar platforms - Cloud & Infrastructure: Understanding of GPU instance options, cloud services (AWS/Azure/GCP), and optimization techniques - Automation: Proficiency in Python, Bash, and infrastructure-as-code tools like Terraform or Ansible - LLM-Specific Frameworks: Experience with tools like TensorBoard, MLFLow, or equivalent for scaling LLMs - Performance Optimization: Knowledge of techniques to monitor and improve inference speed, throughput, and cost - Collaboration: Ability to work effectively across technical teams while adhering to enterprise architecture standards Requirements - Design and implement LLM-specific deployment architectures with Docker containers for both batch and real-time inference - Configure GPU infrastructure on-premises or in the cloud with appropriate CI/CD pipelines for model updates - Build comprehensive monitoring and observability systems with appropriate logging, metrics, and alerts - Implement load balancing and scaling solutions for LLM inference, including model sharding if necessary - Create automated workflows for model retraining, versioning, and deployment - Optimize infrastructure costs through intelligent resource allocation, spot instances, and efficient compute strategies - Collaborate with PG&E's Cyber team on implementing appropriate security controls for GenAI applications - Develop automated testing frameworks to ensure consistent output quality across model updates
Role Description The Senior Machine Learning Engineer will be responsible for: - Building recommendation and search across feed, discovery, search, and content continuation. - Owning retrieval/ranking: candidate generation, embeddings, two-tower models, features, and serving quality. - Designing, launching, and analyzing recommendation/search experiments. Qualifications - 5+ years industry experience building production ML systems with senior ownership. - Bachelor's degree from a recognized university in China is required. - Hands-on experience with recommendation, search, ranking, ads ranking, feed ranking, or content discovery systems. - Experience with consumer apps, entertainment, social, gaming, creator, or engagement-driven products. - Familiarity with two-tower models, embedding retrieval, candidate generation, ranking, and online/offline evaluation. Requirements - 5+ years industry experience building production ML systems with senior ownership. - Hands-on experience with recommendation, search, ranking, ads ranking, feed ranking, or content discovery systems. - Experience with consumer apps, entertainment, social, gaming, creator, or engagement-driven products. - Familiarity with two-tower models, embedding retrieval, candidate generation, ranking, and online/offline evaluation. Benefits - Salary range: 150k-300k+ equity. - This is a remote position. Anti-signals - Cannot show core Senior Machine Learning Engineer, Recommendation experience. - Not comfortable with the listed work mode. - Low ownership, coordination-only, or no shipped examples.

