Job Closed
This listing is no longer active.
Meaningful classroom learning for every student.
AI Operations Engineer
Location
Latin America
Posted
119 days ago
Salary
0
Seniority
Senior
Job Description
AI Operations Engineer
Newsela
• Design and maintain CI/CD pipelines for ML model training, packaging, and deployment across our microservices. • Manage containerized services on AWS ECS, optimizing for cost, latency, and availability. • Automate infrastructure provisioning and service configuration with Terraform. • Work to maintain and scale services that make use of third party LLM providers. • Build and improve data pipelines that feed models from BigQuery, S3, and DynamoDB into training and inference workflows. • Instrument services with observability tooling (Datadog, OpenTelemetry, Langfuse) and establish SLOs for model-serving endpoints. • Collaborate with ML engineers to productionize new models using BentoML, FastAPI, and container-based serving.
Job Requirements
- 2-3 years in ML Ops supporting ML/AI features, systems and workflows with 3-4 years prior experience in DevOps, CloudOps or SRE.
- Strong proficiency in Python.
- Hands-on experience with Docker containerization and container orchestration.
- Solid understanding of CI/CD for ML workflows in an enterprise production environment.
- Experience with Infrastructure as Code, preferably Terraform.
- Familiarity with cloud platforms — specifically AWS (ECS, ECR, S3, DynamoDB, CloudWatch) and GCP (BigQuery, Vertex AI).
- Experience with LLM integration and observability (OpenAI API, Google GenAI, Langfuse tracing).
- Experience building and maintaining data pipelines for ML training and feature engineering
- Familiarity with ML modeling workflows — training, evaluation, experiment tracking (e.g. MLFlow, Weights & Biases), and model versioning
- Experience monitoring and flagging model drift over time.
- Exposure to NLP/NLU models and frameworks such as Hugging Face Transformers, spaCy, or sentence-transformers
- Knowledge of vector databases (LanceDB, FAISS) and embedding-based retrieval systems
- Experience with scaling and maintaining deep learning frameworks (TensorFlow, PyTorch) in production settings
- Familiarity with classical ML libraries (scikit-learn, XGBoost, LightGBM) and model explainability tools (SHAP)
- Working knowledge of ML serving frameworks such as BentoML or similar.
- Comfort working with FastAPI or similar async Python web frameworks.
Benefits
- Competitive salary
- Professional development opportunities
Related Guides
Related Categories
Related Job Pages
More Artificial Intelligence Jobs
• Lead and expand strategic partnerships with GPU and AI ecosystem leaders (NVIDIA, AMD, hyperscalers, OEMs, and AI infrastructure providers). • Develop and execute joint go-to-market strategies and partner programs that position Cologix as a premier destination for AI and GPU deployments. • Work cross-functionally with Solution Engineering, Development/Construction, Product, Marketing, Operations, and Sales to shape GPU-optimized data center offerings, deployment architectures, and technical integration standards. • Develop reference architectures, as well as early-access or beta collaboration programs. • Serve as Cologix’s senior representative in partner roadmap discussions, industry summits, and technical councils focused on AI and high-performance compute. • Track semiconductor and AI industry trends — including emerging accelerator technologies, GPU supply chain shifts, and AI workload evolution — to inform corporate strategy. • Identify and evaluate new partnership opportunities across the AI and data center ecosystem.
AI Agent Trainer & Technical Ops Integrator
Magnolia Home InspectionsWe Prepare and Protect Home Buyers
Role Summary Join our entrepreneurial team and contribute to developing cutting-edge AI systems, while enjoying the flexibility of remote work. We are looking for a proficient Software Engineer - AI Trainer to help advance the build of our human-centered AI Operating System that supports operations of multiple businesses we own. This role owns the execution layer of an AI-driven operating environment. You will implement and manage a multi-agent AI system that supports real-world business operations, translating direction into working infrastructure across logistics, communication, and internal workflows. We are building more than a raw platform: we bring agents to life, giving them clear purpose and specialized training, and orchestrate them into self-governing teams — all while keeping humans firmly at the center of strategy, creativity, relationships, and purpose. You are responsible for improving functioning systems quickly and reliably. The work is hands-on, technical, and operational—ensuring stability while continuously improving how the business runs. Core Responsibilities 1. Logistics Continuity & Coverage - Maintain uninterrupted scheduling and coordination of inspections - Preserve quality and consistency in vendor and client communications - Fully assume logistics responsibilities during coverage periods without escalation 2. Systemization & Redundancy - Document all logistics workflows (SOPs, checklists, contingencies) - Build structured systems that allow seamless role substitution - Reduce dependency on any single individual 3. Technical Development & Automation - Build and maintain internal tools using Linux-based environments and AI coding platforms (Cursor, Codex, ClaudeCode, OpenClaw) - Convert leadership direction into functional systems and workflows - Deliver consistent, production-ready outputs on a weekly cadence 4. Executive Technical Execution - Act as a technical executor for Benjamin - Translate strategy into completed systems, tools, and deliverables - Eliminate bottlenecks between planning and execution - Execute assigned initiatives within 24–48 hour turnaround 5. Continuous Improvement - Identify inefficiencies across logistics and internal systems - Implement structured improvements on a monthly basis - Enhance reliability, speed, and clarity across operations Success Outcomes - Zero disruption in logistics during coverage periods - Fully documented and transferable logistics system - Weekly delivery of functional technical outputs - Reduced execution lag between idea and implementation - Measurable improvements in operational efficiency Ideal Candidate Profile - Prefers structured, technical work over ambiguous or high-risk decision-making - Strong attention to detail and commitment to accuracy - Comfortable working independently once expectations are clear - Communicates in a direct, factual manner - Values clear processes, defined expectations, and stability - Demonstrates reliability, consistency, and follow-through - Capable of learning systems deeply and executing them precisely - Motivated by producing correct, complete work rather than rapid experimentation Management & Environment Fit - Clear objectives and defined expectations will be provided - Initial guidance and structured onboarding will be extensive - Increasing autonomy will be granted as competency is demonstrated - Feedback will be direct, private, and specific - Role emphasizes trust, consistency, and ownership over time
AI Agent Trainer & Technical Ops Integrator
Magnolia Home InspectionsWe Prepare and Protect Home Buyers
Role Summary Join our entrepreneurial team and contribute to developing cutting-edge AI systems, while enjoying the flexibility of remote work. We are looking for a proficient Software Engineer - AI Trainer to help advance the build of our human-centered AI Operating System that supports operations of multiple businesses we own. This role owns the execution layer of an AI-driven operating environment. You will implement and manage a multi-agent AI system that supports real-world business operations, translating direction into working infrastructure across logistics, communication, and internal workflows. We are building more than a raw platform: we bring agents to life, giving them clear purpose and specialized training, and orchestrate them into self-governing teams — all while keeping humans firmly at the center of strategy, creativity, relationships, and purpose. You are responsible for improving functioning systems quickly and reliably. The work is hands-on, technical, and operational—ensuring stability while continuously improving how the business runs. Core Responsibilities 1. Logistics Continuity & Coverage - Maintain uninterrupted scheduling and coordination of inspections - Preserve quality and consistency in vendor and client communications - Fully assume logistics responsibilities during coverage periods without escalation 2. Systemization & Redundancy - Document all logistics workflows (SOPs, checklists, contingencies) - Build structured systems that allow seamless role substitution - Reduce dependency on any single individual 3. Technical Development & Automation - Build and maintain internal tools using Linux-based environments and AI coding platforms (Cursor, Codex, ClaudeCode, OpenClaw) - Convert leadership direction into functional systems and workflows - Deliver consistent, production-ready outputs on a weekly cadence 4. Executive Technical Execution - Act as a technical executor for Benjamin - Translate strategy into completed systems, tools, and deliverables - Eliminate bottlenecks between planning and execution - Execute assigned initiatives within 24–48 hour turnaround 5. Continuous Improvement - Identify inefficiencies across logistics and internal systems - Implement structured improvements on a monthly basis - Enhance reliability, speed, and clarity across operations Success Outcomes - Zero disruption in logistics during coverage periods - Fully documented and transferable logistics system - Weekly delivery of functional technical outputs - Reduced execution lag between idea and implementation - Measurable improvements in operational efficiency Ideal Candidate Profile - Prefers structured, technical work over ambiguous or high-risk decision-making - Strong attention to detail and commitment to accuracy - Comfortable working independently once expectations are clear - Communicates in a direct, factual manner - Values clear processes, defined expectations, and stability - Demonstrates reliability, consistency, and follow-through - Capable of learning systems deeply and executing them precisely - Motivated by producing correct, complete work rather than rapid experimentation Management & Environment Fit - Clear objectives and defined expectations will be provided - Initial guidance and structured onboarding will be extensive - Increasing autonomy will be granted as competency is demonstrated - Feedback will be direct, private, and specific - Role emphasizes trust, consistency, and ownership over time
• Notion Systems - You are the internal expert. • Automations & Integrations - You design and deploy workflows across our stack. • AI & Workflow Optimization - You experiment with LLMs and emerging tools. • Tech Stack & Access Management - You own user permissions and data hygiene. • Strategic Technology Advisory - You stay current on technology trends and best practices.



