Job Closed
This listing is no longer active.
Remote first tech projects
ML Ops Engineer
Location
Czechia
Posted
89 days ago
Salary
0
Seniority
Senior
Job Description
ML Ops Engineer
Pragmatike
• Build and operate production-grade model serving infrastructure using frameworks such as vLLM, TGI, Triton, or equivalent • Design and implement robust deployment pipelines with blue/green and canary rollout strategies for ML models • Develop and maintain auto-scaling systems, multi-model serving architectures, and intelligent request routing layers • Optimize GPU utilization, memory efficiency, network throughput, and model artifact storage performance • Design observability systems for tracking inference latency, throughput, GPU usage, cost metrics, and system health • Manage model registries and CI/CD pipelines enabling automated and reproducible model deployments • Own the full lifecycle of ML systems from development through production, including operational support and on-call responsibilities • Define engineering best practices and contribute to platform scalability in a fast-moving startup environment
Job Requirements
- 4+ years of experience in ML Ops, Platform Engineering, SRE, or similar infrastructure roles focused on ML systems
- Hands-on experience with model serving frameworks such as vLLM, TGI, Triton, or equivalent
- Strong background in container orchestration and operating GPU-based workloads in production
- Experience with MLOps tooling including model registries, experiment tracking, and automated deployment pipelines
- Proficiency in Python and infrastructure-as-code tools (e.g., Terraform, Helm, or similar)
- Strong understanding of distributed systems, performance tuning, and production reliability engineering
- Ability to effectively use AI coding assistants to accelerate development and debugging workflows
- Ownership mindset with the ability to operate independently in a remote-first environment
Benefits
- Agile working environment
- Personal development opportunities
- Flexible working hours
- Remote work options
Related Guides
Related Job Pages
More Machine Learning Engineer Jobs
Machine Learning Engineer
Leap ToolsLeap Tools is an equal opportunity employer committed to fostering an inclusive, equitable, and accessible environment. Accommodations are available on request for candidates taking part in all aspects of the interview process. If you require any accommodation, please contact us at ta@leaptools.com.
Role Description At Leap Tools, we are building the world's most advanced solutions for the interior décor industry. Our technology lets you preview products in your own room before you buy them. You can redesign your home and ensure that new tiles or rugs will look good in your space through our proprietary cutting-edge technology. What You'll Do - Innovative Vision: Spearhead machine learning initiatives to instantly generate and visualize a vast array of home décor items, transforming browsers into personalized interior design studios. - Holistic Redecoration: Use state-of-the-art algorithms to offer users an all-encompassing redecoration experience, seamlessly integrating wallpaper, flooring, countertops, and a plethora of other home essentials. - Adaptive Product Rendering: Implement machine learning techniques to adapt and present décor products in real-time based on user preferences, room dimensions, and existing furnishings. - User-Centric Design: Work closely with interior design experts, UX/UI designers, and development teams to constantly refine models and enhance realism. Qualifications - Your expertise in Machine Learning isn’t just about models and algorithms; it’s about driving tangible, user-centric solutions. - You are willing and able to tackle unsolved problems without supervision or much guidance after you collaborate with your team members on initial designs. - You are persistent, despite days, weeks, or even months of disappointing results, maintaining the same level of determination as on day one. - You have a strong foundation in computer science and are very comfortable coding in Python. - You have read papers on machine learning topics and implemented your own versions of the papers. - You're hands-on with data: from collection and cleaning to putting together evaluation sets. - Challenges? Bring them on. You thrive when you're pushing the boundaries of what Machine Learning can achieve. Requirements - Mostly PyTorch, some Tensorflow, as well as our own in-house systems. - Python and Django. - gRPC. - Kubernetes on AWS. Benefits - Remote-first company that encourages employees to work from where they're most productive. - Tight-knit teams to cultivate an ownership mentality. - Encouragement of curiosity and attention to detail. - Work anywhere in the world for up to 3 months! - Parental leave program. - Work-from-home stipend. - Your birthday (and our company's birthday) is a day off! Company Description Leap Tools is an equal opportunity employer committed to fostering an inclusive, equitable, and accessible environment. Accommodations are available on request for candidates taking part in all aspects of the interview process.
Senior AI/ML Engineer
eSimplicityAn engineering firm that delivers high-quality Healthcare IT, Cybersecurity, and Telecommunication solutions.
• Architect, implement, and productionize ML solutions (supervised/unsupervised, NLP, deep learning) with robust data preprocessing, feature engineering, and evaluation pipelines. • Lead model selection, training, validation, optimization, and calibration, ensuring reliability, fairness, and performance at scale. • Establish MLOps workflows (CI/CD for ML, experiment tracking, model registry, reproducible builds and deployments). • Implement model monitoring (drift, data/feature quality, bias, and business KPIs), alerting, and automated rollback to keep systems safe and responsive. • Design high-quality data pipelines (ingest, transform, validate) across structured and unstructured sources; enforce data contracts and lineage. • Partner with analytics teams to make datasets discoverable, documented, and performant for iterative model development. • Build AI agents that operationalize safety analytics (Copilot Studio, Python agents, retrieval pipelines) to accelerate triage and decision flow. • Integrate agents with APIs, event streams, dashboards, and case management systems to reduce cycle time from signal to action. • Champion secure-by-design practices, reproducibility, and auditability (model cards, data sheets, deployment records). • Contribute to coding standards, code reviews, and knowledge sharing; mentor engineers and data scientists. • Work in Agile teams; drive iterative delivery, joint problem-solving, and continuous improvement. • Translate mission goals into technical roadmaps and measurable outcomes tied to Sentinel time-to-intervention targets. • Provide technical vision and direction to complex model-related initiatives. • Offer guidance and oversight to junior personnel and contribute as a hands-on expert in the field. • Engage closely with project managers, client representatives, and cross-functional teams to provide timely updates, resolve issues, and ensure alignment with business goals. • Translate technical specifications into code and design documents.
Lead D&T Machine Learning Engineer
General MillsWe exist to make food the world loves. But we do more than that. Our company is a place that prioritizes being a force for good, a place to expand learning, explore new perspectives and reimagine new possibilities, every day. We look for people who want to bring their best — bold thinkers with big hearts who challenge one another and grow together. Because becoming the undisputed leader in food means surrounding ourselves with people who are hungry for what’s next.
Role Description General Mills, Digital and Technology India, is seeking a Lead ML Engineer to join our dynamic and innovative Global Data Science team. In this role, you are a critical member of the data science group focused on leading efforts in migrating ML-based solutions from concept to production-level operational excellence. You will lead initiatives building scalable, resilient, and automated solutions in GCP (Google Cloud Platform) to ensure that models deliver on organizational objectives. - Establish and Implement MLOps practices: - Development of end-to-end MLOps framework and Machine Learning Pipeline using GCP, Vertex AI, and Software tools. - Serving Pipeline with multiple creation Vertex AI and GCP services. - Improve ML pipeline documentation and understandability. - Automate logging of model usage and predictions provided. - Improve logging and diagnostic processes. - Automate monitoring of models both for failures and degradation. - Automate monitoring of data sources to identify issues and/or data changes. - Design and implement dynamic re-training of ML pipelines using event-based or custom logic. - Resource and Infra Monitoring configuration and pipeline development using GCP service. - Branching strategies and Version Control using GitHub. - ML Pipeline orchestration and configuration using Airflow/Kubeflow. - Code refactorization & coding best practices implementation as per industry standard. - Implementing MLOps practices on a project and establishing MLOps best practices: - Lead the investigation and resolution of production issues, perform root cause analysis, and recommend changes to reduce/eliminate re-occurrence of issues. - Optimize deployment and change control processes for models. - Create and operationalize quality assurance processes for ML models. - Lead the execution of ML Solutions @Scale: - Partners with business stakeholders to design the right deliver value-added insights and intelligent solutions through ML and AI. - Collaborates with Data Science Leads, ML System Engineering and Platform teams to ensure the models are deployed in a scaled and optimized way. - Ensure support post-production to manage model performance degradation proactively. - Play a lead role in spearheading the development effort of new standards (design patterns, coding practices, orchestration patterns) and drive value and adoption across the Data Science team. - Is considered an expert in the ML Ops and Model management space. - Research, Evolve and Publish best practices: - Research and operationalize technology and processes necessary to scale ML Ops. - Recommend model changes to optimize cloud spend. - Ability to research and recommend MLOps best practices on new technologies, platforms, and services. - Drive ideation, design, and creation of new ML Architecture patterns in discussion with the Enterprise Architecture team. - MLOps pipeline improvement plan and suggestion. - Communication and Collaboration: - Knowledge sharing with the broader analytics team and stakeholders. - Communicate on the on-goings to embrace the remote and geographical culture. - Ability to communicate the accomplishments, failures, and risks in a timely manner. - Knowledge sharing session with team for specific ML Ops topics. - Coach and Mentor junior ML members in the team. - Foster a collaborative and innovative team environment. - Contribute to the overall effort to educate stakeholders on AI practices. - Closely collaborates with the stakeholders on projects and data science leaders. - Embrace a learning mindset: - Continually invest in one’s knowledge and skillset through formal training, reading, and attending conferences and meetups. Qualifications - Full-time graduate from an accredited University. - Advanced degree in a quantitative field (CS, engineering, statistics, math, data science). - Proven technical leadership in a large, complex matrixed organization. - Relevant Machine Learning experience of 6+ years and overall 12+ years of Industry experience. - Experience in supervised ML algorithms, optimization, and performance tuning. - Track record of producing machine learning models and production infrastructure at scale. - Strong verbal and written communication skills including the ability to interact effectively with colleagues of varying technical and non-technical abilities. - Passionate about agile software processes, data-driven development, reliability, and systematic experimentation. - Passion for learning new technologies and solving challenging problems. - Good understanding of CI, CD, TDD, and tools such as Jenkins. - Strong understanding of orchestration frameworks such as Airflow/Kubeflow/MLFlow. - Agile software development experience such as Kanban and Scrum. - Experience in software version control team practices and tools such as GIT and TFS. - Expertise in Data Transformation and Manipulation through Big-Query/SQL. - Professional experience with Vertex AI and GCP Services. - Strong proficiency in Python. Preferred Qualifications - GCP Machine Learning certification. - Understanding of CPG industry. - Exposure to Deep Learning/RL/LLMs. - Prior experience with CPG industry. - Publications or contributions to the data science and AI community. - Certifications in AI, machine learning, or related fields.
Senior Full-Stack Machine Learning Engineer – GenAI, Agentic Systems
PoppuloBetter Communications, Better Outcomes. Enterprise-grade employee communications and digital signage software.
• Own end-to-end delivery of AI features across model, backend, APIs, UI integration, deployment, monitoring, and iteration in production • Solve complex challenges with AI/ML: Design, develop new AI-powered products that deliver the product roadmap, including agentic AI solutions that orchestrate LLMs, tools, and workflows to solve multi-step problems autonomously. • Implement ML lifecycle - from data engineering and model development to cloud-based deployment, integrations and operationalisation. Incl. MLOps • Productionise full-stack AI/ML solutions: Translate emerging techs like GenAI& agenticAI architectures into innovative, practical solutions that transform customer experiences. • Align with Product Strategy: Create proof of concepts at high cadence to demonstrate/validate potential solutions as per our product strategy. • Optimise Model and system performance: Fine-tune, optimise training and inference performances, including latency, cost, and reliability trade-offs in agent-based and LLM-driven systems. • Wider collaboration: Partner with cross-functional teams to demonstrate and validate the impact of ML innovations before introducing them into the product ecosystem. • Research Savvy: Staying up-to-date with SOTA and industry trends in AI/ML, with a strong awareness of advances in agentic systems, autonomous workflows, and multi-agent architectures.



