Staff Machine Learning Engineer

Machine Learning EngineerMachine Learning EngineerFull TimeRemoteLeadTeam 1,001-5,000Since 2007H1B No SponsorCompany SiteLinkedIn

Location

Texas

Posted

6 days ago

Salary

$206.3K - $330.0K / year

Seniority

Lead

Bachelor Degree10 yrs expEnglish

Job Description

Staff Machine Learning Engineer

Triumph Financial, Inc.

• Work closely in a small, cross-functional team of 3-4 people focused on AI/ML systems. • Collaborate closely with product managers and other stakeholders. • Understand customer pain points and deliver impactful solutions. • Build, deploy, and integrate models that generalize across diverse document formats. • Explore new techniques in deep learning, transfer learning, and model optimization.

Job Requirements

  • 10+ years of software engineering experience
  • 4+ years with production ML at scale
  • Proven success designing and operating distributed ML systems.
  • Strong experience with model deployment patterns and data pipelines.
  • Demonstrated ability to influence technical decisions across teams.
  • Deep understanding of reliability engineering and production support.

Benefits

  • Medical
  • Dental
  • Vision
  • Paid Time Off
  • 401k and much more

Related Job Pages

More Machine Learning Engineer Jobs

24-MAG logo

Machine Learning Research Engineer

24-MAG

24-MAG is a next-gen management consulting partner built to help B2B organizations move faster and operate smarter.

Role Description We are sharing a specialised full-time consulting opportunity for machine learning engineers and research practitioners with hands-on experience training, evaluating, and experimenting with ML models end to end. This role supports the development of advanced agentic evaluation benchmarks for frontier AI systems. Selected professionals will transform real machine learning research ideas into rigorous multi-step tasks, implement and run experiments, analyse training behaviour, and evaluate where model-generated solutions fall short of technically correct results. Key Responsibilities - Machine Learning Task Design - Turn practical ML research ideas into well-defined, multi-step evaluation tasks - Develop assignments involving model training, experimental modifications, and performance analysis - Define clear technical requirements, expected outputs, and success criteria - Ensure tasks assess genuine implementation and experimental reasoning rather than superficial library usage - Experiment Implementation & Execution - Implement reference solutions using Python, scripts, and notebook environments - Configure and run model-training experiments from setup through final evaluation - Modify model components, training procedures, reward functions, or experimental parameters - Validate code, dependencies, datasets, intermediate outputs, and final results - Document complete workflows so experiments can be reproduced independently - Model Evaluation & Analysis - Review how frontier AI models approach complex machine learning tasks - Assess implementation quality, experimental methodology, and technical conclusions - Identify coding errors, unsupported assumptions, weak experimental controls, and misleading interpretations - Determine whether reported improvements are supported by the observed results - Explain clearly where and why a model-generated solution fails - Reinforcement Learning Experiments - Develop selected tasks involving reinforcement learning fundamentals - Evaluate reward-function changes, policy-training behaviour, and experimental outcomes - Assess whether proposed modifications produce the intended training effect - Identify instability, unintended incentives, or incorrect interpretations of RL results - Research Collaboration - Work closely with researchers, task authors, and fellow machine learning specialists - Compare evaluation decisions to maintain consistent and rigorous benchmark standards - Refine task instructions, reference solutions, and grading criteria based on testing outcomes - Share recurring model failure patterns and opportunities for stronger benchmark coverage Qualifications - At least 1 year of experience in machine learning research, research engineering, or a comparable technical role - Hands-on experience training and evaluating ML models through complete experimental workflows - Strong understanding of experiment setup, execution, analysis, and reproducibility - Familiarity with large language model capabilities, limitations, and evaluation techniques - Working proficiency in Python and Git - Comfort using both scripting and notebook-based environments - Strong technical writing, analytical reasoning, and attention to detail - Ability to work independently through ambiguous, open-ended research problems - Reliable availability for approximately 35 hours per week Educational Background - A master's degree or PhD in machine learning, computer science, artificial intelligence, engineering, mathematics, or another relevant STEM discipline is highly relevant - Equivalent practical experience in a research-intensive machine learning role may also be considered - Academic or professional work involving model training, experimentation, or ML systems may strengthen an application - Publications, open-source contributions, technical reports, or substantial research projects may also be valuable Nice to Have - Understanding of reinforcement learning concepts, including reward functions and policy training - Experience in AI training, model evaluation, or benchmark development - Background authoring technical tasks, reference solutions, or grading rubrics - Familiarity with agentic AI systems and multi-step model evaluations - Experience diagnosing model-training failures or unexpected experimental behaviour - Knowledge of experimental design, ablation studies, and performance comparison - Experience reviewing code, notebooks, or research analyses prepared by other practitioners - Familiarity with reproducible ML environments and collaborative Git workflows Why This Opportunity - Apply practical machine learning research expertise to frontier AI evaluation - Design realistic tasks grounded in end-to-end model experimentation - Help improve how AI systems approach implementation, training, and analytical reasoning - Work across Python, ML evaluation, reinforcement learning, and reproducible research - Collaborate closely with AI researchers and machine learning specialists - Participate in a structured full-time remote role with competitive hourly compensation Contract Details - Full-time W-2 contingent employment opportunity - Fully remote within the United States - Expected commitment of approximately 35 hours per week - Competitive rates between $55–$85 per hour depending on expertise and project scope - Individual tasks may require one to two days of focused implementation and experimental work - Work may include task design, model training, experiment execution, notebook development, AI output evaluation, and technical reporting - Engagement scope and duration may evolve according to project requirements and performance About the Platform This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

United States
$55 - $85 / hour
Full TimeRemoteTeam 51-200Since 2022

• Build agent memory systems: not just picking what goes into context, but the mechanisms that generate, curate, refine, and store that information in the first place. • Design memory with real constraints: confidentiality and scoping so agents never leak what they shouldn't. • Build systems that watch how agents behave across the org and turn that into shared best practices at scale. • Build reusable "skills" agents can call on: better reasoning, better financial decisions, better report writing. • Design and run tests and benchmarks that show whether these improvements actually work. • Help shape the technical roadmap for agent memory and reasoning as the team stands up. • Take an undefined problem and design a real, shippable solution for it. • Document your methodology clearly enough that others can build on it.

United States

MLOps Engineer

Bright Vision Technologies

Bright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.

Role Description We are seeking a MLOps Engineer to design, build, and operate high-performance, highly reliable inference platforms for serving large machine learning models in production. The role focuses on the systems engineering side of AI deployment, including: - Request routing - Batching - Caching - Autoscaling - GPU utilization - End-to-end observability across diverse model workloads The ideal candidate brings strong distributed systems and performance engineering expertise, has shipped serving systems at scale, and understands the trade-offs between latency, throughput, cost, and quality in ML serving. Qualifications - Bachelor’s or Master’s degree in Computer Science or a related field. - Six or more years of experience in distributed systems, infrastructure, or ML platform engineering. - Strong proficiency in Python and a systems language such as Go, Rust, or C++. - Deep experience operating high-throughput, low-latency services in production. - Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM. - Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization. - Familiarity with Kubernetes, autoscaling, and modern cloud platforms. - Experience with observability stacks including metrics, tracing, and structured logging. - Solid grounding in performance engineering and capacity planning. - Strong communication and incident response skills. Requirements - Design and operate model serving platforms supporting diverse workloads including LLMs, vision models, and recommendation systems. - Optimize inference performance using continuous batching, paged attention, speculative decoding, and request multiplexing. - Implement multi-tenant routing, rate limiting, and quality-of-service policies across model endpoints. - Build autoscaling and capacity management systems that balance latency, throughput, and cost. - Tune GPU utilization, memory management, and KV cache strategies for LLM serving workloads. - Integrate model serving with API gateways, identity systems, and observability platforms. - Implement caching, prompt deduplication, and response reuse strategies where appropriate. - Drive end-to-end observability including latency histograms, queue dynamics, GPU utilization, and error tracking. - Develop deployment workflows including canary releases, shadow testing, and automated rollback. - Operate incident response for high-availability AI services and drive durable reliability improvements. - Collaborate with ML and product teams to support new model releases and capability rollouts. - Implement security controls including request signing, content filtering, and abuse detection at the serving layer. - Document operational procedures, performance characteristics, and tuning guidance for internal teams. - Stay current with AI serving research and translate advances into production capabilities. Benefits - 100% Remote (U.S.) - Full-time, Direct W2 - Salary Range: $100,000–$150,000 Annually - Sponsorship for U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates. How to Apply Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3544. Learn more about Bright Vision Technologies at www.bvteck.com .

United States
$100K - $150K / year
Job Closed
Vancity logo

Senior Machine Learning Engineer

Vancity

Put your money where your values are.

Full TimeRemoteTeam 1,001-5,000Since 1946

• Applying Data Science and Machine Learning best practices to develop robust models and support data-driven decision-making across business domains • Applying machine learning and data science techniques such as forecasting, predictive modeling, classification, regression, recommendation, and optimization to solve business problems • Conducting experiments and evaluating models using appropriate statistical, technical, and business performance metrics • Architecting, building, deploying, and maintaining scalable machine learning models and AI solutions integrated into enterprise systems, applications, and operational workflows • Designing and implementing end-to-end ML workflows, including data preparation, feature engineering, model training, validation, deployment, optimization, and continuous monitoring in a high-scale production environment • Developing reusable machine learning components, feature pipelines, and model-serving frameworks to support multiple use cases and teams • Designing and implementing production-grade MLOps solutions using Azure ML, Databricks, MLflow, and related cloud technologies • Building and maintaining automated ML pipelines, feature engineering workflows, feature store patterns, and deployment processes for training, testing, monitoring, and retraining machine learning models • Implementing standards and best practices for model versioning, lifecycle management, governance, deployment automation, model performance monitoring, drift detection, data quality, operational health, and retraining triggers • Developing production-quality Python code, APIs, automation workflows, and machine learning services to integrate ML capabilities into business applications and processes

Canada
$113.1K - $153K / year