Staff ML Software Engineer
Location
United States
Posted
4 days ago
Salary
$140K - $293K / year
Seniority
Lead
No structured requirement data.
Job Description
Staff ML Software Engineer
Indeed.com
Role Description As a Software Engineer IV (ML) on the Machine Learning Model Platform team at Indeed, you will be responsible for leading and executing key objectives for the Model Platform team, which includes providing support for critical entities like Matching and Recommendation systems, ML model training portal, etc. You will be expected to design and build high-performance, reusable components utilized by hundreds of ML practitioners to transition from research to production impact, covering the entire ML model lifecycle. In this role, you will operate at the intersection of software engineering and machine learning, developing foundational components that facilitate diverse ML model strategies with exceptional performance and scalability. You will engage in close collaboration with Data Scientists to ascertain their requirements, spearhead technical design choices, contribute to workflow optimization, and directly support the achievement of their key objectives. You'll play a critical role in evolving our tech stack to stay in line with cutting-edge industry technologies. Qualifications - Requires a minimum of 8 years of related experience with a Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or Mathematics; or 6 years and a Master’s degree; or a PhD with 3 years experience; equivalent experience may substitute for degree requirements. - Expertise in Python and modern ML frameworks like PyTorch & Triton. - Experience with AWS SageMaker or other cloud-based ML platforms. - Proficiency in software design, data structures, algorithms, and computer science fundamentals. - Experience designing, building, and operating scalable, reliable software systems or platforms. - Demonstrated ownership and accountability for technical outcomes and system quality. - Excellent collaboration and communication skills, with the ability to influence technical direction across teams. Requirements - Own the design, development, and evolution of complex systems, frameworks, or platforms. - Drive technical decision-making, balancing short-term delivery with long-term maintainability and scalability. - Architect new solutions, evaluate trade-offs, and validate ideas through prototyping, experimentation, or iteration on existing systems. - Participate in and influence code and design reviews across teams to uphold high engineering standards. - Identify performance, reliability, and scalability improvements and drive enhancements to existing systems. - Mentor and guide other engineers, supporting technical growth and best practices across teams. - Communicate clearly and effectively with engineers, product managers, and other business partners to align on technical direction and execution. Benefits - Quarterly bonuses. - Restricted Stock Units (RSUs). - Paid Time Off policy. - Region-specific benefits.
Related Guides
Related Job Pages
More AI Engineer Jobs
Senior AI Engineer - Foundation Models and Transformers-2
MastercardFounded in 1966, Mastercard is a worldwide transaction, payment-processing, and consulting company best known for its line of personal and business credit cards. As an employer, Ma
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we're helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior AI Engineer - Foundation Models and Transformers-2 Overview Mastercard is seeking a Senior AI Engineer to design, build, and deploy high-quality AI solutions that support key business and product initiatives. This role is hands-on and delivery-focused, contributing directly to the development of production AI systems while collaborating closely with product, data, and engineering partners. As a Senior AI Engineer, you will work on well-scoped AI initiatives, applying advanced machine learning and software engineering practices to move models from experimentation into reliable, performant production systems. This role represents a critical technical contributor level, with opportunities to grow toward technical leadership and broader system ownership. Role In this role, you will be responsible for building and operationalizing AI solutions under the guidance of Lead and Principal engineers. Key responsibilities include: Design, develop, and deploy AI and machine learning models to solve defined business and product problems Contribute to the development and optimization of transformer-based and generative AI models, including fine-tuning, evaluation, and inference workflows Build and maintain data pipelines, feature engineering logic, and training workflows in collaboration with data engineering teams Implement model serving and inference solutions, integrating models into downstream applications and APIs Apply MLOps best practices, including experiment tracking, versioning, automated testing, monitoring, and model performance evaluation Participate in code reviews, design discussions, and technical planning to ensure high-quality, maintainable solutions Collaborate with product managers and stakeholders to translate requirements into technical implementations Support troubleshooting and performance tuning of models and AI systems in production environments Continuously improve technical skills and stay current with advances in AI, ML frameworks, and engineering practices All About You Solid experience developing machine learning or AI solutions and deploying them into production environments Strong proficiency in Python and experience with ML frameworks such as PyTorch and/or TensorFlow Hands-on experience with transformer-based models (e.g. BERT-style encoders, generative models, embeddings, or similar architectures) Experience working with data pipelines and datasets, including data preparation, feature engineering, and training data management Familiarity with cloud platforms (AWS, Azure, or GCP) and cloud-based ML tooling Working knowledge of MLOps practices, including model deployment, monitoring, and lifecycle management Strong software engineering fundamentals, including version control, testing, and code quality practices Ability to collaborate effectively within cross-functional teams and follow established architectural and engineering standards Clear communicator with a growth mindset and interest in progressing toward broader technical ownership Bachelor's degree or equivalent practical experience in computer science, engineering, data science, or a related field Corporate Security Responsibility All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must: - Abide by Mastercard's security policies and practices; - Ensure the confidentiality and integrity of the information being accessed; - Report any suspected information security violation or breach, and - Complete all periodic mandatory security trainings in accordance with Mastercard's guidelines.
Production AI Engineer
Check Point Software TechnologiesAs the world’s leading vendor of Cyber Security, facing the most sophisticated threats and attacks, we’ve assembled a global team of the most driven, creative, and innovative people. At Check Point, our employees are redefining the security landscape by meeting our customers’ real-time needs and providing our cutting-edge technologies and services to an ever-growing customer base. Check Point Software Technologies has been honored by Time Magazine as one of the World’s Best Companies and Newsweek’s list of Americas Best Cybersecurity Companies. We've also earned a spot on the Forbes list of the World’s Best Places to Work for five consecutive years and recognized as one of the World’s Top Female-Friendly Companies. If you're passionate about making the world a safer place and want to be part of an award-winning company culture, we invite you to join us.
Role Description Check Point is seeking a promising and talented Production AI Engineer to join our Production and DevOps group. If you thrive in a fast-paced, dynamic environment, can handle multiple requests simultaneously, and enjoy working independently as part of a cutting-edge DevOps group, this is your opportunity to help make the world a safer place! - Act as a Production Engineer within a highly skilled team, responsible for large-scale operations from development to production. - Design, develop, and maintain the production platform, including operating systems, containers, cloud orchestration, and full end-to-end automation. - Implement tools and procedures for monitoring, deployment, and alerting across our SaaS multi-tenant product family. - Participate in the large-scale migration of a highly complex system into a secured, regulation-compliant environment. - Continuously improve our cloud infrastructure to ensure fault tolerance, scalability, and security. - Plan capacity, stabilize, and enhance the performance of application infrastructure with cost efficiency and scaling in mind. - Design and shape our monitoring and logging solutions. - Execute all tasks with top-notch cloud infrastructure security as a guiding principle. Qualifications - Hands-on mindset – we all write code daily! We all handle production issues. - 3+ years of relevant DevOps/Production experience building CI/CD pipelines for both development and production – must. - 2+ years of AWS Cloud experience working with distributed systems and microservices services – must. - Strong scripting skills, with fluency in Python – must. - Experience with AI tools and Agent development – must. - Experience with containers and orchestration tools (Docker, Kubernetes, or ECS) – must. - Experience with CI integration tools such as Jenkins. - Familiarity with AWS CloudFormation – an advantage. - Exposure to a wide range of open-source technologies (Open Telemetry, Nagios, Grafana, Prometheus, etc.). - Knowledge of best practices in security, performance, and monitoring. - Proven ability to research, evaluate, and implement new technologies, including running proof of concepts and cost analysis. - Must be eligible to work in the US without sponsorship from an employer now or in the future. Requirements - EOE M/F/Veterans/Disabled Benefits - Healthcare benefits - 401(k) plan and company match - Short-term and long-term disability coverage - Basic life insurance - Stock awards and an employee stock purchasing plan
• Work on the platform with a strong focus on MLOps/LLMOps, backend, and integrations. • Build and maintain automation pipelines for continuous deployment of ML and LLM models. • Monitor model performance, detect drift, and automate reprocessing/retraining. • Develop and integrate APIs for communication between AI models and internal systems using FastAPI. • Work with advanced orchestration in Kubernetes and Airflow to ensure scalability and resilience.
• Monitor and maintain the organization’s infrastructure, including servers, networks, storage systems, and applications. • Perform routine system checks and preventive maintenance to ensure optimal performance and uptime. • Respond to system alerts and incidents, diagnosing and resolving issues promptly to minimize downtime. • Provide technical support to resolve infrastructure-related issues, working closely with other technical teams. • Troubleshoot and resolve hardware, software, and network issues, escalating to higher-level support when necessary. • Maintain detailed documentation of issues, solutions, and processes to improve the team’s knowledge base. • Plan and execute system upgrades, patches, and configuration changes, ensuring minimal disruption to business operations. • Test and validate updates in development environments before deploying them to production. • Ensure that all systems comply with security standards and best practices. • Identify opportunities to automate routine tasks and processes, improving operational efficiency and reducing manual workload. • Implement scripts, automation tools, and AI skills to streamline system management and monitoring. • Continuously evaluate and optimize infrastructure performance, capacity, and resource utilization. • Support the development and execution of disaster recovery plans to ensure business continuity in case of system failures. • Manage backup and restore processes for critical systems and data, ensuring data integrity and availability. • Participate in regular disaster recovery testing and drills. • Plan and execute decommissioning of legacy infrastructure, coordinating Terraform state cleanup and DNS cutover. • Work closely with development, network, and security teams to ensure alignment and effective communication on infrastructure projects. • Provide input on infrastructure design and architecture to support new projects and initiatives. • Communicate effectively with non-technical stakeholders, providing updates on system status and issues.


