Job Closed

This listing is no longer active.

Domino Data Lab logo
Domino Data Lab

The Enterprise MLOps platform powering over 20% of the Fortune 100

Staff Platform Reliability Engineer

Platform EngineerPlatform EngineerFull TimeRemoteSeniorTeam 201-500Since 2013H1B SponsorCompany SiteLinkedIn

Location

United States

Posted

115 days ago

Salary

$185K - $230K / year

Seniority

Senior

Job Description

Staff Platform Reliability Engineer

Domino Data Lab

Who we are At Domino, we build software that helps the largest, AI-driven organizations build and operate advanced data science and AI solutions at scale. Our platform integrates a streamlined model development environment, MLOps capabilities, and novel features for collaboration, reuse, and reproducibility — all of which make data science teams more productive, reduce time to value, and ensure compliance. Our customers — like Johnson & Johnson, GSK, Bristol Myers, UBS, FINRA and the US Navy — are using our software to solve some of the most important challenges in the world, such as developing new medicines, securing our financial markets, or protecting our country. Backed by Sequoia Capital, Coatue Management, NVIDIA, Snowflake and other leading investors, we have been in business for a decade but are still a small team operating with the spirit of a startup. Especially in the world of AI today, we believe that the future is still being invented — and we want to be the ones building it. For more information, visit www.domino.ai What we are building The Automation Team at Domino acts as a force multiplier for engineering, building the tools and systems that enable teams to ship code confidently and consistently. A core part of this mission is Tempest, an in-house platform that orchestrates realistic, long-duration workloads against live Kubernetes clusters and validates the results against real observability data. Today, when scale testing surfaces a bottleneck, a resource misconfiguration, or a regression in system behavior, the team can identify and report the issue — but we need someone who can take the next step: profiling services, tracing root causes through Prometheus and New Relic data, and partnering with platform engineers to drive durable fixes. Focused on iteration and continuous improvement, the team looks for targeted enhancements that create outsized impact, and this role will close the gap between detection and resolution at the infrastructure level. What your impact will be In your first year, you will: - Serve as the technical owner of Tempest, Domino's scale and reliability platform, ensuring it remains reliable, extensible, and aligned with evolving infrastructure needs - Diagnose and drive resolution of performance bottlenecks and resource misconfigurations surfaced by scale testing — working directly with platform and infrastructure teams to ship fixes, not just file tickets - Deliver accurate, data-driven sizing recommendations for customer-facing documentation based on rigorous empirical testing across deployment sizes - Strengthen observability across scale testing by improving Prometheus and New Relic instrumentation, making it faster to pinpoint root causes during and after multi-day load runs - Establish and operationalize scale testing on cloud platforms, ensuring appropriate sizing and configuration guidance for this increasingly divergent product line - Partner with platform teams to enable effective scale and reliability testing across additional cloud providers, helping position Domino for future multi-cloud success - Increase the efficiency and leverage of a small team by building infrastructure automation that scales operationally as the product and customer base grow What we look for in this role - Background in SRE, platform engineering, or infrastructure with hands-on experience operating and troubleshooting distributed systems in production Kubernetes environments - Strong proficiency in Python and comfort working in a large, modular codebase that spans orchestration, infrastructure automation, and systems integration - Experience with observability stacks (Prometheus, Grafana, New Relic, or similar) — writing queries, building dashboards, and using metrics to diagnose performance and reliability issues at the systems level - Demonstrated ability to go beyond detection to resolution: profiling services, identifying resource bottlenecks, and working with engineering teams to ship durable fixes - Familiarity with performance and load testing methodologies (e.g., Locust, k6, or similar) as part of a broader infrastructure or reliability practice - Clear ownership mindset — self-directed, accountable, and able to communicate priorities and status effectively in a remote, async environment What we value - We value a growth mindset. High-performing creative individuals who dig into problems and see the opportunities for success - We believe in individuals who seek truth and speak the truth and can be their whole selves at work - We value all of you that believe improving is always possible At Domino Everything is a work in progress – we can do better at everything - We emphasize an environment of teaching and learning to equip employees with the tools needed to be successful in their function and the company - We strongly believe in the value of growing a diverse team and encourage people of all backgrounds, genders, ethnicities, abilities, and sexual orientations to apply #LI-Remote The annual US base salary range for this role is listed below. For sales roles, the range provided is the role's On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. This salary range will be narrowed during the interview process based on a number of factors, including the candidate's experience, qualifications, and location. Additional benefits for this role may include: equity, company bonus or sales commissions/bonuses; 401(k) plan; medical, dental, and vision benefits; and wellness stipends. Compensation Range $185,000—$230,000 USD

Benefits

  • 401(K), Childcare benefits, Commuter benefits, Company equity, Company-sponsored outings, Company sponsored family events, Continuing education stipend, Customized development tracks, Dental insurance, Disability insurance, Documented equal pay policy, Volunteer in local community, Family medical leave, Fitness stipend, Flexible Spending Account (FSA), Flexible work schedule, Generous parental leave, Company-sponsored happy hours, Health insurance, Job training & conferences, Open door policy, Life insurance, Mean gender pay gap below 10%, Online course subscriptions available, Onsite gym, Open office floor plan, Paid holidays, Paid industry certifications, Pair programming, Paid sick days, Partners with nonprofits, Pet friendly, Pet insurance, Promote from within, Recreational clubs, Lunch and learns, Remote work program, Free snacks and drinks, Team based strategic planning, OKR operational model, Continuing education available during work hours, Mandated unconscious bias training, Unlimited vacation policy, Vision insurance, Wellness programs, Some meals provided, Mental health benefits, Home-office stipend for remote employees, Diversity employee resource groups, Hiring practices that promote diversity

Related Categories

Related Job Pages

More Platform Engineer Jobs

General Motors logo

Staff AI/ML Engineer - CI Platform

General Motors

General Motors (GM), founded in 1908 by William "Billy" Durant in Flint, Michigan, began with the Buick Motor Company and later acquired brands like Oldsmobile

Platform Engineer115 days ago

Description This role is categorized as hybrid. This means the successful candidate is expected to report to the office in [Austin, Detroit, Warren, Milford, Mountain View or Sunnyvale] a minimum of 3 days per week. This position can be considered Remote if the successful candidate is in the Seattle, Washington area. About Us The AI Cloud and Developer Infrastructure organization is responsible for delivering and maintaining the tools and services engineers here at GM use every day to do their best work and drive our cars forward. Tools and services we work on enhancing the entire development process of engineers at GM - how/where code is checked out, modified, compiled, tested, merged, and eventually deployed. Our goal is to ensure our AV engineers and others here have world class tools and a seamless development experience so that they can focus on the problems that matter most in their domain. The Role We are looking for a Staff Engineer with an extensive engineering background, experience using a variety of developer tools and technologies, and who is passionate about developer productivity. As a leader on this team, we are looking for someone who cares deeply about the technical development of other engineers on the team and can effectively balance the needs and priorities of the business, our users, and the growth of our engineers. The way this engineer will deliver impact may vary depending on the situation, but they will be expected to be able to identify how they can best have impact with minimal guidance. The Team The Continuous Integration (CI) Platform team owns our CI infrastructure along with the tools that improve the CI customer experience. We are part of the AI and Cloud Developer Productivity Organization and own the building blocks to produce high quality CI experiences and workflows. This includes services such as the Remote Build Execution (RBE) service which allows GM AI to reliably test at scale and the build artifact storage which is a FUSE based network file system used in conjunction with RBE. Our goal is to accelerate AV development by supporting developer workflows for building, testing, and releasing software. What You'll Do (Responsibilities) - Identify engineering pain points and propose/design/implement solutions that are reliable, scalable, and maintainable - Influence the team's technical roadmap - Evaluate new tools and technologies through PoCs - Ship improvements to our AV development toolchains and services which have a measurable and direct impact on engineering productivity and our core company metrics - Drive software engineering best practices within your team, and create tooling which encourages these - Help steer the engineering culture on the team - Guide the team to find the right balance between delivering impact and addressing technical debt - Mentor and grow engineers on the team - Set the example for high levels of accountability - Execute and deliver impact both individually and through the team - Set strong boundaries when selecting external requests and pushing back on requests that do not align with our team vision Minimum Qualifications (Must-Have) - 7+ years of experience designing, building and operating production systems at scale in the cloud - Bachelors Degree in Computer Science or related field or equivalent work experience - Experience designing highly scalable, reliable, and maintainable services - Experience writing in Go, Python, or other languages at production scale - Understanding of Unix/Linux, SSH, and networking fundamentals - Attention to detail, and a desire to improve processes and systems around you - Ability to lead and influence others, both internal and external to the team - Ability to research, document, communicate, and defend proposals, and provide and take critical feedback - Ability to effectively make trade-offs and communicate the reasoning - Ability to manage competing priorities, focus on shipping, and work effectively under pressure - Passion for mentoring and growing junior engineers - Passion for self-driving technology and its potential impact on the world Preferred Qualifications (Nice-to-Have) - Experience working with GCP - Experience working with Docker and Kubernetes - Experience owning or contributing to Open-Source projects Compensation: The compensation information is a good faith estimate only. It is based on what a successful applicant might be paid in accordance with applicable state laws. The compensation may not be representative for positions located outside of New York, Colorado, California, or Washington. - The salary range for this role: is $170,000 to $240,000. The actual base salary a successful candidate will be offered within this range will vary based on factors relevant to the position. - Bonus Potential: An incentive pay program offers payouts based on company performance, job level, and individual performance. - Benefits: GM offers a variety of health and wellbeing benefit programs. Benefit options include medical, dental, vision, Health Savings Account, Flexible Spending Accounts, retirement savings plan, sickness and accident benefits, life insurance, paid vacation & holidays, tuition assistance programs, employee assistance program, GM vehicle discounts and more. This job may be eligible for relocation benefits. #GM-AV-1 This role is categorized as hybrid. This means the selected candidate is expected to report to a specific location at least 3 times a week {or other frequency dictated by their manager}. The selected candidate will be required to travel <25% for this role. This job may be eligible for relocation benefits. About GM Our vision is a world with Zero Crashes, Zero Emissions and Zero Congestion and we embrace the responsibility to lead the change that will make our world better, safer and more equitable for all. Why Join Us We believe we all must make a choice every day - individually and collectively - to drive meaningful change through our words, our deeds and our culture. Every day, we want every employee to feel they belong to one General Motors team. Total Rewards | Benefits Overview From day one, we're looking out for your well-being-at work and at home-so you can focus on realizing your ambitions. Learn how GM supports a rewarding career that rewards you personally by visiting Total Rewards resources. Non-Discrimination and Equal Employment Opportunities (U.S.) General Motors is committed to being a workplace that is not only free of unlawful discrimination, but one that genuinely fosters inclusion and belonging. We strongly believe that providing an inclusive workplace creates an environment in which our employees can thrive and develop better products for our customers. All employment decisions are made on a non-discriminatory basis without regard to sex, race, color, national origin, citizenship status, religion, age, disability, pregnancy or maternity status, sexual orientation, gender identity, status as a veteran or protected veteran, or any other similarly protected status in accordance with federal, state and local laws. We encourage interested candidates to review the key responsibilities and qualifications for each role and apply for any positions that match their skills and capabilities. Applicants in the recruitment process may be required, where applicable, to successfully complete a role-related assessment(s) and/or a pre-employment screening prior to beginning employment. To learn more, visit How we Hire. Accommodations General Motors offers opportunities to all job seekers including individuals with disabilities. If you need a reasonable accommodation to assist with your job search or application for employment, email us [email protected] or call us at 1-800-865-7580. In your email, please include a description of the specific accommodation you are requesting as well as the job title and requisition number of the position for which you are applying.

Texas + 2 moreAll locations: Texas | Michigan | California
$170K - $240K / year
Teladoc Health logo

Senior Platform Engineer – Azure, Terraform, Snowflake

Teladoc Health

Teladoc Health is a public company and a global, online care leader working to transform how people access healthcare by helping individuals and organizations r

Platform Engineer115 days ago

• Own reliability, automation, and infrastructure-as-code for our modern Data & AI platform. • Ensure our Azure-based data ecosystem is reliable, scalable, and efficient. • Build Terraform-first infrastructure, improve developer experience, and support a healthcare environment. • Implement monitoring, alerting, and SLO/SLIs for data pipelines and platform components. • Lead incident response, root cause analysis, and postmortems. • Create automation, runbooks, and self-healing capabilities to reduce MTTR. • Design secure connectivity patterns between Azure and AWS vendor systems. • Troubleshoot networking, VPN, private endpoints, DNS, and MFT integrations. • Optimize Snowflake compute usage and Airflow/dbt performance.

United States
$155K - $175K / year
Job Closed
Empower logo

Senior Data Platform Engineer

Empower

We are an equal opportunity employer with a commitment to diversity. All individuals, regardless of personal characteristics, are encouraged to apply. All qualified applicants will receive consideration for employment without regard to age, race, color, national origin, ancestry, sex, sexual orientation, gender, gender identity, gender expression, marital status, pregnancy, religion, physical or mental disability, military or veteran status, genetic information, or any other status protected by applicable state or local law.

Platform Engineer115 days ago
Full TimeRemoteTeam 10,001+H1B Sponsor

• Design and implement data platform solutions including data streaming, CDC, data warehouse, data lake, and ETL/ELT pipelines using modern cloud-based technologies. • Contribute to the development and adoption of best practices for data platform configuration management, observability, testing, and operational readiness. • Build and support end-to-end data integration pipelines from source systems through Kafka and/or Striim to analytical and operational targets. • Participate in engineering design sessions and code reviews for complex data solutions across multiple systems and platforms. • Troubleshoot and resolve production issues related to data platform performance, scalability, and reliability. • Collaborate with architects, platform engineering, security, and governance teams to align solutions with enterprise standards and data strategy. • Support implementation of operational resilience practices, including monitoring, alerting, and data recovery processes. • Evaluate and help adopt new technologies and tools that improve data platform capabilities. • Contribute to documentation, design artifacts, and engineering standards. • Mentor junior and mid-level engineers and support team development.

United States
$105.7K - $149.3K / year
Job Closed
Full TimeRemoteTeam 201-500Since 2014H1B Sponsor

Join phData, a dynamic and innovative leader in the modern data stack. We partner with major cloud data platforms like Snowflake, AWS, Azure, GCP, Fivetran, Pinecone, Glean, and dbt to deliver cutting-edge services and solutions. We're committed to helping global enterprises overcome their toughest data challenges. phData is a remote-first global company with employees based in the United States, Latin America, and India. We celebrate the culture of each of our team members and foster a community of technological curiosity, ownership, and trust. Even though we're growing extremely fast, we maintain a casual, exciting work environment. We hire top performers and allow you the autonomy to deliver results. - 6x Snowflake Partner of the Year (2020, 2021, 2022, 2023, 2024, 2025) - Fivetran, dbt, Atlation, and AWS Partner of the Year - #1 Partner in Snowflake Advanced Certifications - 600+ Expert Cloud Certifications (Sigma, AWS, Azure, Dataiku, etc) Recognized as an award-winning workplace in the US, India, and LATAM Senior AI Automation Engineer - Internal Platform At phData, the Platform team builds and operates our internal Intelligence Platform, powering our Operations, Sales, Delivery, and Finance teams with data, analytics, and AI‑driven insights. We provide the core data and technology foundation that helps phData run efficiently and make better decisions every day. We are seeking an AI Automation Engineer to join our Platform team. This role will serve as a hands‑on technical partner to business groups across the organization. The ideal candidate brings demonstrated experience in both artificial intelligence and process automation, with a proven ability to assess feasibility, design solutions, and deliver measurable outcomes that create efficiencies and solve real business problems. As an AI Automation Engineer, you will work directly with business stakeholders to understand their challenges, evaluate whether AI, automation, or a combination of both is the right approach, and then design and deliver end‑to‑end solutions. This is a growth‑oriented role for those passionate about using AI and automation to drive impact in how we run phData on AI. What You’ll Do: - Business Engagement & Feasibility Assessment: Partner with business groups to identify AI/automation opportunities, assess feasibility (viability, data, integration, ROI, readiness), translate problems into technical plans, communicate recommendations, and help prioritize demand across the portfolio. - AI Solution Development: Design, build, and refine AI‑driven applications and agents (RAG, prompt engineering, agent architectures) for internal and occasional client use on platforms like Glean, Microsoft Copilot, and Snowflake Intelligence. - Automation & Efficiency: Implement RPA, scripting, and low‑code/no‑code workflows to automate repetitive processes, integrate systems and data sources, and operate/optimize automations for reliable, end‑to‑end execution. - Development & Code Quality: Write production‑grade code for AI and automation solutions with strong documentation, testing, and SDLC discipline, using modern engineering practices (Git, code review, CI/CD) to ensure quality and maintainability. - Collaborative Development & Cross‑Functional Engagement: Work with cross‑functional teams to deliver solutions, clearly communicate value to technical and non‑technical stakeholders, and align work with priorities through agile planning. - Strategic Support & Professional Growth: Share best practices, track and experiment with emerging AI/automation tools, inform leadership through prototypes and production learnings, and help select and operate secure, governed cloud AI services. Required Experience: - 4‑year Bachelor’s degree in Computer Science, Engineering, Data Science, or a related technical field. - 5+ years of professional experience developing and deploying automation/RPA solutions and AI/ML solutions in a business environment. - Experience with enterprise integration patterns, APIs, and connecting disparate business systems. - Demonstrated experience across both AI/ML and process automation (e.g., RPA, Power Automate, scripting, workflow tools). - Proven ability to work directly with business stakeholders to assess feasibility, gather requirements, and translate business problems into technical solutions. - Strong proficiency in software development (Python preferred), with hands‑on experience building production‑grade applications and workflows. - Familiarity with professional software development workflows: Git‑based source control, branching strategies, code review, and effective team collaboration on shared codebases. - Exposure to agile development methodologies, CI/CD practices, and collaborative development environments. - Basic experience deploying cloud services and APIs for AI workloads (AWS, Azure, or Google Cloud). - Strong analytical and problem‑solving capabilities with a demonstrated ability to assess business processes and identify automation opportunities. - Excellent communication skills for explaining technical information to both technical and non‑technical audiences. - Ability to adapt to evolving technology landscapes and work collaboratively in multidisciplinary teams. - Experience with generative AI, large language models (LLMs), prompt engineering, and AI agent frameworks/platforms. Prefer Any of the Following: - Familiarity with data engineering concepts (ETL/ELT, data pipelines, data warehousing). - Familiarity with the Snowflake Data Platform and cloud data ecosystems. - Experience with business intelligence tools like Sigma Computing. - Experience in the data and AI professional services industry. - Experience conducting structured opportunity assessments, business case development, or cost‑benefit analysis for technology initiatives. - Certifications in AI/ML, RPA, or cloud platforms (AWS, Azure, GCP). - Knowledge of data governance, responsible AI principles, and regulatory compliance for sensitive or regulated industries. Key Attributes for Success: - Consultative Mindset: Naturally engages with business teams to understand the “why” behind requests, asks the right questions, and proposes solutions that address root causes. - Dual‑Skilled: Equally comfortable building an AI agent and designing an automated workflow; understands when to apply AI, when to apply automation, and when to combine both. - Ownership & Follow‑Through: Takes end‑to‑end responsibility for solutions from initial assessment through deployment and ongoing optimization. - Bias Toward Action: Moves things forward with urgency and purpose; doesn’t let ambiguity stall progress and actively seeks opportunities to improve how work gets done. - Collaborative & Influential: Builds trust across teams and stakeholders; can navigate complex organizational dynamics with professionalism and clarity. - Curious & Growth‑Oriented: Stays current with rapidly evolving AI and automation trends; brings new ideas and approaches to the team. - Resilient & Adaptable: Thrives in a fast‑paced, evolving environment; stays composed under pressure and pivots when circumstances change. Why phData: - Work on a high‑impact internal AI and automation platform that directly shapes how phData operates, sells, and delivers for its customers. - Help define how we run phData on data and AI, collaborating closely with stakeholders across Operations, Sales, Delivery, Finance, and IT. - Build on a modern cloud data and AI stack (Snowflake, AWS, Sigma Computing, Glean, GitHub Copilot, and more) with strong support for experimentation and improvement. - Collaborate with experienced data, analytics, and AI practitioners, and play a key role in expanding our internal AI and automation capabilities. - Enjoy a remote‑friendly culture with a distributed team across the globe. Location & Time Zone Expectations This role is based in LATAM (with Brazil as the preferred location) and operates primarily in Brazil time zones with overlap to US time zones. - We are a remote-first company, and you should be comfortable working with a distributed global team. - Some flexibility may be required to collaborate across time zones with colleagues and clients. - Client needs may occasionally require flexibility in working hours to support key milestones or workshops. Benefits at phData - Remote-First Work Environment - Casual, award-winning small-business work environment - Collaborative culture that prizes autonomy, creativity, and transparency - Competitive comp, excellent benefits, generous PTO plan plus 10 Holidays (and other cool perks) - Accelerated learning and professional development through advanced training and certifications phData celebrates diversity and is committed to creating an inclusive environment for all employees. Our approach helps us to build a winning team that represents a variety of backgrounds, perspectives, and abilities. So, regardless of how your diversity expresses itself, you can find a home here at phData. We are proud to be an equal opportunity employer. We prohibit discrimination and harassment of any kind based on race, color, religion, national origin, sex (including pregnancy), sexual orientation, gender identity, gender expression, age, veteran status, genetic information, disability, or other applicable legally protected characteristics. If you would like to request an accommodation due to a disability, please contact us at People Operations.

Finland