Job Closed
This listing is no longer active.
Improving people’s lives by harnessing the healthcare data explosion through an intelligent data integration platform.
Principal Sustaining Engineer – Forward Deployed
Location
United States
Posted
155 days ago
Salary
0
Seniority
Lead
Job Description
Principal Sustaining Engineer – Forward Deployed
Abacus Insights
• Act as a senior technical escalation point during production incidents • Lead real-time incident triage, mitigation, and recovery efforts • Drive root cause analysis (RCA) with a focus on systemic, long-term fixes • Identify recurring failure patterns and push for architectural or operational improvements • Partner with Customer Success and Engineering to manage customer impact during incidents • Own post-launch reliability, stability, and operational quality of core systems • Investigate and resolve complex field issues and production defects • Ensure fixes developed during incidents or customer escalations are up streamed into the core product • Improve operational readiness of services through better runbooks, monitoring, and alerting • Reduce operational toil by converting repeated manual work into automation • Engage directly with strategic customers to solve real-world, production-grade technical challenges • Support complex deployments, integrations, and escalations in customer environments • Act as a trusted technical partner to customers during high-impact issues • Translate customer learnings into concrete product, platform, and operational improvements • Contribute to reusable tools, playbooks, and best practices that accelerate future deployments • Serve as a subject matter expert for AWS-hosted production systems • Troubleshoot and resolve issues across: • AWS compute, storage, networking, IAM, and security • Databricks jobs, clusters, and Spark-based data pipelines • Debug performance degradation, scalability issues, job failures, and data correctness problems • Partner with platform and data teams to harden systems for reliability, scale, and operability • Write production-quality code to: • Automate operational workflows • Improve reliability and observability • Eliminate manual intervention and reduce incident frequency • Contribute primarily in Python, with exposure to JVM-based systems as needed • Review code with a strong emphasis on operability, resiliency, and maintainability • Provide technical leadership without formal authority, influencing design and operational decisions • Mentor engineers through pairing, reviews, and incident leadership • Collaborate closely with Product, Engineering, Data, and Customer teams • Operate effectively in high-pressure, ambiguous environments, especially during customer-impacting incidents
Job Requirements
- 10+ years of experience in software engineering, SRE, sustaining engineering, or production operations
- Deep hands-on experience operating production systems in AWS
- Strong experience troubleshooting Databricks and large-scale data platforms
- Proficiency in Python and experience building production services or tooling
- Strong understanding of:
- Distributed systems
- Incident management and RCA practices
- Monitoring, alerting, and observability
- CI/CD Pipelines that leverage Infrastructure as Code.
- Proven ability to own problems end-to-end, from detection to permanent resolution
- Excellent communication skills, especially during incidents and customer escalations
- Ability to work backward from customer impact to root cause across systems and codebases, delivering fixes in environments with minimal documentation.
- Strong instinct for operational risk, with the ability to proactively identify failure modes and harden systems before they impact customers.
Benefits
- Unlimited paid time off – recharge when you need it
- Work from anywhere – flexibility to fit your life
- Comprehensive health coverage – multiple plan options to choose from
- Equity for every employee – share in our success
- Growth-focused environment – your development matters here
- Home office setup allowance – one-time support to get you started
- Monthly cell phone allowance – stay connected with ease
Related Guides
Related Categories
Related Job Pages
More Engineer Jobs
• Data gathering, validation and analysis. • Assist with the creation of MOPs (Method of Procedures), SOPs (Standard Operating Procedures) and EOPs (Emergency Operating Procedures). • Replicate MOPs, SOPs, EOPs across the platform where similar equipment exists. • Work with existing Facility Engineers to update equipment records. • Collaborate with Compliance team members who will provide direction, feedback, and mentorship throughout the internship.
• Lead the development of scalable, robust analytical models using DBT. • Ensure the implementation of modeling standards (dimensional and/or relational) and analytics engineering best practices. • Create, test and document transformations, ensuring data quality, reliability, traceability and governance. • Develop and maintain complex data pipelines using Airflow (or equivalent orchestration tools), applying versioning, modularity, observability and DataOps best practices. • Integrate different data sources (structured and unstructured), ensuring efficient, resilient and monitorable processes. • Optimize queries, structures and costs in cloud Data Warehouses (such as BigQuery, Redshift, Snowflake or similar), applying advanced performance tuning techniques. • Ensure governance, security, access control (IAM), partitioning, clustering and dataset versioning. • Collaborate with Data Analysts and Data Scientists to enable the consumption of reliable, well-documented and semantically organized datasets. • Evangelize technical standards, transformation best practices and governance across the organization. • Create and support dashboards and analytical products using BI tools (Looker, Power BI, Tableau or similar). • Work with Engineering, Product, Business and Data Governance teams to ensure data integration, consistency and quality. • Participate in defining and evolving modern cloud data architecture (GCP or AWS), proposing continuous improvements.
• Provide on-site support during implantation and follow-up procedures involving Class III medical devices. • Troubleshoot technical issues during procedures and provide immediate, clinically sound recommendations. • Act as the subject matter expert (SME) on device function and clinical applications. • Educate physicians, nurses, and allied health professionals on device functionality, therapy benefits, and procedural workflows. • Conduct in-service trainings, workshops, and ongoing clinical education for new and existing users. • Support collection of clinical data and user feedback to support regulatory submissions, product development, and continuous improvement efforts. • Document field interactions in compliance with quality and regulatory standards. • Work closely with R&D, Regulatory, Safety, and Commercial teams to relay device deficiencies, clinical insights and support product refinements. • Assist with investigational device trials, product evaluations, and site initiation visits (as applicable).
Lead Tools Engineer – Build, Infrastructure
XsollaXsolla's video game business engine helps game developers and publishers operate more efficiently and sell more games.
• Lead the design and development of internal tools that support game development, build pipelines, and developer workflows • Design, implement, and maintain build systems and CI/CD pipelines as first-class tools used daily by the engineering team • Develop robust, maintainable tools and automation using C#, Python, and Bash • Gather requirements from engineers and other disciplines, translate them into practical tooling solutions, and iterate based on feedback • Improve and optimize existing tools, build processes, and pipelines for reliability, performance, and usability • Work with command-line–driven build systems, including Unity batch mode, Xcode CLI, and similar tooling • Design and maintain cloud-backed tooling and infrastructure using AWS CDK and CloudFormation where appropriate • Take ownership of critical shared tools and systems, ensuring they remain stable, well-documented, and easy to evolve • Provide technical leadership and mentorship to other engineers, particularly around tooling, build systems, and CI/CD best practices. • Identify systemic workflow or productivity issues and propose well-thought-out tooling solutions • Participate in and help guide code reviews, architectural discussions, and engineering best practices • Collaborate with all departments to ensure our tools and systems make teams efficient and our games great




