Modern Data Orchestration
Customer Reliability Engineer – Infrastructure
Location
California + 7 moreAll locations: California | Florida | New York | North Carolina | Ohio | Pennsylvania | Texas | Washington
Posted
3 days ago
Salary
$125K - $130K / year
Seniority
Senior
Job Description
Customer Reliability Engineer – Infrastructure
Astronomer
• Provide solutions to customers to make them successful using our products. • Troubleshoot customer environments and engage in active triaging with customers • Participate in on-call rotation for weekend coverage • Provide feedback to the product development teams on customer needs and pain points. • Build out our monitoring and alerting systems. • Build and maintain automation to ensure daily operational tasks are handled as efficiently as possible. • Help direct the architecture of the products and contribute where possible. • Own the customer experience, working directly with customers to prioritize and solve issues, meet SLAs, and provide “white glove” guidance on the path to production. • Participate remotely within a fully distributed team. • Enhance and enrich customer documentation • Work with the latest technology and multi-cloud implementations
Job Requirements
- 5 years of experience, preferably with large, complex cloud infrastructures operating at scale
- 3 years of experience with Kubernetes
- Experience managing a Production distributed system with at least one major cloud provider (one or all: AWS, GCP, Azure)
- Strong Linux experience
- Knowledge of how to operate and monitor issues for distributed systems
- Previous experience in handling customers issues (internal or external)
- Strong communication skills
- DevOps or CI/CD experience
- Python scripting
- Good troubleshooting Skills
Benefits
- equity component
- comprehensive benefits package
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Configure, deploy, and maintain security tools across cloud-native environments. • Integrate security tooling into existing software development and deployment workflows. • Partner with engineering teams to implement security best practices throughout the software development lifecycle. • Manage and optimize security controls within AWS cloud environments and Kubernetes clusters. • Maintain and improve application, infrastructure, and container security posture. • Implement automated security scanning, monitoring, alerting, and remediation processes. • Support compliance, vulnerability management, and incident response efforts. • Continuously evaluate and introduce new security technologies and practices to strengthen platform security.
Senior Electronic Component Reliability Engineer
ALTEN Technology USAWe help transform ideas into innovations with offices across the US, including Denver, CO; Troy, MI; and Greensboro, NC.
• Provide hands-on component reliability support across the full set of Electronic Control Units (ECUs) • Support directed component selection for upcoming designs, coordinating with the client's Component Engineering team and vendors to help meet vehicle mission and reliability targets. • Assist Component Reliability and ECU Reliability teams with Bill of Materials (BoM) reviews against vehicle life requirements and design guidelines, including coordinating vendor actions. • Maintain the list of major component risks and controllers as directed, with routine vendor follow-ups for updates and data. • Prepare and send vendor communications on reliability requirements, targets, and test plans, and manage follow-up responses. • Collect, organize, and format component reliability test data from vendors, including routine follow-ups for missing or updated information. • Execute root-cause coordination tasks, including meeting logistics, data collection, and regular Contract Manufacturer / vendor follow-ups to close action items. • Engage vendors directly in discussions on parts, specifications, and sourcing.
Cloud Operations Engineer
MongoDBMongoDB, originally called 10gen, is a software development company. Since 2007, MongoDB has created an open-source, document-oriented database to help clients
MongoDB Atlas is the premier multi-cloud database-as-a-service built and operated by the makers of MongoDB. The Cloud Operations Engineering team at MongoDB is a worldwide team responsible for the consistent operational success of every MongoDB Atlas customer. As a Cloud Operations Engineer, you will help ensure the success of our Atlas customers, whether they are early startups or large multinational companies, cloud-native or just getting started with a digital transformation to the cloud. You are excited about the core mission of MongoDB, and the opportunity to join the team responsible for operating Atlas, the fastest-growing multi-cloud database-as-a-service in the world. You are prepared to be one of the early members of a 24/7/365 global cloud operations team. Cloud Operations Engineers will be responsible for day-to-day duties such as creating and monitoring system’s alert dashboards, reviewing critical events and system logs, accessing customer instances that underpin their production databases and performing server administration duties including performance troubleshooting. Applicants must be critical thinkers who are quick to detect, resolve, or escalate issues that are sometimes broad in scope and difficult to trace. At MongoDB you will grow your career and skills, wear multiple hats, and be part of an operations team that works at the frontier of Cloud services and database systems. We are looking to speak to candidates who will be based remotely in Ireland. Due to the 24/7 nature of our support organization, certain events throughout the year will require volunteering for coverage outside one’s normal work days or work hours (i.e. regional offsites, regional holidays, etc). These are typically announced weeks in advance with a sign-up system that considers equitability. Responsibilities - Successfully coordinate and collaborate with a global team of Cloud Operations Engineers who are tasked with ensuring our uptime guarantees to our Atlas customer base - Help scale the worldwide Cloud Operations Engineering team with the strategic implementation and refinement of new processes and tools - Assist in scoping, designing and deploying systems that reduce Mean Time to Resolve for customer incidents - Monitor and detect emerging customer-facing incidents on the Atlas platform; assist in their proactive resolution - Automate routine monitoring and troubleshooting tasks - Diagnose live incidents, differentiate between platform issues versus usage issues, and take the next steps toward resolution - Assist in performing root cause analysis after incident recovered; identifying any breakdowns in processes or workflows that contributed to the event and what changes need to be made to prevent similar events - Contribute to documentation of corner case scenarios, troubleshooting workflows and SOPs. - Work alongside our product management, cloud engineering and support organizations by identifying areas for improvement in the management applications powering the Atlas infrastructure - Inform executive leadership and escalation management personnel of major outages - Coordinate and participate in a weekly on-call rotation, where you will handle short term customer incidents (proactively from automated monitoring or through reactive alerts via our Technical Services team) Requirements - Experience with being an on call DevOps, SRE, or Cloud Operations engineer (at least 2 years) - Expertise with Linux system administration, configuration, troubleshooting - Experience in monitoring, system performance data collection and analysis, and reporting - Knowledge of database operations and concepts - Expertise with networking technologies like DNS, TCP/IP, etc. - Familiarity with Amazon Web Services and other Cloud infrastructure platforms (e.g. GCP, Azure) - Knowledgeable about a wide range of web and internet technologies - Capability to write small programs/scripts to solve both short-term systems problems - A CS/CE degree or equivalent experience - At least 1 of the following programming languages: Java, Go, Javascript - A keen interest in learning new things Nice To Have - MongoDB - Splunk - Kubernetes Benefits include - Competitive salary, equity, pension and health insurance - Regular performance, compensation and development reviews - 20 weeks Maternity & Paternity leave to spend time with new arrivals About MongoDBMongoDB is built for change, empowering our customers and our people to innovate at the speed of the market. We have redefined the data platform for the AI era, enabling builders to create, transform, and disrupt industries with software. MongoDB’s unified data platform, the most widely available, globally distributed data platform on the market, helps organizations modernize legacy workloads, embrace innovation, and unleash AI. Our cloud-native platform, MongoDB Atlas, is the only globally distributed, multi-cloud data platform and is available across AWS, Google Cloud, and Microsoft Azure. With offices worldwide and over 67,000 customers, including 75% of the Fortune 100 and AI-native startups, relying on MongoDB for their most important applications, we’re powering the next era of software. Our compass at MongoDB is our Leadership Commitment, guiding how and why we make decisions, show up for each other, and win. It’s what makes us MongoDB. To drive the personal growth and business impact of our employees, we’re committed to developing a supportive and enriching culture for everyone. From employee affinity groups, to fertility assistance and a generous parental leave policy, we value our employees’ wellbeing and want to support them along every step of their professional and personal journeys. Learn more about what it’s like to work at MongoDB, and help us make an impact on the world! MongoDB is committed to providing any necessary accommodations for individuals with disabilities within our application and interview process. To request an accommodation due to a disability, please inform your recruiter. MongoDB is an equal opportunities employer. Req ID 4263312827
Role Description To further strengthen our Corporate Unit we are looking for a: Senior Linux DevOps/Systems Engineer - Administer and optimize Enterprise Linux systems across global data centers, ensuring uptime, performance, and adherence to security best practices (e.g., SELinux, firewallD, OpenSCAP). - Deploy, configure, and maintain Linux-based applications such as HAProxy, NGINX, and other critical services, focusing on high-availability, scalability, and performance optimization. - Design, implement, and operate monitoring and logging solutions using Prometheus, Sensu, OpsGenie, Grafana, and Loki to ensure visibility, proactive alerting, and rapid incident resolution. - Drive automation with Infrastructure as Code (IaC) using Ansible to enable repeatable, reliable, and scalable deployments and configurations. - Manage Identity Management solutions like FreeIPA and integrate them securely into the infrastructure, supporting authentication and access control across services. - Assist in operating containerized workloads (Docker/Podman) and Kubernetes environments as part of a broader infrastructure strategy. Qualifications - Solid, hands-on expertise in Linux system administration (RHEL-based and Debian-based), including troubleshooting, tuning, and lifecycle management — this is the foundation of the role. - Proven experience with Linux application services such as HAProxy and NGINX, including high-availability design and performance tuning. - Practical experience with centralized monitoring and logging tools (Prometheus, Grafana Loki, Sensu, OpsGenie) in production environments. - Strong skills in automation using Ansible and familiarity with Infrastructure as Code principles. - Knowledge of Linux security and compliance frameworks (SELinux, firewallD, OpenSCAP) and distributed failover/high-availability strategies. - Fluent English communication skills (written & spoken) with the ability to document clearly, collaborate effectively, and solve complex problems. Benefits DO YOU FEEL ADDRESSED? Then we look forward to receiving your detailed application stating your earliest possible starting date and your salary expectations.



