Job Closed
This listing is no longer active.
Infrastructure for an Evolving World
Lead DevOps Engineer
Location
New York
Posted
141 days ago
Salary
$160K - $180K / year
Seniority
Senior
Job Description
Lead DevOps Engineer
Xpansiv
• Deploy and maintain critical applications on cloud-native microservices architecture • Implement automation, effective monitoring, and infrastructure-as-code • Deploy and maintain CI/CD pipelines across multiple environments • Support and work alongside a cross-functional engineering team on the latest technologies • Iterate on best practices to increase the quality & velocity of deployments • Have on call responsibilities in rotation with the engineering team • Increase the sophistication of our alerting and escalation mechanisms • Help increase system performance with a focus on high availability and scalability • Propose, scope, design, and implement various infrastructure architectures • Develop and maintain solutions for operational administration, system/data backup, disaster recovery, and security/performance monitoring • Continuously evaluate existing systems with industry standards, and make recommendations for improvement • Perform root cause analysis for production errors • Continue to keep the lights on (day-to-day administration)
Job Requirements
- BSc in Computer Science, Engineering or relevant field
- 12+ years of professional experience as a DevOps/System Engineer
- Experience maintaining and deploying highly-available, fault-tolerant systems at scale
- A drive towards automating repetitive tasks (e.g. scripting via Terraform, Powershell, Bash, Python, Ruby, etc.)
- Practical experience with Docker containerization and clustering (Kubernetes/ECS)
- Expertise with Azure, AWS or GCP
- Version control system experience (e.g. Git)
- Working knowledge of databases and SQL
- Experience implementing CI/CD (e.g. Azure DevOps, Jenkins)
- Operational (e.g. HA/Backups)
- Experience with configuration management tools (e.g. Powershell with DSC, Ansible, Chef)
- Experience with infrastructure-as-code
- Experience in managing on-premise Active Directory and Domain Controller
- Experience in managing on-premise infrastructure
- Effective communication skills
- Professional experience and a high-level understanding of working with various operating systems and their implications
- Problem-solving attitude
- Collaborative team spirit
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Provide technical and line-management leadership to your development team. • Take responsibility for the successful delivery of projects. • Identify and resolve blockers before they become issues. • Ensure best practices in DevOps, software development and agile methodologies are upheld within the team. • Work directly with clients, translating requirements into technical briefs. • Shape and define architectural decisions ensuring scalability, security, and maintainability. • Provide updates to client and Nearform leadership to ensure clear understanding of project status and drive good decision-making.
DevOps Engineer
Bluelight ConsultingBluelight is a leading software consultancy dedicated to designing and developing innovative technology that enhances users' lives. With a steadfast commitment to delivering exceptional service to our clients, Bluelight excels in its focus on quality and customer satisfaction. Our mission is not only to create cutting-edge applications but also to foster a collaborative and enriching work environment where each team member can grow and thrive. With a presence across the United States and Central/South America, Bluelight is in an exciting phase of expansion, continually seeking exceptional talent to join its dynamic and diverse community.
This description is a summary of our understanding of the job description. Click on 'Apply' button to find out more. Role Description We are looking for a skilled individual to join our rapidly growing team at Bluelight Consulting. This position is ideal for someone who thrives in a fast-paced, dynamic environment where everyone's opinions and efforts are valued and appreciated. You will have the opportunity to contribute to challenging and meaningful projects, developing high-quality applications that stand out in the market. We value continuous learning, personal growth, and hard work, offering a collaborative environment that promotes professional development. If you are passionate about software development and eager to be part of a growing software consultancy, we invite you to apply and join us on this exciting journey. Qualifications - Cloud Engineering (cloud computing) experience with AWS, GCP, and/or Azure to include load balancing - Infrastructure as a code (Terraform / Pulumi / Cloudformation) - Designed and maintained CI/CD process and tools (CircleCI, GitLab, Jenkins) - In-depth experience with the orchestration tools (Kubernetes) - In-depth experience with the config management tools (Helm, Ansible, Chef Puppet) - Testing, code review, good communication skills Benefits - Competitive salary and bonuses, including performance-based salary increases - Generous paid-time-off policy - Technology / Office stipend - Health Coverage - Flexible working hours - Work remotely - Continuing education, training, conferences - Company-sponsored coursework, exams, and certifications
• Design and evolve reliability architecture for distributed and cloud-hosted systems. • Define and implement SRE best practices, including SLIs, SLOs, error budgets, and capacity planning. • Partner with platform and application teams to design systems for reliability, scalability, and operability. • Identify and mitigate systemic reliability risks across infrastructure and services. • Lead incident response processes including on-call rotations, escalation, and post-incident reviews. • Conduct root cause analysis for complex production incidents and drive long-term improvements. • Improve operational readiness through runbooks, automation, and resilience testing. • Reduce operational toil through tooling, automation, and process improvements. • Design and maintain observability systems for metrics, logging, tracing, and alerting. • Ensure services and data pipelines are observable, debuggable, and performant in production. • Drive performance analysis and tuning across infrastructure and service layers. • Build automation to improve system reliability, deployment safety, and recovery processes. • Partner with DevOps and Cloud Platform teams on CI/CD reliability, rollout strategies, and safe deployment patterns. • Support and improve Kubernetes-based environments and containerized workloads. • Collaborate with security teams to ensure secure and resilient system design. • Participate in disaster recovery planning and testing. • Maintain strong operational practices around access control, secrets management, and change management.
This description is a summary of our understanding of the job description. Click on 'Apply' button to find out more. Role Description Collate is the creator of the fast-growing open-source OpenMetadata project, and we’re passionate about transforming the way data teams work together. Our mission is to help every company realize the fullest potential of data through AI Agents via open-source, and unified metadata, by solving problems around data discovery, observability, and governance. What you'll be doing: - Design and build OpenMetadata as a SaaS platform that is iterating and evolving rapidly. - Put automation at the forefront of our engineering culture. Automate CI/CD and our integration tests. - Define the release methodologies and implement CI/CD for our SaaS and On-Prem releases. - Build the observability platform for our growing SaaS customers. - Provide leadership for the devops team. - Set priorities/roadmap and lead initiative on SaaS and OpenSource infrastructure. - Be a part of the team that'll change the way data products are being built in OSS. - Implement corporate Governance on Cloud service usage and Security measures. - Lead the effort to develop cloud mechanisms to improve scalability and reduce costs. - Work with our current technical stack including Java, JSON, Rest, Python, Docker, Reach/Node/TypeScript. Qualifications - 5+ years of experience in building scalable SaaS. - Experience with CI/CD setup for Cloud (ECS) and On-premises environments, as well as with Orchestration tools: Dockers, Kubernetes is beneficial. - Knowledge of DevSecOps concepts and infrastructure-as-a-code, as well as hands-on experience in security related to cloud-based infrastructure. - Provide guidance on operational concerns to development and business teams. - Experience managing and maintaining load balancers, proxies, webservers, Queues, Caches etc. - A desire to learn and grow as an engineer. - A passion to uncover and solve real-user problems. - Excitement about being an early engineer. You'll be defining our engineering culture, choosing our shared tools, and helping us build a world-class team. - Previous open-source contributions are a plus. Benefits - Competitive salary and Seed stage startup equity. - Fully remote. - Work with smart, motivated, like-minded peers, whom you can teach and learn from to grow together. - You'll be joining a small team with no bureaucracy or politics and get to work directly with the founders.


