Senior Development Operations Engineer
Location
United States
Posted
6 days ago
Salary
0
Seniority
Senior
No structured requirement data.
Job Description
Senior Development Operations Engineer
Rapta, Inc
Role Description We're seeking an experienced senior development operations (DevOps) engineer to own the build, test, and deployment pipeline for our distributed computer vision platform running at edge sites across customer manufacturing floors. This is a full-time remote position working directly with our engineering team to harden our release process, drive deployment automation, and pioneer LLM-driven test generation and validation at Rapta. - Own and evolve the end-to-end release pipeline — branching strategy, build orchestration, artifact promotion, and rollback — across our Bazel monorepo and Python deployable units. - Design and maintain Ansible-driven fleet automation for heterogeneous Linux edge nodes (Ubuntu LTS, NVIDIA driver stacks, Docker with NVIDIA runtime). - Manage all update tooling, currently written in Golang. - Build LLM-powered automated testing systems: test generation from specs, flake triage, log/failure analysis, regression diffing, and release-note synthesis from commit and ticket history. - Harden CI/CD for offline and bandwidth-constrained deployment targets (airgap wheel distribution, signed artifacts, deterministic builds). - Drive observability for releases — deployment telemetry, version drift detection, and post-deploy health validation across the fleet. - Mentor engineers on release hygiene, reproducible builds, and infrastructure-as-code practices. Qualifications - 10+ years of professional experience in release engineering, DevOps, or SRE roles shipping production Linux systems. - Deep curiosity for software, infrastructure, and applied AI — particularly using LLMs as production engineering tools, not just chat assistants. - Expert-level Python (3.8+) with a strong grasp of packaging, dependency resolution, and PEP 440 versioning discipline. - Demonstrated ownership of Linux fleets at scale — kernel, systemd, networking, package management. - Excellence in technical communication, runbook authorship, and post-incident documentation. - Strong systems thinking — comfortable reasoning about failure modes across hardware, OS, container, and application layers. Requirements - Expert proficiency with Ansible (roles, dynamic inventory, idempotent design); working knowledge of Terraform. - Expert proficiency with Docker, including creation and lifecycle management of containers, image hardening, registry management and installing & configuring the NVIDIA container runtime. - Production experience with Linux administration: systemd, networking (VLANs, DHCP, DNS), kernel/driver management (especially NVIDIA/DKMS), package and APT internals. - Strong Python skills focused on tooling, automation, packaging (wheels, pip, private indexes), and subprocess/CI integration. - Proficiency with Git workflows, branching strategies, and modern CI/CD systems (GitHub Actions, GitLab CI, or equivalent). - Experience designing and operating automated test infrastructure — unit, integration, hardware-in-the-loop, and end-to-end. - Practical experience using LLMs (Anthropic, OpenAI, or local) as part of engineering workflows — test generation, code review augmentation, log analysis, or agentic tooling. Nice to Have - Bazel or similar monorepo build systems. - Edge or embedded deployment experience. - Tailscale, WireGuard, or zero-trust networking in production. - gRPC/protobuf service ecosystems. - Vault, PKI, or secrets management at fleet scale. - Background in regulated or compliance-driven environments (CMMC, ISO 27001, SOC 2). Benefits - Work on cutting-edge AI infrastructure with real-world impact on American manufacturing. - Build the release engineering foundation for a late-seed company actively scaling. - Direct collaboration with the CTO and engineering leadership. - Remote-first culture with flexible hours. - Meaningful equity in a company solving hard problems. Location Requirements - Remote (US). Equal Opportunity Rapta is committed to hiring and retaining a diverse workforce. We are proud to be an Equal Opportunity/Affirmative Action Employer, making decisions without regard to race, color, religion, creed, sex, sexual orientation, gender identity, marital status, national origin, age, veteran status, disability, or any other protected class. How to Apply No recruiters or agencies — we only accept applications directly from applicants.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Role Description TESTING-INOVANCE Broadcom is proud to be an equal opportunity employer. We will consider qualified applicants without regard to: - Race - Color - Creed - Religion - Sex - Sexual orientation - National origin - Citizenship - Disability status - Medical condition - Pregnancy - Protected veteran status - Any other characteristic protected by federal, state, or local law We will also consider qualified applicants with arrest and conviction records consistent with local law. If you are located outside USA, please be sure to fill out a home address as this will be used for future correspondence. Company Description
DevOps Automation Engineer
Rootshell Enterprise Technologies, Inc.Rootshell Enterprise Technologies Inc. is a recognized provider of professional IT Consulting services in the US.
Role Description We are actively seeking a DevOps Automation Engineer for one of our clients. Responsibilities - Serve as a primary point of contact for infrastructure automation configuration via Terraform, Ansible, Python, Cloud Formation. - Responsible for the overall health, performance, and capacity of our Internet-facing services. - Create dashboards and other tools for DevOps to use in day-to-day monitoring and troubleshooting. - Perform advanced troubleshooting and monitoring of the systems to ensure SLAs are met. - Collaborate with other engineering teams on design and expansion of our infrastructure. - Assist in the roll-out and deployment of new product features and installations to facilitate our rapid iteration and constant growth. - Work closely with development teams to ensure that platforms are designed with operability in mind. - Drive standardization efforts across multiple services. - Identify and lead efforts to improve automation. - Participate in an on-call rotation. - Function well in a fast-paced, rapidly-changing environment. Skills - 5+ years of DevOps or Site Reliability experience in Infra Automation. - Experience in designing and building Test Automation Frameworks for cloud-based products from the ground up. - Experience with Automation Tools such as Ansible, Terraform, Puppet. - Experience with Continuous Integration Tools - Jenkins. - Solid Experience with AWS - EC2, S3, ELB, Auto Scale, CloudFormation. - Experience with Data Monitoring Tools - Nagios, Grafana, Cloudwatch, etc. - Strong interpersonal communication skills. - Experience with Apache/Nginx. - Experience with Kernel and network-based tuning for performance and stability. Company Description Rootshell Enterprise Technologies Inc. is a recognized provider of professional IT Consulting services in the US.
Role Description Support and maintain business-critical production environments, ensuring high availability and system reliability. - Monitor infrastructure and applications, proactively identifying and resolving issues before they impact users. - Participate in incident response activities, troubleshooting production outages and coordinating recovery efforts. - Perform Root Cause Analysis (RCA) and contribute to postmortems, corrective actions, and continuous improvement initiatives. - Manage and optimize Kubernetes clusters and cloud infrastructure. - Develop and maintain monitoring dashboards, alerts, and observability solutions. - Automate operational processes and infrastructure deployments using Infrastructure as Code (IaC) and scripting. - Collaborate with engineering and product teams to improve scalability, performance, and operational excellence. - Support and enhance CI/CD pipelines to ensure reliable and efficient software delivery. Qualifications - 4+ years of experience in Site Reliability Engineering, DevOps, Cloud Operations, or Infrastructure Engineering. - Strong hands-on experience with Linux administration, troubleshooting, and production support. - Experience managing and supporting Kubernetes and containerized workloads (Docker/OpenShift is a plus). - Solid knowledge of AWS, Azure, or GCP cloud environments. - Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, Splunk, ELK, or CloudWatch. - Experience with Infrastructure as Code (Terraform preferred) and CI/CD pipelines. - Ability to troubleshoot complex production issues, perform Root Cause Analysis (RCA), and drive preventive improvements. - Working knowledge of automation and scripting using Bash, Python, or Go. - Intermediate to advanced English (B2+). Company Description
SAP NS2 Sr. DevOps Engineer
SAPSAP, an acronym for Systemanalyse und Programmentwicklung, or Systems Analysis and Program Development in English, provides clients in 180+ countries with enter
Role Description We are seeking an Expert DevOps Engineer to join our team and play a critical role in building and automating the deployment processes for SAP PCE (Private Cloud Edition) products. This position requires a skilled engineer who can design, implement, and optimize infrastructure as code across multiple cloud platforms while driving automation and process improvements. Key Responsibilities: - Infrastructure & Automation - Design, build, and maintain automated build and deployment pipelines for SAP PCE products - Develop and manage infrastructure as code using Terraform across AWS, GCP, and Azure environments - Create and maintain Ansible playbooks for configuration management and application deployment - Implement automation solutions using Python to improve operational efficiency and reduce manual processes - Ensure consistency, reliability, and repeatability of deployment processes across all cloud platforms - Cloud Platform Management - Deploy and manage infrastructure across multi-cloud environments (AWS, GCP, Azure) - Optimize cloud resource utilization and implement cost-effective solutions - Maintain security best practices and compliance requirements across all cloud platforms - Troubleshoot and resolve complex infrastructure and deployment issues - Collaboration & Leadership - Work cross-functionally with development, security, operations, and product teams - Represent the DevOps team in technical discussions with stakeholders and leadership - Provide technical guidance and mentorship to team members - Communicate complex technical concepts effectively to both technical and non-technical audiences - Participate in architecture and design reviews - Continuous Improvement - Identify opportunities for process optimization and automation - Implement monitoring, logging, and alerting solutions for infrastructure and applications - Develop and maintain documentation for processes, procedures, and infrastructure - Stay current with industry trends and emerging technologies in DevOps and cloud computing Qualifications - Bachelor’s degree in Computer Science, Information technology or related. Years of experience may be used in lieu of a degree. - 7+ years of experience in DevOps, Site Reliability Engineering, or related field - Proficiency with Terraform for infrastructure as code - Strong experience with Ansible for configuration management and automation - Hands-on experience deploying and managing infrastructure across public clouds such as AWS, GCP, and Azure - Proficiency in Python for automation and scripting - Deep understanding of CI/CD principles and tools - Experience with version control systems (Git) - Understanding of networking, security, and cloud architecture principles - Excellent communication and interpersonal skills - Strong problem-solving and analytical abilities - Ability to manage multiple priorities and projects simultaneously - Proven ability to work effectively in cross-functional teams - Experience representing technical teams to stakeholders and leadership Benefits - Constant learning and skill growth - Great benefits - A team that wants you to grow and succeed


