Job Closed
This listing is no longer active.
Leidos is an innovation company rapidly addressing the world’s most vexing challenges in national security and health.
Site Reliability Engineer
Location
Florida + 2 moreAll locations: Florida | Hawaii | Virginia
Posted
135 days ago
Salary
$87.1K - $157.5K / year
Seniority
Senior
Job Description
Site Reliability Engineer
Leidos
• Work alongside the development and operations teams to ensure speedy and reliable software deployments • Monitor systems and improve overall reliability of the platform • Develop features utilizing the AI coding tool and repository of scripts to automate, scale, test, and secure the cloud infrastructure and the pipelines • Enhance performance monitoring of the various systems via Splunk or other dashboard reporting tools • Identify performance bottlenecks and optimize the performance of cloud infrastructure • Contribute to continuing the SRE journey by suggesting ways to improve engineering build, maintenance, automation and reliability across the platform with SRE/DevOps tools and Infrastructure-as-Code • Develop and code high-quality pipeline automation workflows • Create, script, and run performance tests to measure system behavior under varying levels of load and traffic • Design, implement, and maintain automated test suites for infrastructure and application components • Ensure that testing is integrated into the CI/CD pipeline to validate system reliability with every release • Build automated systems for continuous performance testing, stress testing, and load testing • Work closely with SREs, developers, and operations teams to define reliability goals and develop appropriate testing strategies to validate those goals.
Job Requirements
- Requires BS degree and with 4+ years of prior relevant experience or equivalent experience to substitute for education
- Currently possess and ability to maintain an active DoD Secret security clearance
- Minimum of DoD 8570.01 IAT Level II Certification required and must maintain certification while supporting the SMIT Contract
- Experience with automated script design, coding, debugging, and maintenance skills (using bash, python, etc.) preferred
- Experience in CI/CD toolsets (e.g. Jenkins, GitLab, etc.)
- Good command of Linux/Unix and command line knowledge
- Experience in application administration, configuration, and integration
- Familiarity with agile development methodologies
- Knowledge of Agile and DevSecOps /SRE concepts and best practices, with a desire to grow that knowledge
- Hand-on experience with Atlassian products (Jira, Confluence, Bitbucket, etc.)
- Experience creating JIRA and/or Azure DevOps workflows, projects, custom configurations
- Experience administrating/maintaining SRE platform via Ansible playbooks (e.g. upgrading Jenkins)
- Experience with PaaS using Red Hat OpenShift/Kubernetes and Docker containers
- Experience with commercial cloud infrastructure deployment environments such as AWS and Azure.
- Experience with automated provisioning and configuration tools like Terraform, Cloud Formation, Chef, Puppet, Ansible, or similar technologies.
Benefits
- Health and Wellness programs
- Income Protection
- Paid Leave
- Retirement
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Automate CI/CD pipelines to build and deploy our software. • Provisioning and maintenance of systems, load balancer configs, firewall rules, databases, and other automation driven infrastructure tasks. • Troubleshoot infrastructure and application related issues. • Work towards implementing platform stability and cost-effective practices. • Take ownership of platform standardization initiatives, involving migration to Granicus standard tools and technologies. • Escalation support for the SRE team.
Senior DevOps Software Engineer
eClinical SolutionsWe bring people and data together to support tomorrow’s breakthroughs
Senior DevOps Software Engineer eClinical Solutions is transforming clinical development with elluminate®, our Clinical Data Cloud, helping life sciences organizations unify, analyze, and unlock the value of their data faster than ever before. By combining a modern cloud platform with expert data services, we empower smarter decisions across the clinical trial lifecycle—accelerating innovation that ultimately improves patient outcomes. Our engineering teams build enterprise-grade, revenue-generating products at the intersection of cloud, data, analytics, and emerging AI technologies. If you’re excited about building sophisticated software that makes a real-world impact in healthcare, this is the place to do it. You will make an impact: As a Senior DevOps Software Engineer, you will serve as a technical leader within the product development organization. You will be responsible for designing and architecting AWS-based infrastructure that powers a reliable, scalable, secure, high-performing Clinical Data Intelligence Platform as well as operationalizing AI/ML capabilities across the platform and engineering team. You will drive infrastructure design and implementation from the ground up, building fully automated testing and code delivery mechanisms with minimal manual intervention. The environment requires strong SDLC traceability, governance, and controlled promotion of software artifacts across environments. In this role, you will influence engineering standards, infrastructure strategy, and technical direction across teams. You will collaborate closely with experienced engineers and leverage tools such as Bitbucket, TeamCity, Octopus Deploy, and Jira to enable automated CI/CD pipelines, artifact traceability, and controlled releases. This role is ideal for an experienced, self-driven engineer who enjoys hands-on architecture, mentoring others, and introducing modern capabilities—including AI-assisted and agentic workflows—into production-grade systems. Accelerate your skills and career within a fast-growing company while impacting the future of healthcare. Your day to day: - Design, develop, test, and deploy scalable, secure, and highly interactive web applications - Own and evolve core platform modules, from concept through release and support - Influence application and system architecture with a focus on performance, reliability, security, and maintainability - Lead by example through clean, well-tested code, thoughtful design reviews, and pragmatic technical decisions - Collaborate closely with Product Management, QA, and other engineers throughout the SDLC - Provide technical mentorship and guidance to other engineers on the team - Diagnoses and resolves complex production issues across distributed systems - Ensure solutions meet eClinical Solutions quality standards and applicable industry regulations - Contribute to technical documentation including design specs, acceptance criteria, and release notes Take the first step towards your dream career. Here is what we are looking for in this role. Qualifications: - Bachelor’s degree or higher preferred (Computer Science, Data Science, Engineering, or related field) and/or equivalent work experience preferred - 10+ years of hands-on enterprise software engineering experience in scalable production environments preferred - 8+ years of practical AWS experience preferred - Experience managing and operating SaaS platforms on Kubernetes (EKS) and ECS, including container build, deployment, lifecycle management, scaling, and troubleshooting - Proven experience fully automating software release processes, scaling from infrequent releases to continuous delivery with multiple daily deployments, including automated rollback based on test validation. - Implement and support MLOps pipelines for model training, validation, deployment, monitoring, and retraining. - Strong experience designing and implementing Infrastructure as Code using Terraform and CloudFormation - Advanced scripting experience (Python, PowerShell, Bash) for automation and platform tooling - Responsible for end-to-end support of automated test execution within CI/CD pipelines using TeamCity and Octopus Deploy, including test orchestration, environment provisioning, parallel execution optimization, reporting, and enforcement of quality gates - Maintain integration with Bitbucket for source control and pull request workflows, and Jira for automated traceability between code changes, builds, deployments, and work items - Strong experience with application monitoring and observability using CloudWatch, Datadog, and/or New Relic - Experience with building infrastructure prepared to meet rigorous Disaster Recovery requirements - Strong analytical and problem-solving skills - Excellent planning, coordination, and organizational skills - AWS: - Networking - VPC, subnets, route tables, NAT, ALB/NLB, Route 53) - Compute - EC2, Auto Scaling, Lambda, ECS, EKS) - Storage - S3, EBS, EFS, FSx) - Security - Solid understanding of security and compliance principles, including IAM, WAF, encryption (KMS), and integration of SAST, SCA, and DAST into CI/CD pipelines - Programming/Scripting: Python, PowerShell, bash, SQL, .NET - DevOps Tools: TeamCity, Octopus Deploy, Git-based workflows - Containerization & Orchestration: Docker, Kubernetes and AWS EKS and ECS - Infrastructure as Code: Terraform or equivalent - Databases: SQL Server or equivalent - AI/ML Frameworks: TensorFlow, PyTorch, scikit-learn (preferred) - MLOps Tools: MLflow or equivalent - Monitoring & Observability tools: DataDog, NewRelic Accelerate your skills and career within a fast-growing company while impacting the future of healthcare. We have shared our story, now we look forward to learning yours! eClinical is a winner of the 2025 Top Workplaces USA Award for Remote Work! We have also received numerous Top Workplaces Culture Excellence Awards celebrating our exceptional company vision, values, and work-life balance. See all the details here: https://topworkplaces.com/company/eclinical-solutions/ eClinical Solutions is a people first organization. Our inclusive culture values the contribution that diversity brings to our business. We celebrate individual experiences that connect us and that inspire innovation in our community. Our team seeks out opportunities to learn, grow and continuously improve. Bring your authentic self, you are welcome here! We are proud to be an equal opportunity employer that values diversity. Our management team is committed to the principle that employment decisions are based on qualifications, merit, culture fit and business need. #LI-AB1 Pay Range US Pay Ranges $132,000—$165,000 USD
Director of DevOps, Site Reliability Engineering
CargoSprintEmpowering the people that make global commerce happen.
• Lead and mentor a distributed team of DevOps, SRE, and Database engineers. • Architect and operate secure, scalable, and cost-efficient Azure Cloud environments. • Implement and optimize CI/CD pipelines, infrastructure as code (IaC), and observability platforms. • Champion AIOps and AI-driven tooling (e.g., GitHub Copilot, Azure DevOps AI, intelligent alerting) to improve developer productivity and operational efficiency. • Establish and enforce SRE practices — SLIs/SLOs, incident response, on-call processes, and postmortems. • Oversee performance, scalability, and reliability of PostgreSQL, MySQL, SQL Server, CosmosDB, and Redis databases in production. • Partner cross-functionally with product and engineering teams to align infrastructure with business priorities. • Drive cost optimization, disaster recovery, and security compliance initiatives.
Deployment Engineer
Simbe RoboticsFounded in 2014, Simbe builds automation solutions for retailers. The company's first product, Tally, is the world's first entirely autonomous shelf-auditing an
• Manage the full deployment lifecycle for assigned retail locations: pre-deployment verification, physical installation coordination, network configuration, digital setup, and handoff to operations. Own quality and timeline for each deployment. • Confirm all site requirements are met prior to deployment — including network credentials, docking locations, traversal area definitions, and hardware readiness. Identify and resolve blockers early to protect deployment timelines. • Coordinate remote contractors via chat and voice through the physical installation process. Provide timely, clear support to field teams and ensure deployment procedures are followed consistently. • Configure robots for new sites including traversal pattern setup, schedule management, and route mapping. Validate initial performance against quality thresholds across decode accuracy, OOS detection, price verification, location coverage, and upload speed. • Diagnose and resolve issues across network connectivity, hardware, and software encountered during deployment. Triage root causes systematically and escalate with clear documentation when needed. • Build scripts, workflow automations, or lightweight internal tools to reduce manual effort in the deployment process — including pre-deployment checklists, status tracking, configuration validation, and reporting. Use AI-assisted development to accelerate tooling where applicable. • Develop and maintain QA checks for deployment readiness and early operational performance. Identify patterns in deployment failures and build detection or prevention mechanisms. • Maintain accurate records of deployments, issues, and resolutions in Jira and Confluence. Contribute to deployment playbooks and continuously improve procedures based on field learnings.



