Job Closed

This listing is no longer active.

Fathom Management LLC logo
Fathom Management LLC

Fathom Management, Inc. is an Equal Opportunity Employer committed to fostering a diverse and inclusive workplace. All employment decisions are made without regard to any protected characteristic under applicable law.

Kubernetes / AWS DevOps Engineer – VA Lighthouse Platform (Remote)

DevOps EngineerDevOps EngineerOtherRemoteTeam 1

Location

United States

Posted

145 days ago

Salary

$130K - $150K / year

No structured requirement data.

Job Description

Kubernetes / AWS DevOps Engineer – VA Lighthouse Platform (Remote)

Fathom Management LLC

Kubernetes / AWS DevOps Engineer – VA Lighthouse Platform (Remote) Remote (U.S.) Supporting the U.S. Department of Veterans Affairs (VA) Full-Time Salary: $130,000 – $150,000 annually (based on experience) Public Trust Clearance Required (or ability to obtain) Mission Impact This role supports modernization of digital healthcare infrastructure for the U.S. Department of Veterans Affairs (VA) by enabling secure, scalable cloud-native platforms and DevSecOps pipelines that power VA Lighthouse APIs and digital health services. Your work will help expand 24/7 healthcare access, secure data exchange, and mission-critical Veteran services nationwide. Position Overview Fathom Management, Inc. is seeking an experienced DevOps Engineer to support a federal healthcare modernization program. In this role, you will design and implement cloud-native infrastructure, DevSecOps pipelines, and Kubernetes-based platforms that support secure and scalable application deployment. You will collaborate with software engineers, security teams, product owners, and federal stakeholders to build reliable infrastructure that supports VA Lighthouse platform capabilities, healthcare APIs, and digital service delivery for Veterans. Technology Stack Kubernetes, AWS Cloud, DevSecOps CI/CD, Terraform / OpenTofu, ArgoCD, Service Mesh (Istio/Cilium), Docker, GitOps, API Platforms, Monitoring & Observability Key Responsibilities Cloud Architecture & Platform Engineering - Support the architectural vision for VA Lighthouse platforms, focusing on scalability, automation, developer experience, and security. - Build automated self-service platforms for deploying containerized applications and exposing secure API endpoints to federal data consumers. - Deliver Platform-as-a-Service (PaaS) capabilities, including low-code/no-code platforms and enterprise-grade API engineering solutions. DevSecOps & Automation - Design and maintain secure CI/CD pipelines supporting automated build, test, deployment, and infrastructure provisioning. - Implement DevSecOps best practices, including shift-left security and automated vulnerability scanning. - Automate infrastructure using Infrastructure-as-Code (IaC) frameworks and scripting. Kubernetes & Platform Operations - Design, deploy, and maintain Kubernetes clusters in production environments. - Support L4–L7 networking in AWS environments, including technologies such as Cilium, Istio, or similar service mesh operators. - Implement monitoring, logging, and observability solutions to ensure platform health and operational insights. Governance & Compliance - Ensure solutions meet federal compliance standards, including audit readiness, security documentation, and platform governance. - Collaborate with product owners, engineering leads, and security teams to define platform roadmaps. - Establish and promote DevOps and engineering best practices across development teams. Continuous Platform Innovation - Continuously improve the VA Lighthouse platform infrastructure through automation and emerging cloud technologies. - Deliver reliable, scalable, and secure platform capabilities supporting healthcare modernization. Required Qualifications - 5+ years of experience with AWS cloud platforms and DevOps engineering - Deep hands-on experience with Kubernetes including design, implementation, troubleshooting, and scaling - Strong understanding of networking layers (L4–L7) within AWS environments - Experience with service mesh technologies such as Cilium or Istio - Experience implementing Infrastructure-as-Code (IaC) frameworks and automation scripts - Experience defining monitoring and observability solutions for platform reliability - Experience leading engineering efforts or defining platform standards - Must be legally authorized to work in the United States without employer sponsorship - Ability to obtain and maintain a VA Public Trust clearance Preferred Qualifications - 8+ years of experience in software, DevOps, or platform engineering - Experience with Terraform or OpenTofu - Experience with configuration management tools such as Ansible, Puppet, or Chef - Experience implementing DevSecOps pipelines and containerization tools such as Docker and Kubernetes - Experience with Blue/Green deployments using ArgoCD - Experience implementing end-to-end encryption with managed mTLS - Familiarity with federal compliance frameworks such as FedRAMP or FISMA - Experience working within government or public sector environments - Familiarity with VA Lighthouse APIs Education Requirements - Bachelor's degree in Computer Science, Engineering, or related technical field, or equivalent professional experience. Benefits & Career Growth At Fathom Management, Inc., we value our employees and offer a competitive benefits package designed to support health, financial well-being, and professional growth. Employee Benefits Include - Paid vacation, sick leave, and company holidays - Medical, dental, and vision insurance - Life insurance coverage - Short-term and long-term disability insurance - 401(k) retirement plan with company match and immediate vesting - Military leave benefits - Training and professional development opportunities - Tuition reimbursement - Employee wellness initiatives - Commuter benefits - Additional voluntary benefits About Fathom Management, Inc. Fathom Management, Inc. supports federal healthcare modernization initiatives, including transformation programs for the U.S. Department of Veterans Affairs (VA). Our teams specialize in cloud modernization, healthcare technology integration, DevSecOps engineering, and large-scale program delivery to improve healthcare access, operational efficiency, and the Veteran experience. Joining Fathom Management means contributing to mission-critical programs that strengthen healthcare delivery for Veterans nationwide. Equal Employment Opportunity (EEO) Statement Fathom Management, Inc. is an Equal Opportunity Employer committed to fostering a diverse and inclusive workplace. All employment decisions-including recruitment, hiring, training, promotion, compensation, benefits, and termination-are made without regard to race, color, religion, creed, national origin, sex, age, marital status, sexual orientation, gender identity, citizenship status, veteran status, disability, or any other characteristic protected by applicable federal, state, or local law.

Related Categories

Related Job Pages

More DevOps Engineer Jobs

OZ logo

Azure DevOps

OZ

A leading consulting company whose Intelligent Automation expertise accelerates the way you do business.

DevOps Engineer145 days ago
Full TimeRemoteTeam 201-500H1B Sponsor

• Deliver excellent service through actively learning the client’s business and business processes, responding to the needs of the internal customers, and following through on commitments. • Responsible for designing and developing an appropriate solution that conforms to and satisfies the client's business needs. • Responsible for making improvement recommendations to the Senior Manager/Director, Applications Development/Software Engineering concerning changes in business process, internal department process and software development tools. • Building and implementing new development tools and infrastructure. • Understanding the needs of stakeholders and conveying them to developers. • Working on ways to automate and improve development and release processes. • Testing and examining code written by others and analyzing results. • Ensuring that systems are safe and secure against cybersecurity threats. • Identifying technical problems and developing software updates and fixes. • Working with software developers and software engineers to ensure that development follows established processes and works as intended. • Planning projects and being involved in project management decisions. • Mentor and act as a technical role model for junior resources. • Ensure complete issue tracking and reporting are maintained. • Ensure code compliance and versioning using the Company’s dedicated source code management solution. • Provide guidance with technical design. • Work with Business Analysts, QA Analysts, Process Owners, and other cross-functional resources to define and deliver business-impacting projects. • Work directly with stakeholders to capture business requirements and translate them into technical approaches and designs that comply with the client’s technical requirements. • Collaborate with development team members to ensure proper implementation and integration of the solutions. • Support deployments or troubleshoot production issues outside of work hours and participate in an on-call rotation as needed. • Assist with the implementation of new software enhancements, system processes, and/or 3rd party products within the software applications ecosystem. • Maintain appropriate software on server and client development computers. • Maintain proper documentation throughout the software development lifecycle of assigned projects. • Participates in the creation of the supported software and hardware lists. • Accountable for adherence to IT dept. standards, including proper design, project documentation, coding standards, and approval processes. • Provide timely and accurate responses to local and remote users’ concerns relating to applications. • Must provide excellent customer service through end-user training and SOX documentation. • Participate in design and development of physical and logical application frameworks that are extensible, stable, and re-usable. • Ensures Sarbanes Oxley compliance on all initiatives. • Maintains sign-off documentation and other SDLC methodology documentation as required.

Argentina
Job Closed
Full TimeRemoteTeam 10,001+H1B Sponsor

• Ensure DevSecOps, security, observability and FinOps practices are aligned with the Central Bank of Brazil's regulations applicable to the fintech's operations. • Support risk, compliance and audit teams in the technical interpretation and implementation of regulatory requirements. • Design and maintain monitoring, metrics and distributed tracing solutions. • Build and evolve dashboards, alerts and reliability indicators (SLOs, SLAs and SLIs). • Implement instrumentation standards using OpenTelemetry. • Integrate logs, metrics and traces for proactive incident detection. • Support teams with performance analysis and advanced troubleshooting of applications and infrastructure. • Monitor, analyze and optimize cloud infrastructure costs (Azure or GCP). • Create reports, define tagging standards and support cost governance. • Propose architectural improvements with a focus on financial efficiency. • Implement anomaly alerts and resource consumption forecasting. • Support squads in cost-aware decision making.

Brazil
Job Closed
AWP Safety logo

Site Reliability Engineering (SRE) Intern

AWP Safety

The AWP Safety FP&A Internship Program provides a hands‑on, high‑impact learning experience designed for early‑career professionals who want to build a future in Financial Planning & Analysis. Interns will partner directly with corporate and operational finance leaders on critical projects that support organizational performance, financial accuracy, and strategic decision‑making. While this internship is primarily project‑based and can be remote depending on location, interns will also have opportunities to collaborate closely with cross‑functional teams to understand how financial insights drive real‑world business outcomes. Receive one‑on‑one mentorship from Senior FP&A leaders. Attend workshops, panels, and intern networking events. Participate in our “Journey‑to‑the‑Job” series to hear from seasoned executives, sharing their diverse career paths within the organization.

DevOps Engineer145 days ago
OtherRemoteTeam 5,001-10,000

Company Description The AWP Safety IT Internship Program immerses you in provides a hands‑on, high‑impact learning experience designed for early‑career professionals who want to build a future in IT Site Reliability Engineering. In this role, you won't just be watching application performance monitoring dashboards; you will be building the observability pipelines that keep our infrastructure and applications resilient, highly available, and robust. You will work at the intersection of Software Engineering and Systems Operations, using Dynatrace as your primary lens to diagnose performance bottlenecks and automate "toil" out of existence. While this internship is primarily project‑based and can be remote depending on location, interns will also have opportunities to collaborate closely with cross‑functional teams to understand how technical insights drive real‑world business outcomes. What You’ll Experience - Full‑Stack Observability: Trace requests from browser to code to database. - Incident Lifecycle: Join blameless post‑mortems and help implement “never‑twice” fixes. - AIOps: Use Dynatrace’s predictive AI to find “the needle in the haystack” before an outage occurs. - Scalable Infrastructure: How to manage monitoring for thousands of hosts without manual intervention. Professional & Team‑Building Activities - Attend workshops, panels, and intern networking events. - Participate in our “Journey‑to‑the‑Job” series to hear from seasoned executives, sharing their diverse career paths within the organization. Job Description This 10-week internship places interns at the center of our IT operations, offering meaningful work with real organizational impact. You’ll thrive if you have a passion for “measuring everything”. You’ll collaborate closely with Platform, AppDev, and Security teams on production‑grade outcomes for our business. Core Responsibilities - Observability‑as‑Code: Help deploy and configure Dynatrace OneAgent and ActiveGates with automated tooling. - SLI/SLO Implementation: Define and instrument user‑centric metrics and objectives in Dynatrace. - AI‑Assisted Troubleshooting: Combine Davis® AI with Copilot/Claude to identify root causes and reduce MTTR. - Dashboard Engineering: Build actionable, real‑time dashboards for application and cloud health. - Automation & Scripting: Write Python/Bash to trigger self‑healing or response playbooks from alerts. Qualifications - Rising junior/senior or current master’s student. - Clear communication and teamwork skills in fast‑moving ops environments. - Systems Thinking: Understand how web apps, databases, and networks interact. - SRE Mindset: Care deeply about reliability, scalability, and error budgets. - Scripting Proficiency: Familiarity with Python, Go, or PowerShell. - Cloud Basics: Exposure to containers (Docker/Kubernetes) and microservices patterns. - Data Fluency: Read metrics, logs, and traces to tell a story about system health. - Clear communication and teamwork in fast‑moving ops environments. Additional Information - Full‑time, 10‑week temporary internship; non‑benefits eligible - Compensation: $30-34/hour based on location Join us for an IT internship that strengthens your technical abilities, builds your professional confidence, and prepares you for a future in high‑impact SRE roles. Apply today and help shape the technical insights that power AWP Safety. AWP Safety is an Equal Opportunity Employer (EOE). Women, minorities, veterans, and individuals with disabilities are encouraged to apply. Qualified applicants will receive consideration for employment without regard to their race, color, age, religion, national origin, sex, sexual orientation, gender identity, protected veteran status or disability. - Compensation: USD 30 - USD 34 - hourly

United States
$30 - $34 / hour
Job Closed
Leidos logo

Azure Cloud Infrastructure Ops Support

Leidos

Leidos is an innovation company rapidly addressing the world’s most vexing challenges in national security and health.

DevOps Engineer145 days ago
OtherRemoteTeam 10,001+Since 1969H1B Sponsor

Leidos was awarded the U.S. Air Force Cloud One Architecture and Common Shared Services contract, and currently has an opening for an Azure Cloud Infrastructure Ops Support Engineer across AWS, Azure, Google, and Oracle clouds. This is an exciting opportunity to use your experience to modernize a leading, global-scale multi-cloud environment in support of a critical mission, supporting USAF system resiliency, security, and cost effectiveness. Location: This position will be remote, candidate will be required to work onsite as needed. Primary Responsibilities: We are seeking an Azure Cloud Infrastructure Ops Support Engineer with expertise in multiple cloud platforms. A successful individual will be responsible for developing in a scalable cloud-native solutions, and ensuring best practices across architecture, development, deployment, and security from design, test, integration, production, sustainment and maintenance. This is a hands-on technical role that requires rolling up your sleeves to architect, code, debug, and mentor.  - Perform cloud infrastructure operations and engineering tasks to enhance, sustain, and maintain scalable, resilient, and secure cloud solutions for CloudOne multi-cloud environment - Perform specified cloud operations, sustainment, and maintenance tasks to maintain optimum cloud capabilities - Utilize DevOps practices, infrastructure as code, and automation frameworks  - As tasked, perform development activities to optimize application performance and reliability in cloud environments  - Participate on DevSecOps teams to design, implement and sustain secure cloud architectures and networks implementing zero-trust principles and defense-in-depth strategies  - Perform tasks as directed to implement and maintain cloud networking security controls including STIG requirements utilizing Assured Compliance Assessment Solution (ACAS), Tenable Security Center, Nessus - Support development of migration methodologies and application containerization to ensure minimal organizational disruption during transitions  - Utilize CI/CD workflows and infrastructure-as-code development using tools such as, Jenkins, Terraform, Ansible, Kubernetes, Jira, Confluence, Artifactory, and Guacamole to support DevSecOps practices. - Support the design, development, implementation, and sustainment of Shared Services. - As tasked perform configuration and troubleshooting of cloud, virtual, physical hardware and software systems. - Support preparation of detailed technical documentation of development and operational processes. - Work in cross-functional teams including development, operations, security, and product management  Minimum Qualifications - Bachelors and less than 2 years of experience; Masters degree a plus. Additional experience may be accepted in lieu of degree. - Secret Clearance - Certifications: CompTIA Security+ or equivalent (IAT-2) - Practiced verbal and written communications skills - Ability to participate in team efforts to accomplish assigned tasks  - Demonstrated experience in cloud operations and sustainment and performing tasks and actions described in the primary responsibilities section ​ Preferred Qualifications - Experience with USAF Cloud One or Platform 1 - Knowledge of Zero Trust Architecture. Experience a plus. - Capable of working in high powered teams and maintaining positive interpersonal relationships while delivering products and services to the customer - Understanding of tools such as Ansible, Azure console, Elastic, Jira, Confluence, Git, Bitbucket and various cloud Software as a Service (SaaS) offerings to conduct DEV/SEC/OPS pipeline development activities - ​Experience administering Windows Server, and related services  - Cloud certifications in AWS, Azure, Google, or Oracle clouds - Certification Examples - ​CompTIA Security+, Azure Certified Cloud Practitioner, ITIL 4 Foundation C1NACSS If you're looking for comfort, keep scrolling. At Leidos, we outthink, outbuild, and outpace the status quo — because the mission demands it. We're not hiring followers. We're recruiting the ones who disrupt, provoke, and refuse to fail. Step 10 is ancient history. We're already at step 30 — and moving faster than anyone else dares. Original Posting: March 6, 2026 For U.S. Positions: While subject to change based on business needs, Leidos reasonably anticipates that this job requisition will remain open for at least 3 days with an anticipated close date of no earlier than 3 days after the original posting date as listed above. Pay Range: Pay Range $57,850.00 - $104,575.00 The Leidos pay range for this job level is a general guideline only and not a guarantee of compensation or salary. Additional factors considered in extending an offer include (but are not limited to) responsibilities of the job, education, experience, knowledge, skills, and abilities, as well as internal equity, alignment with market data, applicable bargaining agreement (if any), or other law.

United States
$57.9K - $104K / year
Job Closed