Job Closed
This listing is no longer active.
Lead DevOps Engineer – Data, AWS, Kubernetes, Cloud Infrastructure, MSSQL Server
Location
Texas
Posted
56 days ago
Salary
$173.5K - $281.0K / year
Seniority
Senior
Job Description
Lead DevOps Engineer – Data, AWS, Kubernetes, Cloud Infrastructure, MSSQL Server
Triumph Financial, Inc.
• Design, build, and maintain AWS-based cloud infrastructure, CI/CD pipelines, and automation tools • Partner with development teams to ensure applications are scalable, reliable, and secure • Optimize and manage Kubernetes clusters for performance, scalability, and consistency • Develop and maintain Helm charts for containerized applications • Manage and support Kafka clusters and streaming infrastructure • Collaborate with security teams to meet compliance standards (SOC2, SOX, FFIEC) • Monitor system performance using tools like Grafana, Loki, and other observability platforms • Automate deployments, configurations, and operational processes • Recommend and implement improvements to infrastructure design and DevOps practices • Define and implement SLOs and SLIs to enhance system reliability • Continuously improve DevOps workflows and platform efficiency.
Job Requirements
- 5+ years of experience in DevOps, cloud engineering, or infrastructure roles
- Strong hands-on experience with AWS services (EC2, S3, EKS, RDS, IAM, VPC, and more)
- Deep expertise in Kubernetes and Helm
- Experience with Terraform (Terragrunt is a plus) and CI/CD tools like Argo CD or GitHub Actions
- Familiarity with Kafka, Redis, and Postgres in production environments
- Experience managing MSSQL Server and performing database migrations
- Solid understanding of networking in cloud and hybrid environments
- Comfortable working in Linux environments
- Experience supporting large-scale, multi-account AWS environments
- Familiarity with Snowflake or Looker is a plus
- Strong troubleshooting and problem-solving skills
- Ability to communicate technical concepts clearly to different audiences
- Experience working in Agile teams
- Motivated to learn, grow, and pursue technical certifications.
Benefits
- Medical
- Dental
- Vision
- Paid Time Off
- 401k
- and much more.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Design and maintain Kubernetes-based infrastructure, including cluster provisioning, RBAC configuration, network policy, and workload management • Package and deploy applications using Helm charts; maintain chart repositories and manage release lifecycle across environments • Implement and enforce policy controls using Istio service mesh, OPA Gatekeeper, Kyverno, and related Kubernetes admission controllers • Build and maintain CI/CD pipelines using GitLab CI, GitHub Actions, Jenkins, or equivalent tooling; integrate automated security scanning and compliance gates • Deploy and operate workloads on AWS GovCloud and Azure Government; architect for high availability, disaster recovery, and cross-region compliance requirements • Manage and harden container images; integrate with Iron Bank, Platform One, and other DoW-approved registry sources • Configure and maintain observability stacks including Prometheus, Grafana, and Datadog; develop alerting, dashboards, and SLO frameworks • Participate in ATO processes, support STIG/CIS compliance scanning, and contribute to System Security Plans (SSPs) and documentation artifacts • Collaborate with development, security, and program teams to establish and refine DevSecOps practices across the software delivery lifecycle • Support air-gapped and classified environment deployments; design solutions for offline image transfer, registry mirroring, and artifact management • Coordinate with government platform teams and managed service providers to integrate and sustain vendor tooling within approved DoD software factories
• Own and operate the production environment, maintaining a holistic view of system health, availability, and performance • Design, build, and maintain infrastructure and platform systems using automation-first principles • Proactively create monitors and observability strategies (e.g., Sumo Logic, New Relic, Prometheus) to prevent issues before they occur • Apply AI tools (Claude, Cursor) to enhance infrastructure operations, including debugging, monitoring, and performance optimization at the system level • Measure, analyze, and optimize system performance to continuously improve uptime and user experience • Provide operational support for large, distributed systems, including participation in on-call and incident response • Diagnose and resolve complex production issues, and mentor others in debugging and troubleshooting approaches • Partner with engineering teams to improve reliability through testing, release processes, and system design • Contribute to and influence system architecture, ensuring scalability, resilience, and long-term maintainability • Participate in capacity planning, platform management, and infrastructure roadmap discussions • Drive automation and continuous improvement across infrastructure and deployment workflows • Balance speed and reliability through clear service level objectives (SLOs)
• Define the long-term technical vision for our cloud platform, advising leadership on architectural strategy, investment priorities, and systemic risk • Set organizational standards for cloud reliability, observability, and operational excellence, including SLO/SLI frameworks, incident management practices, and the platform tooling that underpins them • Serve as the senior solutions engineering authority for partner teams: translating cross-functional requirements into platform strategy and driving prioritization of developer experience investments • Own the developer experience roadmap, identifying and closing systemic gaps in self-service infrastructure, CI/CD workflows, and operational visibility across engineering • Proactively surface systemic risks in our AWS infrastructure, IaC practices, and delivery pipelines, and drive organizational action before issues become incidents • Translate complex infrastructure risk and platform strategy into clear, business-aligned narratives for Director-level and executive stakeholders to drive resourcing and prioritization decisions • Mentor Staff (T4) and Senior (T3) engineers and model the culture of engineering rigor, documentation, and cross-functional accountability expected across the organization
Lead Deployment Engineer
MashginMashgin is working to revolutionize the checkout experience with cutting-edge technology that blends AI and 3D vision for seamless, lightning-fast transactions.
• Mentor and develop team members at all levels, providing coaching, honest feedback, and the kind of direct conversations that help people grow professionally and technically. • Become the go-to expert on Mashgin's hardware, software, and deployment systems, a reliable resource for the team, customers, and internal partners alike. • Own and improve standard operating procedures, identifying what's working, what isn't, and what needs to be built from scratch. • Serve as a clear and reliable conduit between the team and cross-functional partners including Product, Engineering, and vendors, surfacing recurring issues and keeping the right people informed. • Travel up to 25% of the time to support new location launches and assist customers across the country.



