ACI Infotech

ACI Infotech | Global IT and Business Transformation Partner | www.aciinfotech.com Consulting Industry Solutions Insights & Analytics Digital Cloud Business Transformation Security Engage. Innovate. Experience.

Site Reliability engineer (SRE)

Location

United States

Posted

2 days ago

Salary

$65 - $70 / hour

Seniority

Mid Level

No structured requirement data.

Job Description

Site Reliability engineer (SRE)

ACI Infotech

Role Description - Design and manage multi-account AWS infrastructure (VPC, Route Tables, EC2, ECS, EKS 1.33, RDS, DynamoDB, Elasticache, S3, Transit Gateway, Resource Access Manager, Lambda, CloudFormation, AWS Backup) - Configure load balancing and traffic management (ELB, NLB, Target Groups with gRPC, Route53, Global Accelerator, CloudFront) - Implement security and compliance controls (IAM, IAM Identity Center, SCP, Guard Duty, WAF, CloudTrail, ACM, Secrets Manager, OKTA integration) - Manage Cloudflare infrastructure (Zero Trust, Argo Smart Routing, DNS, Workers, Load Balancer, Bot Management, WAF, Rules & Policies, Cache) - Manage S3 with Access Policies, Lifecycle Policies, S3 Storage Lens optimization, and cross-region replication - Operate messaging and notification services (SNS, SES, SQS) - Architect and manage multi-cluster EKS environments with HA and cross-region DR scenarios using Istio service mesh, Network Policies, Karpenter, HPA, KEDA, Argo CD, Argo Rollouts - Implement and maintain Argo CD for multi-cluster application management with HA and cross-region DR configurations - Configure Argo CD Application Sets for managing applications across multiple EKS clusters - Implement ECR with global cross-region replication for container image distribution and disaster recovery - Implement Aurora Global Database for cross-region DR, manage Aurora RDS (MySQL and PostgreSQL) and standalone MySQL/PostgreSQL instances for development - Design and maintain RDS cross-region replication, automated backups, failover strategies, and upgrade procedures - Establish and maintain DevOps practices including change management, release management and deployment strategies - Build resilient CI/CD pipelines with cross-region artifact replication, automated testing, and failover capabilities - Develop and maintain GitHub Actions shared internal workflows and reusable actions for standardized deployments - Implement change approval workflows, deployment gates, and release coordination processes - Implement Crossplane for automated feature environment creation, upgrades, and AWS resource provisioning - Deploy applications using Helm, Customize with Overlay Patches, Jsonnet, and Crossplane for infrastructure orchestration - Maintain platform operators (External DNS, External Secrets, Reloader) and custom CRDs - Build comprehensive observability stack & Dashboards (Grafana, Thanos/Prometheus, Loki, Alert manager, Open Telemetry Alloy/Tempo/Beyla/Pyro scope) - Configure exporters (Blackbox, MySQL, Redis, YACE CloudWatch, Cloudflare, Node Exporter, Prometheus Push Gateway) - Support data platforms (Kafka/Kafka UI, Minion, Airflow, JupyterHub, DASK, Superset, Imply, AWS Glue, Athena, Quick Sight, Bedrock) - Optimize CI/CD with GitHub Actions, Actions Runner Controller (ARC), runs-on.com, GitHub Rulesets - Manage mobile app delivery pipelines (Unity Build Management, Fastlane, Google Play Developer, Apple Developer/Enterprise, Applivery) - Implement and maintain all infrastructure using Terraform/Open Tofu with Scalr, backporting existing resources into code - Automate operational tasks wherever possible; create comprehensive runbooks for non-automatable procedures - Conduct thorough post-mortem analysis after incidents, documenting learnings and implementing preventive measures - Drive cost optimization initiatives using S3 Storage Lens, CloudWatch metrics, rightsizing recommendations, and resource lifecycle management - Develop automation in Bash, Python, Go, C#/.NET (Unity Game Engine) - Maintain developer experience (Backstage, Click Up, Miro, Shared GitHub Action/Workflows) - Integrate monitoring and alerting (PagerDuty, Cronitor, Wiz, CloudWatch) Qualifications - Multi-account AWS architecture with Transit Gateway, Resource Access Manager, VPC design, and Route Tables - Kubernetes/EKS high availability with cross-region disaster recovery scenarios - Multi-cluster EKS management with service mesh (Istio), autoscaling (Karpenter, KEDA), GitOps (Argo CD) - Argo CD enterprise deployment for multi-cluster application management with HA and cross-region DR - Argo CD Application Sets, app-of-apps patterns with Helm, and cluster management strategies - ECR global cross-region replication strategies for container image distribution and DR - Cloudflare enterprise features (Zero Trust, Argo Smart Routing, DNS management, Workers, Load Balancer, Bot Management, Cache optimization, WAF Rules & Other Security Policies) - Aurora Global Database implementation and management for cross-region DR - Aurora RDS (MySQL and PostgreSQL engines) and standalone MySQL/PostgreSQL instance management - RDS cross-region replication, automated failover, disaster recovery, and version upgrade strategies - DevOps best practices including change management, release management, and deployment coordination - Resilient CI/CD pipelines with automated testing, cross-region artifact distribution, and failover - GitHub Actions shared workflows and reusable actions development for internal use - Crossplane for Kubernetes-native infrastructure provisioning, feature environment automation, and upgrade orchestration - Expert-level Terraform/Open Tofu with enterprise policy management (Scalr) - Infrastructure backporting and migration from ClickOps to IaC - Complete observability stack (Prometheus, Grafana, Loki, Open Telemetry, distributed tracing) - Data pipeline orchestration (Kafka, Airflow) and analytics platforms (Superset, Imply) - GitHub Actions with self-hosted runners (ARC, runs-on.com) - Proficiency in Python, Bash, Go, and C#/.NET for automation development - Security implementations (IAM, SCP, OKTA, WAF, Guard Duty, Wiz) - Mobile CI/CD (Unity, Fastlane, Apple/Google distribution & Applivery during Development) - Disaster recovery planning, testing, and automation (AWS Backup, cross-region strategies) - AI/ML infrastructure experience (AWS Bedrock) - Cost optimization strategies and Quick Sight for AWS Cost Review - Post-mortem facilitation and blameless incident analysis - Runbook creation and maintenance for operational procedures Technical Skills - Container orchestration with advanced networking and progressive delivery - Infrastructure as Code and GitOps methodologies with automation-first mindset - Change management workflows, approval gates, and release orchestration - CI/CD pipeline design with automated testing, security scanning, and deployment strategies - Incident response, on-call management, post-mortem analysis, DR execution - Crossplane composition design and custom resource definitions - Custom CRD and operator development in Kubernetes - Event-driven architecture (Lambda, SQS, SNS, SES) - Real-time analytics and BI platforms - Developer portal management (Backstage) - Multi-region failover automation and orchestration - Cost analysis and optimization using native AWS tools - Automation of repetitive operational tasks - Technical documentation and runbook authoring - Database performance tuning and optimization (Aurora, MySQL, PostgreSQL) - Argo CD backup, restore, and disaster recovery procedures - Cloudflare Workers development & deployment using Wrangler Soft Skills - Strong troubleshooting - Cross-functional communication - Self-directed - Documentation-focused - Cost-conscious - Continuous improvement mindset

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Vena Solutions logo

Site Reliability Developer

Vena Solutions

Take your entire business from reactive to proactive with the leading AI-Powered Complete FP&A Platform.

DevOps Engineer2 days ago
Full TimeRemoteTeam 501-1,000Since 2011H1B No Sponsor

• Support key ITIL processes, including Incident management, request management, problem management and change management. • Define and document runbooks and standard operating procedures. • Field operational requests from our Application Support team and other internal stakeholders • Triage and solve issues within defined SLA’s to ensure an excellent customer experience and to unblock other development and support teams • Maintain services once they are live by measuring and monitoring availability, latency and overall system health. • Identify and troubleshoot problems, investigate root causes, and champion fixes across the organization. • Work with infrastructure-as-Code (IaC) with a focus on continuous improvement. • Collaborate with cross-functional team members on features and implementation within an agile environment. • Report on SLAs and performance metrics as part of the Operations function. • Participate in on-call rotation.

Canada
$85K - $115K / year
Full TimeRemoteTeam 11-50H1B No Sponsor

• proaktywne dostrzeganie wyzwań i ich adresowanie wspólnie z zespołami backendowymi i devops • tworzenie szablonów i standardów infrastrukturalnych dla najczęstszych potrzeb zespołów developerskich (np. nowe serwisy, projekty Cloud Run) • rozwój i utrzymanie stacku observability (Loki, Grafana, Prometheus) oraz optymalizacja wykorzystania zasobów i kosztów chmury (GCP) • utrzymanie i rozwój wewnętrznych narzędzi self-hosted (m.in. NocoDB, n8n, Outline) • wsparcie infrastrukturalne zespołów Data Engineering i Data Science (m.in. Airflow, pipeline'y forecastingowe czy serwery pod ML) • zarządzanie infrastrukturą sieciową w darkstore'ach (sieć, CCTV, drukarki paragonów, urządzenia handheld) we współpracy z IT Managerem • budowanie i utrzymanie infrastruktury on-prem (serwery pod ML i workloady wymagające dużych zasobów, runnery CI/CD dla projektów mobilnych) • utrzymanie i usprawnianie procesów CI/CD (GitLab) oraz developer experience • budowanie infrastruktury pod rozwiązania AI (autonomiczne agenty, boty Slack, agenty code review, remote coding agents) • rozwój praktyk SRE: alerting, procesy zgłaszania i obsługi incydentów (PagerDuty, Slack)

Poland
zł23K - zł27K / month
MKS2 Technologies logo

Senior DevOps Engineer

MKS2 Technologies

Austin-based SDVOSB delivering application development, cybersecurity, instructional design and training to DOD and VA.

DevOps Engineer2 days ago
Full TimeRemoteTeam 201-500Since 2008H1B No Sponsor

• Conduct DevOps and DevSecOps activities for an Azure integration platform • Create and maintain GitHub CI/CD pipelines • Conduct Azure DevOps activities • Conduct Azure system administration • Build platform automation • Create and maintain PowerShell scripting • Understanding and provisioning of Azure Infrastructure services and network • Create and maintain containers and infrastructure (IaC, Terraform) • Deploy solutions through CI/CD pipelines • Embed DevOps best practices and process/tech improvements across the team and platform • Leverage AI tools to accelerate activities • Collaborate across technical teams • Create and maintain detailed technical documentation • Support troubleshooting and resolution of Production issues • Participate in Agile ceremonies • Report on progress and status of development

United States
$130K - $140K / year
Fiserv logo

Senior DevOps Engineer

Fiserv

Founded in 1984, Fiserv is a global provider of ecommerce and information management systems for the financial services industry. In 2013, Fiserv acquired Open

DevOps Engineer2 days ago

• Design, implement, and support secure Azure IaaS and PaaS solutions across enterprise environments • Build and maintain CI/CD pipelines using Azure DevOps and GitLab to enable efficient, automated software delivery • Develop and manage Infrastructure as Code solutions using Terraform and Ansible for provisioning and configuration management • Implement and maintain Zero Trust security architectures, including identity and access management (IAM), data encryption, and security controls • Identify, assess, and remediate infrastructure and application security vulnerabilities • Create and maintain monitoring, alerting, and operational dashboards using Azure Monitor, Log Analytics, and Application Insights • Troubleshoot production issues, restore services, and maintain operational runbooks and knowledgebase documentation • Collaborate with cross-functional teams to support cloud modernization initiatives

Maryland
$109K - $152K / year