itD Tech

About itD: We are part of a new generation of consulting and software development company that blends diversity, innovation, and integrity with real business results. Our structure rejects any strong hierarchy, empowering us to deliver excellent results. We are a woman- and minority-led firm. Every day, we challenge ourselves to be considerate, fair and to re-think what great outcomes mean for our customers. This permeates down to how we approach every interaction, on every project, for every client. You’ll thrive here if you are a dynamic self-starter, a difference-maker or someone who wants to deliver great results, without constraints. The itD Digital Experience: Joining us means you’ll be part of our global community, you have a say about your own career journey, and you’ll get a chance to give back to causes that matter. You will experience working with Fortune 500 companies and high-performance teams across numerous industries. itD offers our employees excellent benefits such as medical, dental, vision, life insurance, paid holidays, 401K + matching.

Site Reliability Engineer

Location

United States

Posted

16 hours ago

Salary

0

Seniority

Mid Level

No structured requirement data.

Job Description

Site Reliability Engineer

itD Tech

Role Description itD is seeking a Site Reliability Engineer to develop and enhance automation solutions that improve the reliability, scalability, and operational efficiency of large-scale cloud infrastructure. The ideal candidate will bring hands-on experience in site reliability engineering, infrastructure automation, and cloud operations, with a proven track record of delivering reliable automation, scalable deployment pipelines, and resilient production environments. Location: 100% Remote within the United States. Duration: 24 months Responsibilities - Develop and maintain infrastructure automation solutions using Ansible to improve the reliability, scalability, and operational efficiency of cloud environments. - Design, implement, and enhance CI/CD pipelines, testing frameworks, and operational tooling to support infrastructure growth and software delivery. - Troubleshoot complex Linux-based infrastructure and distributed systems issues to maintain high availability and platform performance. - Build automation that enables the rapid, repeatable deployment of regional, sovereign, and purpose-built cloud environments. - Collaborate with engineering teams, product management, and cross-functional stakeholders to identify opportunities for operational improvements and increased platform reliability. - Monitor infrastructure performance and implement enhancements that reduce operational overhead and improve system scalability. - Contribute to automation and engineering best practices that support reliable, efficient cloud platform operations. Internal Responsibilities - Attend regular internal practice community meetings. - Collaborate with your itD practice team on industry thought leadership. - Complete client case studies and learning material (blogs, media material). - Build out material to contribute to the Digital Transformation practice. - Attend internal itD networking events (in person and virtual). - Work with leadership on career fast-track opportunities. Qualifications - 2+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a related role supporting cloud-based production environments. - Experience developing and maintaining infrastructure automation using Ansible. - Experience programming in Ruby and developing automated tests using RSpec or comparable testing frameworks. - Experience administering and troubleshooting Linux-based systems and distributed infrastructure environments. - Experience designing, implementing, and maintaining CI/CD pipelines, including GitLab CI. - Experience supporting large-scale infrastructure environments consisting of hundreds or thousands of systems. - Must be eligible to work on FedRAMP projects. - Must be a U.S. citizen working from U.S. soil. Preferred Qualifications and Skills - Experience with AWS or other public cloud platforms and hybrid infrastructure environments. - Knowledge of monitoring, observability, and site reliability engineering practices and tools. - Familiarity with Kubernetes concepts and containerized application platforms. - Experience leveraging AI-assisted development tools to improve software development, infrastructure automation, operational analysis, and engineering productivity. Education - Bachelor's degree in a relevant field or equivalent work experience required. Benefits - Comprehensive medical benefits. - 401(k) plan. - Paid holidays. - Networking & career learning and development programs. Company Description About itD: We are part of a new generation of consulting and software development company that blends diversity, innovation, and integrity with real business results. Our structure rejects any strong hierarchy, empowering us to deliver excellent results. We are a woman- and minority-led firm. Every day, we challenge ourselves to be considerate, fair and to re-think what great outcomes mean for our customers. The itD Digital Experience: Joining us means you’ll be part of our global community, you have a say about your own career journey, and you’ll get a chance to give back to causes that matter. You will experience working with Fortune 500 companies and high-performance teams across numerous industries.

Related Categories

Related Job Pages

More DevOps Engineer Jobs

BeReal. logo

Senior SRE

BeReal.

Your friends for real.

DevOps Engineer16 hours ago
Full TimeRemoteTeam 51-200Since 2020

• Apply and help maintain SRE practices across your scope, including SLIs, SLOs, error budgets, incident management, and postmortem processes • Design, implement, and optimize infrastructure for availability, scalability, reliability, and cost efficiency • Contribute to our observability stack, improving monitoring, alerting, logging, and distributed tracing • Automate infrastructure and operational workflows (e.g., Terraform, Terragrunt, Kubernetes) • Support FinOps initiatives, helping build tools and insights to optimize cloud costs • Partner closely with development squads to improve service reliability, performance, and operational excellence • Contribute to architectural decisions and help apply best practices for building resilient distributed systems • Share knowledge with other Infrastructure engineers, helping raise the bar on reliability and operational excellence • Analyze performance bottlenecks and work on solutions such as scaling strategies, service optimizations, and system debugging

France
Full TimeRemoteTeam 1,001-5,000H1B No Sponsor

• Develop and maintain Infrastructure as Code (IaC) using Terraform • Create and manage configuration and provisioning automations with Ansible • Design, deploy, and manage cloud and hybrid infrastructure environments • Develop reusable, standardized modules for resource provisioning • Integrate infrastructure automations into CI/CD pipelines • Perform configuration management, version control, and code reviews using Git • Ensure compliance with security standards, governance, and best practices • Monitor, optimize, and troubleshoot incidents in infrastructure environments • Automate operational tasks to increase efficiency and reduce manual work • Automate deployment, configuration, and update processes for servers and applications • Collaborate with DevOps, Cloud, Development, and Security teams • Document architectures, processes, and operational procedures • Participate in the continuous evolution of the platform, promoting scalability, availability, and reliability of environments • Implement governance, compliance, and access control best practices • Participate in technical meetings and discussions

Brazil
Outpost logo

Site Reliability Engineer

Outpost

The Backbone of Your Freight Network

DevOps Engineer17 hours ago
ContractRemoteTeam 51-200Since 2021H1B No Sponsor

• Own reliability targets across our backend/API, worker services, applications and CV pipeline; MTD, MTM, MTR, and follow-through on root causes. • Level up our monitoring and alerting, and build out auto-remediation, so on-call load scales with automation, not headcount. • Partner with our agentic engineering work to build agents that triage alerts and handle routine remediation. • Harden and optimize our GCP infrastructure (Cloud Run, Cloud SQL, GCS) for cost and performance as load scales. • Own database scale and performance; connection pooling, query optimization and indexing, read replicas, and capacity planning, so Postgres doesn't become the bottleneck as data volume grows. • Improve the reliability of our ML training and monitoring infrastructure, in partnership with the CV/ML team. • Run blameless postmortems and drive fixes for root causes, not just symptoms. • Participate in on-call rotation.

United States
Full TimeRemoteTeam 10,001+Since 1936H1B Sponsor

• Support all US Fresh and Packaged Meat Facilities as required • Work with facilities teams and project engineering on Capital Infrastructure Plans • Engage professional refrigeration engineering resources during project design and development • Assist in compliance-related issues as requested • Develop designs including selecting equipment and obtaining quotes • Implement resolutions to issues related to facility refrigeration systems • Assist with planning and budgeting for utilities engineering improvement projects • Research and test new technology for system improvements • Identify and correct deficiencies within existing systems

Kentucky + 2 moreAll locations: Kentucky | North Carolina | Virginia
$85K - $120K / year