Infrastructure Engineer
Location
United States
Posted
15 days ago
Salary
$79.5K - $100K / year
Seniority
Mid Level
Job Description
Infrastructure Engineer
Save A Lot
Role Description The Infrastructure Engineer performs a critical role in defining and implementing cloud-based solutions and strategies. This role is responsible for leading complex efforts, mentoring others, and leading level-3 support. This individual is highly effective and experienced with various development tools, IT infrastructure components, and has proven experience designing, developing, and implementing technology solutions using Agile methodologies. The candidate will be supporting Infrastructure and Operations services, with a focus on cloud technologies. - Work with business and technology partners to understand needs from the client perspective - Drive technology discussions and facilitate decision with other teams within the organization - Collaborate cross-functionally to develop our target state cloud platform and infrastructure - Gather technical requirements, assess capabilities, and analyze findings to provide appropriate cloud platform solution recommendations and adoption strategy - Design, build, and deploy efficient solutions that improve efficiencies and add value to the business - Develop creative and efficient solutions for migration to cloud environments - Design and build technology solutions, including the process, integrations, and automation - Automate testing and execution of playbook actions - Identify and execute on opportunities to automate and simplify lifecycle management - Implement best-practice standards that will be leveraged by business units to streamline and simplify cloud platform operations - Operate and maintain cloud-based solutions to ensure high performance and reliability - Define and govern operational processes, guidelines, and standards that facilitate robust, scalable, cost-effective, and innovative services - Manage risk identification within the technical architecture in partnership with security teams - Ensure alerts and events are flowing across various cloud-based services and partners - Normal business hours and some after-hours/on-call responsibilities - Support projects and initiatives to deliver on strategic goals, including building automation to improve efficiencies - Work closely with agile teams to prioritize the implementation of user stories while balancing priorities, ensuring continuous delivery of business value - Track time and complete assignments on schedule - Build, maintain, and promote strong technical documentation - Mentor and cross-train team members - Keep abreast of and gain expertise in the evolving technology and understand how new technologies could be applied to our Cloud platform offerings - Develop and maintain an operational dashboard to track and report on the health of IT systems, including alignment or impact to business services - Execute POCs and feasibility studies to validate next-gen product concepts, technologies, and services leveraging results to guide business and technology decisions - Research, test, and understand the relevant products and product capability - Participate in the vendor community on relevant products and product capability Qualifications - Bachelor's degree in information systems, or related field, OR at least 3+ years of applicable experience - Experience with Windows Server and/or Linux distributions, Windows Server 2012 or above, Linux distributions (Red Hat-based or Debian-based) - Experience developing infrastructure as code and with at least a couple of the following tools: - PowerShell - JSON - Python - Azure CLI - AWS CLI - Terraform - PowerCLI - Javascript - OAuth, webhooks, and REST API - A solid understanding of SNMP, Syslog, RMON, and other messaging and agent facilities - Experience supporting cloud-native applications in an agile manner using DevOps concepts and principles - Distributed systems design experience and integrating multiple systems using enterprise integration patterns and standard methodologies - Proficiency in implementing Group Policy - Experience maintaining Active Directory - Basic networking skills used in troubleshooting client issues Physical Requirements - Ability to travel up to ~10% of the time, which may include weekends and evenings, as needed - Most work is performed in a temperature-controlled environment - Incumbent may sit for long periods of time at a desk or computer terminal - Incumbent may use calculators, keyboards, telephone, and other office equipment in the course of a normal workday - Stooping, bending, twisting, and reaching may be required in completion of job duties Our Values - Ability to demonstrate, understand and apply our workplace values. - Simplicity (operate) – the drive to identify root cause and innovate to remove complexity to deliver the best outcome - Heart (emotion) – the passion that drives you to get up every day and work hard to strive for excellence - Performance Excellence (mindset) – clearly defining high expectations, driving ownership of key roles and responsibilities, executing with integrity and emphasis while creating a culture of accountability - Respect (philosophy) – taking pride in being inclusive and treating everyone who comes through the doors with respect Benefits - 401K company match up to 4% - Paid Time Off - Medical Insurance options including FSA & HSA - Vision Insurance - Dental insurance - Employee Assistance Programs - Team Member Referral Program - Tuition Reimbursement - Wellbeing Program - Career development opportunities
Related Guides
Related Categories
Related Job Pages
More Infrastructure Engineer Jobs
• Design, deploy, administer, and maintain Microsoft Azure cloud infrastructure in accordance with corporate architecture, cybersecurity, and compliance standards. • Administer Azure subscriptions, resource groups, virtual networks, storage accounts, virtual machines, monitoring tools, and related cloud services. • Support hub-and-spoke network architectures, hybrid connectivity, and secure integration between on-premises and cloud environments. • Monitor Azure environments for performance, reliability, availability, utilization, and cost optimization. • Assist with cloud migration, modernization, and infrastructure transformation initiatives. • Administer Microsoft Entra ID (Azure AD), Active Directory, and related identity services. • Manage user lifecycle activities including onboarding, transfers, role changes, license assignments, access reviews, and offboarding. • Support Single Sign-On (SSO), Multi-Factor Authentication (MFA), Conditional Access, identity governance, role-based access control (RBAC), and privileged access management activities. • Administer Active Directory objects, groups, organizational units, Group Policy, and hybrid identity synchronization services. • Troubleshoot authentication, directory, access, and identity synchronization issues. • Develop and maintain Infrastructure-as-Code solutions using Terraform and related deployment tools. • Create, maintain, and optimize PowerShell scripts and automation workflows to reduce manual administrative work. • Support Azure DevOps pipelines, deployment automation, configuration management, and infrastructure lifecycle activities. • Support Microsoft 365 tenant administration and integrated cloud-based services. • Assist with new employee account creation, provisioning, access updates, and technology onboarding activities. • Execute and maintain standard automation scripts, operational procedures, and infrastructure support practices. • Participate in change management, incident response, problem management, disaster recovery, and business continuity activities. • Maintain accurate infrastructure diagrams, technical documentation, operational procedures, and support records. • Support cloud-based applications and services used across the organization.
Senior Cloud Infrastructure Engineer
DragosDragos is a computer and network security company specializing in industrial cybersecurity, incident response, threat intelligence, and security software. Past
Role Description We are seeking an experienced Senior Cloud Infrastructure Engineer to join our Delivery Team. This role will design, build, and maintain the AWS and Azure infrastructure that powers the Dragos platform. You will own automation, deployment pipelines, and production reliability, bringing strong Infrastructure as Code practices to a fast-moving environment. - Design, build, and maintain scalable cloud infrastructure services in AWS and Azure - Contribute production-quality code (Go, Python, or similar) to existing cloud services - Develop and own automation and software deployment pipelines for maximum efficiency - Implement Infrastructure as Code practices using Terraform, Packer, or similar tools - Design and improve secure, reliable software update and release capabilities - Manage production release pipelines, ensuring stability and efficiency - Configure and optimize cloud networking components (firewalls, VPNs, load balancers, routing tables) - Monitor and improve service availability, uptime, and resilience - Collaborate with engineering teams on observability, logging, and security practices - Deliver production deployments and respond to incidents as required Qualifications - 3-7 years of experience working with containerized applications (Docker, Kubernetes) - Cybersecurity experience - Strong proficiency in Go and/or Python - Experience provisioning infrastructure and automating with Terraform, Packer, Ansible, or similar tools - Hands-on experience deploying and managing infrastructure in AWS and Azure - Experience with identity and access management (IAM) in AWS and/or Azure - Strong knowledge of production release pipeline and asset management concepts - Familiarity with cloud security best practices and ability to author security documentation - Experience maintaining system security plans for cloud services and dependencies - Proven track record building and maintaining cloud-based infrastructure for high availability - Strong ownership mindset and accountability for production stability - Excellent problem-solving, communication, and collaboration skills Preferred Qualifications - Experience maintaining OS package mirrors, package deployment infrastructure, and custom DEB or RPM packaging - Expertise in Linux networking and Bash scripting - Experience with secrets management tools (Hashi Corp Vault, AWS Secrets Manager, Azure Key Vault) - Experience with Postgres database management and redundancy - Experience with Makefile-driven build systems - Passion for automation and continuously improving infrastructure efficiency - Experience with modern coding languages (Rust, Go, Typescript, etc) Compensation - Salary: $165,000.00 - Competitive Equity Package - Comprehensive Benefits Plan Company Description Dragos is an Equal Opportunity Employer and considers applicants for employment without regard to race, color, religion, sex, orientation, national origin, age, disability, genetics, or any other basis forbidden under federal, state, or local laws. All new hires must pass a background check as a condition of employment.
IT Infrastructure Operations Engineer I
AstreyaIT services that put people at the center of your business
• Monitor server and network infrastructure health using established monitoring tools and dashboards, identifying alerts and anomalies requiring attention. • Provide first-level response to service tickets within the organization's ticketing system, ensuring 2-hour initial response SLA compliance for all incoming requests. • Perform basic troubleshooting of Dell PowerEdge server hardware issues using iDRAC interfaces, RAID, escalating complex problems to L2 support with detailed documentation. • Execute routine health checks on Cisco routers and switches, documenting status and flagging any deviations from normal operating parameters. • Follow standardized runbooks and operational procedures for common infrastructure issues, ensuring consistent resolution approaches across the team. • Log and document all incidents, actions taken, and resolutions in the ticketing system with accurate and detailed information for knowledge management. • Escalate unresolved or complex issues to L2/L3 engineers with comprehensive handover notes including symptoms, actions attempted, and relevant logs. • Assist with routine maintenance activities including scheduled reboots, basic configuration backups, and pre-approved firmware update executions under supervision. • Participate in shift handover meetings, providing clear status updates on open tickets and ongoing issues to incoming team members. • Coordinate with on-site technicians and vendors for basic hardware replacement activities, ensuring proper ticketing and tracking of all dispatch requests. • Maintain awareness of scheduled maintenance windows and change activities, monitoring for any unexpected impacts during and after implementation. • Contribute to the continuous improvement of runbooks and documentation by identifying gaps and suggesting updates based on real-world incident handling experience.
IT Infrastructure Operations Engineer II
AstreyaIT services that put people at the center of your business
Role Description We are looking for an experienced L2 IT Infrastructure Operations Engineer to provide advanced technical support for our enterprise server and network infrastructure. This mid-level position bridges the gap between frontline support and expert-level engineering, handling escalated incidents, performing complex troubleshooting, and contributing to operational excellence. The ideal candidate will possess hands-on experience with Dell PowerEdge servers, Cisco networking equipment, and enterprise monitoring solutions. - Mentor L1 engineers. - Participate in change management activities. - Collaborate with cross-functional teams to ensure high availability and performance of critical infrastructure in a 24x7 global environment. Qualifications - 5+ years of hands-on experience in enterprise IT infrastructure operations. - Strong proficiency with Dell PowerEdge server administration, including hardware troubleshooting, iDRAC/Redfish management, and firmware lifecycle management. - Solid experience with Cisco networking equipment (routers, switches), including IOS/NX-OS configuration, troubleshooting, and upgrade procedures. - Working knowledge of monitoring and logging tools, with ability to create dashboards, configure alerts, and analyze performance metrics for proactive issue detection. - Excellent problem-solving abilities with demonstrated experience in incident management, root cause analysis, and implementing corrective actions in production environments. - Industry certifications such as Dell Server certifications or ITIL Foundation. - Ability to work rotating shifts in a 24x7 global support model. Requirements - Provide advanced troubleshooting and fault isolation for escalated server and network incidents, utilizing iDRAC, Redfish, and Cisco CLI tools to diagnose and resolve complex issues. - Execute firmware, BIOS, and driver updates on Dell PowerEdge servers following standardized procedures, ensuring minimal service disruption and maintaining system stability. - Perform IOS/NX-OS firmware and software updates on Cisco routers and switches, adhering to change management protocols and conducting post-update validation. - Manage hardware break/fix procedures for server infrastructure, coordinating with Dell support for warranty claims, parts ordering, and scheduling on-site technician dispatch. - Conduct regular network health audits and performance analysis, identifying potential bottlenecks and recommending optimization measures to prevent service degradation. - Collaborate with the SRE team to enhance monitoring dashboards and refine alerting thresholds, ensuring proactive detection of infrastructure instability or security events. - Mentor and provide technical guidance to L1 engineers, conducting knowledge transfer sessions and assisting with complex ticket resolution to build team capability. - Participate in blameless post-mortems following major incidents, contributing to root cause analysis and implementing preventative actions to improve system reliability. - Maintain and update operational runbooks, network diagrams, and technical documentation to reflect current configurations and best practices. - Support hardware lifecycle management activities including equipment provisioning, asset tracking, and coordination with vendors for hardware returns and repairs. - Provide 24x7 on-call support for critical escalations, ensuring rapid response to high-priority incidents affecting production systems. - Collaborate with the FTE IT Team Lead on capacity planning activities, providing data-driven insights on infrastructure utilization trends and growth projections. Tools Required - Server & Hardware Tools: Dell iDRAC, Lifecycle Controller, OpenManage, RAID/PERC utilities for server provisioning, firmware baselining, and remote management. - OS Deployment Tools: PXE boot infrastructure, iDRAC Virtual Media, Windows Server & Linux ISOs with hardening and automation scripts. - Network Tools: Cisco IOS CLI, PoE management, VLAN/QoS configuration tools, network monitoring, and bandwidth/latency testing utilities. - Automation & Operations Tools: Ansible, Python, CMDB systems, configuration backup tools, and documentation/diagramming platforms for global 24x7 operations.


