Dragos logo
Dragos

Dragos is a computer and network security company specializing in industrial cybersecurity, incident response, threat intelligence, and security software. Past

Senior Cloud Infrastructure Engineer

Infrastructure EngineerInfrastructure EngineerFull TimeRemoteSeniorTeam 295Since 2016Company Site

Location

Northern America + 3 moreAll locations: Northern America | Europe | Asia Pacific | Western Asia (Middle East)

Posted

1 day ago

Salary

$165K / year

Seniority

Senior

Job Description

Senior Cloud Infrastructure Engineer

Dragos

Role Description We are seeking an experienced Senior Cloud Infrastructure Engineer to join our Delivery Team. This role will design, build, and maintain the AWS and Azure infrastructure that powers the Dragos platform. You will own automation, deployment pipelines, and production reliability, bringing strong Infrastructure as Code practices to a fast-moving environment. - Design, build, and maintain scalable cloud infrastructure services in AWS and Azure - Contribute production-quality code (Go, Python, or similar) to existing cloud services - Develop and own automation and software deployment pipelines for maximum efficiency - Implement Infrastructure as Code practices using Terraform, Packer, or similar tools - Design and improve secure, reliable software update and release capabilities - Manage production release pipelines, ensuring stability and efficiency - Configure and optimize cloud networking components (firewalls, VPNs, load balancers, routing tables) - Monitor and improve service availability, uptime, and resilience - Collaborate with engineering teams on observability, logging, and security practices - Deliver production deployments and respond to incidents as required Qualifications - 3-7 years of experience working with containerized applications (Docker, Kubernetes) - Cybersecurity experience - Strong proficiency in Go and/or Python - Experience provisioning infrastructure and automating with Terraform, Packer, Ansible, or similar tools - Hands-on experience deploying and managing infrastructure in AWS and Azure - Experience with identity and access management (IAM) in AWS and/or Azure - Strong knowledge of production release pipeline and asset management concepts - Familiarity with cloud security best practices and ability to author security documentation - Experience maintaining system security plans for cloud services and dependencies - Proven track record building and maintaining cloud-based infrastructure for high availability - Strong ownership mindset and accountability for production stability - Excellent problem-solving, communication, and collaboration skills Preferred Qualifications - Experience maintaining OS package mirrors, package deployment infrastructure, and custom DEB or RPM packaging - Expertise in Linux networking and Bash scripting - Experience with secrets management tools (Hashi Corp Vault, AWS Secrets Manager, Azure Key Vault) - Experience with Postgres database management and redundancy - Experience with Makefile-driven build systems - Passion for automation and continuously improving infrastructure efficiency - Experience with modern coding languages (Rust, Go, Typescript, etc) Compensation - Salary: $165,000.00 - Competitive Equity Package - Comprehensive Benefits Plan Company Description Dragos is an Equal Opportunity Employer and considers applicants for employment without regard to race, color, religion, sex, orientation, national origin, age, disability, genetics, or any other basis forbidden under federal, state, or local laws. All new hires must pass a background check as a condition of employment.

Related Categories

Related Job Pages

More Infrastructure Engineer Jobs

Astreya logo

IT Infrastructure Operations Engineer I

Astreya

IT services that put people at the center of your business

Full TimeRemoteTeam 1,001-5,000Since 2001H1B Sponsor

• Monitor server and network infrastructure health using established monitoring tools and dashboards, identifying alerts and anomalies requiring attention. • Provide first-level response to service tickets within the organization's ticketing system, ensuring 2-hour initial response SLA compliance for all incoming requests. • Perform basic troubleshooting of Dell PowerEdge server hardware issues using iDRAC interfaces, RAID, escalating complex problems to L2 support with detailed documentation. • Execute routine health checks on Cisco routers and switches, documenting status and flagging any deviations from normal operating parameters. • Follow standardized runbooks and operational procedures for common infrastructure issues, ensuring consistent resolution approaches across the team. • Log and document all incidents, actions taken, and resolutions in the ticketing system with accurate and detailed information for knowledge management. • Escalate unresolved or complex issues to L2/L3 engineers with comprehensive handover notes including symptoms, actions attempted, and relevant logs. • Assist with routine maintenance activities including scheduled reboots, basic configuration backups, and pre-approved firmware update executions under supervision. • Participate in shift handover meetings, providing clear status updates on open tickets and ongoing issues to incoming team members. • Coordinate with on-site technicians and vendors for basic hardware replacement activities, ensuring proper ticketing and tracking of all dispatch requests. • Maintain awareness of scheduled maintenance windows and change activities, monitoring for any unexpected impacts during and after implementation. • Contribute to the continuous improvement of runbooks and documentation by identifying gaps and suggesting updates based on real-world incident handling experience.

India
Astreya logo

IT Infrastructure Operations Engineer II

Astreya

IT services that put people at the center of your business

Full TimeRemoteTeam 1,001-5,000Since 2001H1B Sponsor

Role Description We are looking for an experienced L2 IT Infrastructure Operations Engineer to provide advanced technical support for our enterprise server and network infrastructure. This mid-level position bridges the gap between frontline support and expert-level engineering, handling escalated incidents, performing complex troubleshooting, and contributing to operational excellence. The ideal candidate will possess hands-on experience with Dell PowerEdge servers, Cisco networking equipment, and enterprise monitoring solutions. - Mentor L1 engineers. - Participate in change management activities. - Collaborate with cross-functional teams to ensure high availability and performance of critical infrastructure in a 24x7 global environment. Qualifications - 5+ years of hands-on experience in enterprise IT infrastructure operations. - Strong proficiency with Dell PowerEdge server administration, including hardware troubleshooting, iDRAC/Redfish management, and firmware lifecycle management. - Solid experience with Cisco networking equipment (routers, switches), including IOS/NX-OS configuration, troubleshooting, and upgrade procedures. - Working knowledge of monitoring and logging tools, with ability to create dashboards, configure alerts, and analyze performance metrics for proactive issue detection. - Excellent problem-solving abilities with demonstrated experience in incident management, root cause analysis, and implementing corrective actions in production environments. - Industry certifications such as Dell Server certifications or ITIL Foundation. - Ability to work rotating shifts in a 24x7 global support model. Requirements - Provide advanced troubleshooting and fault isolation for escalated server and network incidents, utilizing iDRAC, Redfish, and Cisco CLI tools to diagnose and resolve complex issues. - Execute firmware, BIOS, and driver updates on Dell PowerEdge servers following standardized procedures, ensuring minimal service disruption and maintaining system stability. - Perform IOS/NX-OS firmware and software updates on Cisco routers and switches, adhering to change management protocols and conducting post-update validation. - Manage hardware break/fix procedures for server infrastructure, coordinating with Dell support for warranty claims, parts ordering, and scheduling on-site technician dispatch. - Conduct regular network health audits and performance analysis, identifying potential bottlenecks and recommending optimization measures to prevent service degradation. - Collaborate with the SRE team to enhance monitoring dashboards and refine alerting thresholds, ensuring proactive detection of infrastructure instability or security events. - Mentor and provide technical guidance to L1 engineers, conducting knowledge transfer sessions and assisting with complex ticket resolution to build team capability. - Participate in blameless post-mortems following major incidents, contributing to root cause analysis and implementing preventative actions to improve system reliability. - Maintain and update operational runbooks, network diagrams, and technical documentation to reflect current configurations and best practices. - Support hardware lifecycle management activities including equipment provisioning, asset tracking, and coordination with vendors for hardware returns and repairs. - Provide 24x7 on-call support for critical escalations, ensuring rapid response to high-priority incidents affecting production systems. - Collaborate with the FTE IT Team Lead on capacity planning activities, providing data-driven insights on infrastructure utilization trends and growth projections. Tools Required - Server & Hardware Tools: Dell iDRAC, Lifecycle Controller, OpenManage, RAID/PERC utilities for server provisioning, firmware baselining, and remote management. - OS Deployment Tools: PXE boot infrastructure, iDRAC Virtual Media, Windows Server & Linux ISOs with hardening and automation scripts. - Network Tools: Cisco IOS CLI, PoE management, VLAN/QoS configuration tools, network monitoring, and bandwidth/latency testing utilities. - Automation & Operations Tools: Ansible, Python, CMDB systems, configuration backup tools, and documentation/diagramming platforms for global 24x7 operations.

India
US Anesthesia Partners logo

Infrastructure Engineer

US Anesthesia Partners

Quality Anesthesia Care: We're raising the bar for the industry.

Full TimeRemoteTeam 5,001-10,000Since 2012H1B No Sponsor

• Leads and participates in IT infrastructure projects, including the design, deployment, and maintenance of server systems, cloud services, and virtualization platforms. • Ensures projects are completed on time and within budget. • Oversees the deployment of IT infrastructure components, including Windows, Linux, and virtualized environments. • Implements system updates, patches, and upgrades to ensure optimal performance and security. • Works closely with the architecture team to design and implement scalable and secure IT infrastructure solutions. • Ensures alignment with architectural standards and best practices. • Develops and implements backup and disaster recovery plans to ensure data integrity and business continuity. • Scripts in PowerShell for automation of routine tasks and configuration management. • Understands coding and scripting languages to support infrastructure automation. • Implements and manages infrastructure using IaC principles and tools like Terraform to automate deployment and configuration processes. • Participates in Agile workflow methodologies, including sprint planning, daily stand-ups, and retrospectives to ensure efficient project delivery. • Setups monitoring systems and tools to ensure the performance, availability, and reliability of our IT infrastructure. • Responds to IT infrastructure-related issues and provides timely resolution to minimize downtime. • Collaborates with other IT teams and vendors to troubleshoot complex problems and implement solutions. • Maintains accurate documentation of IT infrastructure configurations, processes, and procedures. • Generates regular reports on system performance, capacity, and security metrics.

Texas
$73.6K - $125.1K / year

AI Infrastructure Engineer

Bright Vision Technologies

Bright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.

Role Description We are seeking an AI Performance Optimization Engineer to focus on extracting maximum throughput, minimizing latency, and reducing cost across training and inference workloads for large neural network systems. The role spans the full stack from low-level kernel optimization to distributed system tuning, requiring deep understanding of GPU architecture, model parallelism, memory management, and compiler-level optimization. The ideal candidate has demonstrated an impact on production of AI workloads, with strong instrumentation and measurement discipline that enables rigorous, data-driven optimization decisions. In this role you will work closely with cross-functional partners — product, design, engineering, operations, and business stakeholders — to translate ambiguous requirements into well-engineered solutions, and will be expected to raise the bar through code review, design review, and mentorship of more junior engineers. The successful candidate brings strong engineering discipline, a clear communication style, and a track record of shipping meaningful work that holds up well in production. Key Responsibilities - Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, and cost. - Identify and eliminate bottlenecks across data loading, model compute, communication, and memory. - Implement and tune quantization, sparsity, and pruning strategies to reduce model footprint and accelerate inference. - Optimize distributed training using tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding. - Tune attention implementations using Flash Attention, paged attention, and related techniques. - Implement KV cache optimization, continuous batching, and speculative decoding for LLM serving. - Drive compiler-level optimizations using Triton, XLA, Torch Inductor, or TVM, working with the broader ML framework community to land improvements that translate into measurable end-to-end performance gains. - Optimize data pipelines, sharding strategies, and storage access patterns for high-throughput training. - Build and maintain rigorous benchmark suites and regression frameworks across workloads. - Collaborate with ML and platform engineering teams to embed best practices in standard pipelines. - Drive cost-efficiency improvements through model architecture, hardware selection, and scheduling strategies. - Evaluate new hardware and software offerings and advise on adoption. - Document performance tuning playbooks and share findings broadly across engineering teams. - Stay current with AI systems to research and translate advances into production improvements. Qualifications - Bachelor's or master's degree in computer science, Computer Engineering, or related field. - Six or more years of experience in performance engineering, ML systems, or HPC. - Strong proficiency in Python and C++. - Hands-on experience optimizing deep learning workloads on modern GPUs. - Deep understanding of distributed training and inference techniques. - Experience with profiling tools across CPU, GPU, and distributed systems. - Familiarity with model compression techniques and their accuracy implications. - Strong grasp of memory hierarchies, communication primitives, and parallelism strategies. - Excellent measurement, debugging, and analytical reasoning skills. - Strong communication and collaboration skills. Preferred Qualifications - Experience optimizing LLM inference at production scale. - Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects. - Familiarity with custom kernel authoring in Triton or CUTLASS. - Experience with FinOps for AI workloads. - Publications or talks on AI systems performance. How to Apply Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 650-6699. Learn more about Bright Vision Technologies at www.bvteck.com . Equal Employment Opportunity (EEO) Statement Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall. BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.

United States
$100K - $150K / year