IT services that put people at the center of your business
IT Infrastructure Operations Engineer II
Location
India
Posted
5 days ago
Salary
0
Seniority
Mid Level
Job Description
IT Infrastructure Operations Engineer II
Astreya
Role Description We are looking for an experienced L2 IT Infrastructure Operations Engineer to provide advanced technical support for our enterprise server and network infrastructure. This mid-level position bridges the gap between frontline support and expert-level engineering, handling escalated incidents, performing complex troubleshooting, and contributing to operational excellence. The ideal candidate will possess hands-on experience with Dell PowerEdge servers, Cisco networking equipment, and enterprise monitoring solutions. - Mentor L1 engineers. - Participate in change management activities. - Collaborate with cross-functional teams to ensure high availability and performance of critical infrastructure in a 24x7 global environment. Qualifications - 5+ years of hands-on experience in enterprise IT infrastructure operations. - Strong proficiency with Dell PowerEdge server administration, including hardware troubleshooting, iDRAC/Redfish management, and firmware lifecycle management. - Solid experience with Cisco networking equipment (routers, switches), including IOS/NX-OS configuration, troubleshooting, and upgrade procedures. - Working knowledge of monitoring and logging tools, with ability to create dashboards, configure alerts, and analyze performance metrics for proactive issue detection. - Excellent problem-solving abilities with demonstrated experience in incident management, root cause analysis, and implementing corrective actions in production environments. - Industry certifications such as Dell Server certifications or ITIL Foundation. - Ability to work rotating shifts in a 24x7 global support model. Requirements - Provide advanced troubleshooting and fault isolation for escalated server and network incidents, utilizing iDRAC, Redfish, and Cisco CLI tools to diagnose and resolve complex issues. - Execute firmware, BIOS, and driver updates on Dell PowerEdge servers following standardized procedures, ensuring minimal service disruption and maintaining system stability. - Perform IOS/NX-OS firmware and software updates on Cisco routers and switches, adhering to change management protocols and conducting post-update validation. - Manage hardware break/fix procedures for server infrastructure, coordinating with Dell support for warranty claims, parts ordering, and scheduling on-site technician dispatch. - Conduct regular network health audits and performance analysis, identifying potential bottlenecks and recommending optimization measures to prevent service degradation. - Collaborate with the SRE team to enhance monitoring dashboards and refine alerting thresholds, ensuring proactive detection of infrastructure instability or security events. - Mentor and provide technical guidance to L1 engineers, conducting knowledge transfer sessions and assisting with complex ticket resolution to build team capability. - Participate in blameless post-mortems following major incidents, contributing to root cause analysis and implementing preventative actions to improve system reliability. - Maintain and update operational runbooks, network diagrams, and technical documentation to reflect current configurations and best practices. - Support hardware lifecycle management activities including equipment provisioning, asset tracking, and coordination with vendors for hardware returns and repairs. - Provide 24x7 on-call support for critical escalations, ensuring rapid response to high-priority incidents affecting production systems. - Collaborate with the FTE IT Team Lead on capacity planning activities, providing data-driven insights on infrastructure utilization trends and growth projections. Tools Required - Server & Hardware Tools: Dell iDRAC, Lifecycle Controller, OpenManage, RAID/PERC utilities for server provisioning, firmware baselining, and remote management. - OS Deployment Tools: PXE boot infrastructure, iDRAC Virtual Media, Windows Server & Linux ISOs with hardening and automation scripts. - Network Tools: Cisco IOS CLI, PoE management, VLAN/QoS configuration tools, network monitoring, and bandwidth/latency testing utilities. - Automation & Operations Tools: Ansible, Python, CMDB systems, configuration backup tools, and documentation/diagramming platforms for global 24x7 operations.
Related Guides
Related Categories
Related Job Pages
More Infrastructure Engineer Jobs
Infrastructure Engineer
USAP - US Anesthesia PartnersFounded in 2012 to help anesthesiologists create positive patient outcomes, USAP - U.S. Anesthesia Partners serves as a strategic partner to high-quality groups
• Leads and participates in IT infrastructure projects, including the design, deployment, and maintenance of server systems, cloud services, and virtualization platforms. • Ensures projects are completed on time and within budget. • Oversees the deployment of IT infrastructure components, including Windows, Linux, and virtualized environments. • Implements system updates, patches, and upgrades to ensure optimal performance and security. • Works closely with the architecture team to design and implement scalable and secure IT infrastructure solutions. • Ensures alignment with architectural standards and best practices. • Develops and implements backup and disaster recovery plans to ensure data integrity and business continuity. • Scripts in PowerShell for automation of routine tasks and configuration management. • Understands coding and scripting languages to support infrastructure automation. • Implements and manages infrastructure using IaC principles and tools like Terraform to automate deployment and configuration processes. • Participates in Agile workflow methodologies, including sprint planning, daily stand-ups, and retrospectives to ensure efficient project delivery. • Setups monitoring systems and tools to ensure the performance, availability, and reliability of our IT infrastructure. • Responds to IT infrastructure-related issues and provides timely resolution to minimize downtime. • Collaborates with other IT teams and vendors to troubleshoot complex problems and implement solutions. • Maintains accurate documentation of IT infrastructure configurations, processes, and procedures. • Generates regular reports on system performance, capacity, and security metrics.
AI Infrastructure Engineer
Bright Vision TechnologiesBright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.
Role Description We are seeking an AI Performance Optimization Engineer to focus on extracting maximum throughput, minimizing latency, and reducing cost across training and inference workloads for large neural network systems. The role spans the full stack from low-level kernel optimization to distributed system tuning, requiring deep understanding of GPU architecture, model parallelism, memory management, and compiler-level optimization. The ideal candidate has demonstrated an impact on production of AI workloads, with strong instrumentation and measurement discipline that enables rigorous, data-driven optimization decisions. In this role you will work closely with cross-functional partners — product, design, engineering, operations, and business stakeholders — to translate ambiguous requirements into well-engineered solutions, and will be expected to raise the bar through code review, design review, and mentorship of more junior engineers. The successful candidate brings strong engineering discipline, a clear communication style, and a track record of shipping meaningful work that holds up well in production. Key Responsibilities - Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, and cost. - Identify and eliminate bottlenecks across data loading, model compute, communication, and memory. - Implement and tune quantization, sparsity, and pruning strategies to reduce model footprint and accelerate inference. - Optimize distributed training using tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding. - Tune attention implementations using Flash Attention, paged attention, and related techniques. - Implement KV cache optimization, continuous batching, and speculative decoding for LLM serving. - Drive compiler-level optimizations using Triton, XLA, Torch Inductor, or TVM, working with the broader ML framework community to land improvements that translate into measurable end-to-end performance gains. - Optimize data pipelines, sharding strategies, and storage access patterns for high-throughput training. - Build and maintain rigorous benchmark suites and regression frameworks across workloads. - Collaborate with ML and platform engineering teams to embed best practices in standard pipelines. - Drive cost-efficiency improvements through model architecture, hardware selection, and scheduling strategies. - Evaluate new hardware and software offerings and advise on adoption. - Document performance tuning playbooks and share findings broadly across engineering teams. - Stay current with AI systems to research and translate advances into production improvements. Qualifications - Bachelor's or master's degree in computer science, Computer Engineering, or related field. - Six or more years of experience in performance engineering, ML systems, or HPC. - Strong proficiency in Python and C++. - Hands-on experience optimizing deep learning workloads on modern GPUs. - Deep understanding of distributed training and inference techniques. - Experience with profiling tools across CPU, GPU, and distributed systems. - Familiarity with model compression techniques and their accuracy implications. - Strong grasp of memory hierarchies, communication primitives, and parallelism strategies. - Excellent measurement, debugging, and analytical reasoning skills. - Strong communication and collaboration skills. Preferred Qualifications - Experience optimizing LLM inference at production scale. - Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects. - Familiarity with custom kernel authoring in Triton or CUTLASS. - Experience with FinOps for AI workloads. - Publications or talks on AI systems performance. How to Apply Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 650-6699. Learn more about Bright Vision Technologies at www.bvteck.com . Equal Employment Opportunity (EEO) Statement Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall. BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
Enterprise Architect - Infrastructure
BJC HealthCareBJC HealthCare is one of the largest healthcare organizations in the U.S. focused on delivering "the world's best medicine," made better by its 30,000+ clinical
Role Description The Enterprise Architect provides strategic architectural leadership across the enterprise. This role owns enterprise direction, architectural coherence, and cross-domain alignment by translating business strategy into architecture principles, guardrails, reference architectures, lifecycle decisions, and roadmaps. The Enterprise Architect guides and governs rather than designing project-specific solutions or performing implementation work. We are looking for an Infrastructure Enterprise Architect. This is a remote position. We are looking for candidates in MO or IL. Qualifications - Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related field; master’s degree is a plus. - 10+ years of IT experience with 5+ years in architecture roles spanning multiple technology or business domains. - Demonstrated experience with architecture principles, governance models, reference architectures, portfolio analysis, modernization planning, and business capability-based planning. - Broad knowledge across cloud, integration, data, security, infrastructure, and applications, with the ability to operate with breadth rather than deep implementation ownership. - Strong working knowledge of architecture governance, cross-domain dependency management, and arbitrating architectural conflicts. - Strong executive communication, facilitation, and influence skills in complex environments. Requirements - Education: Bach Deg and/or Equivalent Exp - Experience: 5-10 years - Supervisor Experience: No Experience Preferred Qualifications - Education: Bachelor's Degree - Experience: 10+ years - Licenses & Certifications: MS Azure Solns Arch Expert, The Open Group Arch Framework Responsibilities - Collaborates to develop current state, future state architectures and implementation plans. - Mentors solution architects in architecture development and execution of architecture processes. - Reviews proposed architectures and provides guidance for mitigation of critical issues. - Researches and collaborates with teams to identify relevant technology change drivers and opportunities that will impact the enterprise environment. - Leads and mentors enterprise proof-of-concept efforts to ensure alignment with business requirements and key principles. - Develops technology roadmaps which have organizational support and commitment. - Develops and presents planning, status and guidance documents that are audience appropriate. - Develops enterprise and domain reference architectures and processes to provide technical guidance. - Facilitates and mentors solution development and implementation of strategic efforts as needed. - Collaborates with technical teams and SMEs in the development, maintenance and publication of technology standards. - Collaborates with multi-domain technology working groups to identify key issues and develop solution plans. - Collaborates in events focused on information and perspective sharing with internal and external technology leaders and experts. - Actively identifies and engages in efforts to assist with process guidance and information evaluation. - Develops and shares processes that allow for consistent execution of analysis and decision making efforts. Benefits - Comprehensive medical, dental, vision, life insurance, and legal services available first day of the month after hire date. - Disability insurance paid for by BJC. - Annual 4% BJC Automatic Retirement Contribution. - 401(k) plan with BJC match. - Tuition Assistance available on first day. - BJC Institute for Learning and Development. - Health Care and Dependent Care Flexible Spending Accounts. - Paid Time Off benefit combines vacation, sick days, holidays and personal time. - Adoption assistance.
Senior Engineer, Infrastructure
ZencoderZencoder is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
• We’re looking for an Engineer to help build and operate the infrastructure behind Zencoder’s AI-powered products. • You’ll work across our production platform, improving its reliability, security, scalability and cost efficiency. • This includes our Kubernetes foundations, cloud infrastructure, networking, data systems and the internal tooling that enables engineers to deploy and operate services confidently. • This is a hands-on engineering role rather than a traditional operations position. • You’ll write code, automate infrastructure, investigate production issues and design systems that reduce operational complexity as the company grows. • The exact problems will evolve quickly. You should be comfortable taking ownership of unfamiliar systems, identifying the highest-leverage improvements and moving between immediate production needs and longer-term platform investments.



