IT services that put people at the center of your business
IT Infrastructure Operations Engineer I
Location
India
Posted
3 days ago
Salary
0
Seniority
Mid Level
Job Description
IT Infrastructure Operations Engineer I
Astreya
• Monitor server and network infrastructure health using established monitoring tools and dashboards, identifying alerts and anomalies requiring attention. • Provide first-level response to service tickets within the organization's ticketing system, ensuring 2-hour initial response SLA compliance for all incoming requests. • Perform basic troubleshooting of Dell PowerEdge server hardware issues using iDRAC interfaces, RAID, escalating complex problems to L2 support with detailed documentation. • Execute routine health checks on Cisco routers and switches, documenting status and flagging any deviations from normal operating parameters. • Follow standardized runbooks and operational procedures for common infrastructure issues, ensuring consistent resolution approaches across the team. • Log and document all incidents, actions taken, and resolutions in the ticketing system with accurate and detailed information for knowledge management. • Escalate unresolved or complex issues to L2/L3 engineers with comprehensive handover notes including symptoms, actions attempted, and relevant logs. • Assist with routine maintenance activities including scheduled reboots, basic configuration backups, and pre-approved firmware update executions under supervision. • Participate in shift handover meetings, providing clear status updates on open tickets and ongoing issues to incoming team members. • Coordinate with on-site technicians and vendors for basic hardware replacement activities, ensuring proper ticketing and tracking of all dispatch requests. • Maintain awareness of scheduled maintenance windows and change activities, monitoring for any unexpected impacts during and after implementation. • Contribute to the continuous improvement of runbooks and documentation by identifying gaps and suggesting updates based on real-world incident handling experience.
Job Requirements
- Bachelor's degree in Computer Science, Information Technology, or related field, or equivalent practical experience with 2+ years in IT support roles.
- Basic understanding of server hardware components, networking fundamentals (TCP/IP, DNS, DHCP), and familiarity with enterprise infrastructure concepts.
- Foundational knowledge of ticketing systems, monitoring tools, and remote management interfaces; exposure to Dell iDRAC or similar BMC tools is advantageous.
- Strong verbal and written communication skills with the ability to document issues clearly and escalate effectively to senior engineers.
- Willingness to work in a 24x7 rotational shift environment, including nights, weekends, and holidays, supporting a global infrastructure.
- Industry certifications such as CompTIA A+, Network+, or CCNA (or actively pursuing) are preferred; demonstrated eagerness to learn and grow technically.
Benefits
- Professional development opportunities
Related Guides
Related Categories
Related Job Pages
More Infrastructure Engineer Jobs
Infrastructure Engineer
US Anesthesia PartnersQuality Anesthesia Care: We're raising the bar for the industry.
• Leads and participates in IT infrastructure projects, including the design, deployment, and maintenance of server systems, cloud services, and virtualization platforms. • Ensures projects are completed on time and within budget. • Oversees the deployment of IT infrastructure components, including Windows, Linux, and virtualized environments. • Implements system updates, patches, and upgrades to ensure optimal performance and security. • Works closely with the architecture team to design and implement scalable and secure IT infrastructure solutions. • Ensures alignment with architectural standards and best practices. • Develops and implements backup and disaster recovery plans to ensure data integrity and business continuity. • Scripts in PowerShell for automation of routine tasks and configuration management. • Understands coding and scripting languages to support infrastructure automation. • Implements and manages infrastructure using IaC principles and tools like Terraform to automate deployment and configuration processes. • Participates in Agile workflow methodologies, including sprint planning, daily stand-ups, and retrospectives to ensure efficient project delivery. • Setups monitoring systems and tools to ensure the performance, availability, and reliability of our IT infrastructure. • Responds to IT infrastructure-related issues and provides timely resolution to minimize downtime. • Collaborates with other IT teams and vendors to troubleshoot complex problems and implement solutions. • Maintains accurate documentation of IT infrastructure configurations, processes, and procedures. • Generates regular reports on system performance, capacity, and security metrics.
AI Infrastructure Engineer
Bright Vision TechnologiesBright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.
Role Description We are seeking an AI Performance Optimization Engineer to focus on extracting maximum throughput, minimizing latency, and reducing cost across training and inference workloads for large neural network systems. The role spans the full stack from low-level kernel optimization to distributed system tuning, requiring deep understanding of GPU architecture, model parallelism, memory management, and compiler-level optimization. The ideal candidate has demonstrated an impact on production of AI workloads, with strong instrumentation and measurement discipline that enables rigorous, data-driven optimization decisions. In this role you will work closely with cross-functional partners — product, design, engineering, operations, and business stakeholders — to translate ambiguous requirements into well-engineered solutions, and will be expected to raise the bar through code review, design review, and mentorship of more junior engineers. The successful candidate brings strong engineering discipline, a clear communication style, and a track record of shipping meaningful work that holds up well in production. Key Responsibilities - Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, and cost. - Identify and eliminate bottlenecks across data loading, model compute, communication, and memory. - Implement and tune quantization, sparsity, and pruning strategies to reduce model footprint and accelerate inference. - Optimize distributed training using tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding. - Tune attention implementations using Flash Attention, paged attention, and related techniques. - Implement KV cache optimization, continuous batching, and speculative decoding for LLM serving. - Drive compiler-level optimizations using Triton, XLA, Torch Inductor, or TVM, working with the broader ML framework community to land improvements that translate into measurable end-to-end performance gains. - Optimize data pipelines, sharding strategies, and storage access patterns for high-throughput training. - Build and maintain rigorous benchmark suites and regression frameworks across workloads. - Collaborate with ML and platform engineering teams to embed best practices in standard pipelines. - Drive cost-efficiency improvements through model architecture, hardware selection, and scheduling strategies. - Evaluate new hardware and software offerings and advise on adoption. - Document performance tuning playbooks and share findings broadly across engineering teams. - Stay current with AI systems to research and translate advances into production improvements. Qualifications - Bachelor's or master's degree in computer science, Computer Engineering, or related field. - Six or more years of experience in performance engineering, ML systems, or HPC. - Strong proficiency in Python and C++. - Hands-on experience optimizing deep learning workloads on modern GPUs. - Deep understanding of distributed training and inference techniques. - Experience with profiling tools across CPU, GPU, and distributed systems. - Familiarity with model compression techniques and their accuracy implications. - Strong grasp of memory hierarchies, communication primitives, and parallelism strategies. - Excellent measurement, debugging, and analytical reasoning skills. - Strong communication and collaboration skills. Preferred Qualifications - Experience optimizing LLM inference at production scale. - Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects. - Familiarity with custom kernel authoring in Triton or CUTLASS. - Experience with FinOps for AI workloads. - Publications or talks on AI systems performance. How to Apply Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 650-6699. Learn more about Bright Vision Technologies at www.bvteck.com . Equal Employment Opportunity (EEO) Statement Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall. BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
Enterprise Architect - Infrastructure
BJC HealthCareBJC HealthCare is one of the largest healthcare organizations in the U.S. focused on delivering "the world's best medicine," made better by its 30,000+ clinical
Role Description The Enterprise Architect provides strategic architectural leadership across the enterprise. This role owns enterprise direction, architectural coherence, and cross-domain alignment by translating business strategy into architecture principles, guardrails, reference architectures, lifecycle decisions, and roadmaps. The Enterprise Architect guides and governs rather than designing project-specific solutions or performing implementation work. We are looking for an Infrastructure Enterprise Architect. This is a remote position. We are looking for candidates in MO or IL. Qualifications - Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related field; master’s degree is a plus. - 10+ years of IT experience with 5+ years in architecture roles spanning multiple technology or business domains. - Demonstrated experience with architecture principles, governance models, reference architectures, portfolio analysis, modernization planning, and business capability-based planning. - Broad knowledge across cloud, integration, data, security, infrastructure, and applications, with the ability to operate with breadth rather than deep implementation ownership. - Strong working knowledge of architecture governance, cross-domain dependency management, and arbitrating architectural conflicts. - Strong executive communication, facilitation, and influence skills in complex environments. Requirements - Education: Bach Deg and/or Equivalent Exp - Experience: 5-10 years - Supervisor Experience: No Experience Preferred Qualifications - Education: Bachelor's Degree - Experience: 10+ years - Licenses & Certifications: MS Azure Solns Arch Expert, The Open Group Arch Framework Responsibilities - Collaborates to develop current state, future state architectures and implementation plans. - Mentors solution architects in architecture development and execution of architecture processes. - Reviews proposed architectures and provides guidance for mitigation of critical issues. - Researches and collaborates with teams to identify relevant technology change drivers and opportunities that will impact the enterprise environment. - Leads and mentors enterprise proof-of-concept efforts to ensure alignment with business requirements and key principles. - Develops technology roadmaps which have organizational support and commitment. - Develops and presents planning, status and guidance documents that are audience appropriate. - Develops enterprise and domain reference architectures and processes to provide technical guidance. - Facilitates and mentors solution development and implementation of strategic efforts as needed. - Collaborates with technical teams and SMEs in the development, maintenance and publication of technology standards. - Collaborates with multi-domain technology working groups to identify key issues and develop solution plans. - Collaborates in events focused on information and perspective sharing with internal and external technology leaders and experts. - Actively identifies and engages in efforts to assist with process guidance and information evaluation. - Develops and shares processes that allow for consistent execution of analysis and decision making efforts. Benefits - Comprehensive medical, dental, vision, life insurance, and legal services available first day of the month after hire date. - Disability insurance paid for by BJC. - Annual 4% BJC Automatic Retirement Contribution. - 401(k) plan with BJC match. - Tuition Assistance available on first day. - BJC Institute for Learning and Development. - Health Care and Dependent Care Flexible Spending Accounts. - Paid Time Off benefit combines vacation, sick days, holidays and personal time. - Adoption assistance.
Senior Engineer, Infrastructure
ZencoderZencoder is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
• We’re looking for an Engineer to help build and operate the infrastructure behind Zencoder’s AI-powered products. • You’ll work across our production platform, improving its reliability, security, scalability and cost efficiency. • This includes our Kubernetes foundations, cloud infrastructure, networking, data systems and the internal tooling that enables engineers to deploy and operate services confidently. • This is a hands-on engineering role rather than a traditional operations position. • You’ll write code, automate infrastructure, investigate production issues and design systems that reduce operational complexity as the company grows. • The exact problems will evolve quickly. You should be comfortable taking ownership of unfamiliar systems, identifying the highest-leverage improvements and moving between immediate production needs and longer-term platform investments.



