AI Developer Cloud
Datacenter Infrastructure Specialist
Location
United States
Posted
3 days ago
Salary
0
Seniority
Senior
Job Description
Datacenter Infrastructure Specialist
Runpod
• Assist in validating new hardware, ensuring partner deployments meet Runpod’s specifications for distributed AI/ML workloads. • Monitor fleet health to identify performance degradation. You will help audit downtime and provide the technical data needed to protect customer SLAs. • We operate with an AI-first mindset, powering our operations with the technology we host. You will work with LLMs and AI agents to help automate network triage and generate dynamic runbooks for our fleet. • Coordinate technical incident communications with clear updates, acting as a steady hand that translates outages into actionable resolutions. • Support the growth of our infrastructure partners.
Job Requirements
- 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering.
- Strong proficiency in standard datacenter networking and performance troubleshooting. Exposure to RDMA, InfiniBand, or RoCE is highly preferred.
- Hands-on experience with the NVIDIA Software Stack (driver installation, performance utilities) and an understanding of multi-node performance tuning.
- Solid Linux system administration skills and experience with containerization (Docker). You are comfortable performing system-level troubleshooting and performance tuning at the kernel and hardware interface layers.
- Clear written and verbal communication skills. You can explain hardware or networking issues to both technical partners and internal leadership.
- As our global fleet scales, this role may require participating in an on-call rotation in the future.
- You are detail-oriented and proactive when it comes to identifying potential failures before they impact customers.
- Experience working in a fast-paced environment where you have contributed to building operational workflows.
- Experience managing or optimizing bare-metal High-Performance Computing environments at massive scale.
- Experience with Grafana, Prometheus, or Datadog to monitor system health.
- Proficiency in Python, Go (Golang), or Bash to automate repetitive infrastructure tasks and interface with internal APIs.
Benefits
- Meaningful equity in a fast-growing AI infra company — everyone on the team receives stock options — your impact drives our growth, and you share in the upside.
- Generous medical, dental & vision plans — we cover 100% for all employees and partial for dependents.
- Flexible PTO — take the time you need to recharge.
- Most roles are remote work first with inclusive, collaborative teams utilizing Slack as the main form of internal communication.
- Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale.
- $1,200 Home Office & Equipment Stipend — We set you up for success from day one with gear and support to create your ideal workspace.
Related Guides
Related Categories
Related Job Pages
More Infrastructure Engineer Jobs
• Design, implement, and maintain IT infrastructure, including networks, servers, storage, and virtualization. • Troubleshoot and resolve IT issues in a timely and efficient manner. • Collaborate with internal teams to identify and prioritize IT projects and initiatives. • Develop and maintain technical documentation and training materials for employees. • Ensure compliance with industry standards, regulations, and company policies. • Stay up-to-date with emerging technologies and trends in the automotive software industry. • Provide End User Support with JDP Issued PC/Laptop hardware and software, monitoring SugarCRM and Service Now ticketing systems. • Onboarding and Offboarding activities, including procuring and retrieving assets. • Financial database servers and client workstations. Including 3rd part plugin support. • Data Center Work with INAP/Horizon IQ/Evocative to ensure proper operation of the hosted environments. • Perform system, database and server patching. • Participate in on call, off hours scheduled and unscheduled work. • Manage and administer datacenter and cloud environments. • Configure monitoring and alerting. • Perform and monitor environment backups . • Implement environment changes moves/adds/deletes. • Support SMTP and SMS Gateways.
Specialist, Technology Infrastructure
BetaNXT IncBetaNXT is a leading provider of frictionless wealth management infrastructure, real-time data solutions, and an enhanced advisor experience. We invest in platforms, products, and partnerships to accelerate growth for the ecosystem we serve. Our connective approach empowers our clients to deliver a comprehensive, end-to-end advisor and investor experience. BetaNXT is a premier provider of technology, data, and operations as services to a rich client base of wealth managers, institutional wealth firms, and digital brokers. It is comprised of three industry-leading businesses which, combined, provide end-to-end solutions across the investment lifecycle.
Role Description The Specialist, Technology Infrastructure supports the planning, coordination, and delivery of technology infrastructure initiatives that enable secure, reliable, and scalable business operations across BetaNXT. This role serves as a key contributor in ensuring infrastructure projects, operational activities, and service delivery efforts are executed efficiently and aligned with organizational priorities. Working closely with Infrastructure, Service Desk, Engineering, Product, and business stakeholders, the Specialist coordinates activities across multiple teams, facilitates communication, tracks project deliverables, and helps drive successful outcomes for infrastructure-related initiatives. This position plays an important role in maintaining operational excellence, supporting continuous improvement efforts, and enhancing the overall effectiveness of technology services across the organization. The ideal candidate is highly organized, collaborative, and detail-oriented, with strong communication and project coordination skills and an interest in technology infrastructure within a fast-paced, regulated environment. Duties and Responsibilities - Coordinate and support the successful delivery of infrastructure projects, initiatives, and operational activities. - Track project milestones, dependencies, risks, and action items to help ensure deliverables are completed on schedule and within established objectives. - Facilitate communication and collaboration across Infrastructure, Service Desk, Engineering, and business teams. - Assist with planning, prioritizing, and monitoring infrastructure work to support both internal and client-facing commitments. - Prepare and maintain project documentation, status updates, meeting materials, and delivery reporting. - Coordinate resources and activities across multiple teams to support efficient project execution. - Communicate project progress, timelines, requirements, and potential constraints to stakeholders. - Partner with technical teams to understand business requirements and help align delivery activities with organizational goals. - Support issue identification, escalation, and resolution efforts to minimize delays and operational impacts. - Contribute to process improvement initiatives that enhance infrastructure delivery effectiveness, consistency, and customer satisfaction. - Ensure compliance with internal controls, operational standards, and regulatory requirements applicable to infrastructure activities. - Foster strong working relationships with internal stakeholders and support a culture of accountability, collaboration, and continuous improvement. Qualifications - Bachelor's degree or equivalent combination of education and relevant work experience. - 2–4 years of experience supporting technology, infrastructure, service delivery, project coordination, or related operational functions. - Strong organizational skills with the ability to manage multiple priorities in a fast-paced environment. - Proficiency with Microsoft Office Suite, including Excel, PowerPoint, Outlook, and Teams. - Experience coordinating activities across multiple stakeholders and business functions. - Strong written and verbal communication skills. - Experience working within technology, fintech, financial services, insurance, or other highly regulated industries preferred. Preferred Qualifications - Experience supporting infrastructure, technology operations, service delivery, or project management initiatives. - Working knowledge of infrastructure concepts including networks, servers, cloud platforms, and enterprise technology environments. - PMP, CAPM, Agile, Scrum, or other project management certifications are desirable but not required. - Experience creating project plans, status reporting, and executive-level communications. - Demonstrated ability to identify process improvement opportunities and drive efficiencies. - Experience working in environments with regulatory, security, and compliance requirements. Company Description BetaNXT is a leading provider of frictionless wealth management infrastructure, real-time data solutions, and an enhanced advisor experience. We invest in platforms, products, and partnerships to accelerate growth for the ecosystem we serve. Our connective approach empowers our clients to deliver a comprehensive, end-to-end advisor and investor experience. BetaNXT is a premier provider of technology, data, and operations as services to a rich client base of wealth managers, institutional wealth firms, and digital brokers. It is comprised of three industry-leading businesses which, combined, provide end-to-end solutions across the investment lifecycle.
Senior Splunk Engineer - Infrastructure Operations
GovCIOGovCIO is a service-disabled-veteran-owned small business (SDVOSB) that offers technology services to improve business performance for government organizations.
Role Description GovCIO is currently hiring for Senior Splunk Engineer - Infrastructure Operations to support our Administrative Office of the US Courts NLS project. The NLS currently ingests an average of 18-20TB of logging data daily across 60 indexers distributed in 2 data centers. This position is located within the United States and is fully remote. - Design, implement, and operate the Splunk Core, Enterprise Security, IT Service Intelligence (i.e., ITSI), Phantom (Security Orchestration, Automation, and Response (SOAR)), Splunk Cloud, Splunk On-Call, and Multi-Site Index Clustering environment. - Monitor overall Splunk health through the Monitoring Console (DMC) including indexer, search head, and cluster master status. - Track indexing rates, license usage, queue health, and search concurrency to identify performance or ingestion issues early. - Monitor CPU, memory, and disk utilization across all Splunk components to ensure optimal resource usage. - Respond promptly to health alerts, DMC warnings, or anomalies observed on monitoring dashboards. - Investigate and resolve common user-reported issues such as access problems, failed searches, or non-triggering alerts. - Troubleshoot data ingestion, parsing, and indexing issues across Universal Forwarders, Heavy Forwarders, and HEC endpoints. - Investigate missing or duplicate logs, timestamp errors, or sourcetype misassignments and escalate complex parsing issues to Engineering. - Validate new data source onboardings by confirming sourcetype assignment, timestamp accuracy, and field extraction integrity. - Support data source owners with forwarder deployment, syslog setup, and connectivity troubleshooting during initial onboarding. - Maintain data flow visibility from source → forwarder → indexer to confirm data completeness and performance. - Rotate and update credentials, API keys, or tokens used in data inputs, integrations, alerts, and scheduled searches. - Manage RBAC user and role mappings, handling access requests, entitlement reviews, and permission troubleshooting. - Provide end-user assistance with SPL searches, reports, alerts, and dashboards, including query optimization tips. - Maintain and update knowledge base articles, SOPs, and FAQs for repeatable issues and troubleshooting steps. - Log and escalate platform or parsing issues to the Engineering team with evidence such as logs, screenshots, and correlation IDs. - Open and manage Splunk Support cases for platform-level bugs, license problems, or critical system faults. - Monitor and manage ITSI service health, including KPIs, correlation searches, NEAP policies, and summary index latency. - Troubleshoot ITSI-related issues such as broken KPIs, delayed episodes, or missing notable events. - Perform capacity management by monitoring index growth, bucket rotation, and frozen data retention policies. - Conduct periodic system maintenance tasks, including orphaned object cleanup and knowledge object review. - Verify and maintain compliance with data governance and retention policies, ensuring secure and auditable configurations. - Participate in DR testing and validation to ensure Splunk data recovery and HA configurations are functioning as expected. - Document incidents, RCA findings, and preventive actions for future reference. - Collaborate closely with the Engineering team for escalations, root-cause investigations, and deployment verifications. Qualifications - Bachelor's with 10 years (or commensurate experience) OR - Masters Degree or higher (in a related discipline) with 7 years experience Requirements - Expert skills in Enterprise Security, ITSI, SOAR, and the Splunk product line. - Able to design, implement, and operate the Splunk Core, Enterprise Security, IT Service Intelligence (i.e., ITSI), Phantom (Security Orchestration, Automation, and Response (SOAR)), Splunk Cloud, Splunk On-Call, and Multi-Site Index Clustering environment. - Clearance Required: Must be able to obtain and maintain AOUSC Public Trust. Posted Salary Range USD $105,000.00 - USD $145,000.00 /Yr.
Senior Infrastructure Engineer
NavaBuilding simple, effective government services. Want to contribute? We're hiring!
• The Infrastructure Engineer will work on cross functional teams to build scalable infrastructure for our government -- designing, implementing, and delivering services that millions of Americans depend on. • You may be responsible for technical leadership of platform infrastructure projects supporting dev teams and large-scale systems with a strong focus on automation, organization, and communication. • Your work enables application development teams to run effectively in the cloud to deliver users a modern experience. • You will be responsible for delivering on client requirements as well as providing long-term oversight and vision to help shape future work. • Your strong cross-functional communication skills will complement your technical leadership skills to provide the highest value for government systems.


