Located in St. Louis, Missouri, Washington University in St. Louis is an award-winning institution of higher education dedicated to excellence in learning, teac
Systems Engineer II – Shared Infrastructure Data Center
Location
Missouri
Posted
4 days ago
Salary
$65.9K - $112.7K / year
Seniority
Senior
Job Description
Systems Engineer II – Shared Infrastructure Data Center
Washington University in St. Louis
• Perform operational duties and capacity planning efforts for enterprise systems. • Create and maintain documentation related to services, solutions and interfaces. • Perform testing and changes while minimizing risk through disciplined change management practices. • Perform research, analyze technology, consult vendors and apply best practices to design technical solutions by utilizing systems analysis techniques and procedures, including consulting with users, to determine hardware, software or system functional specifications. • Validate, test and implement new products and services. • Respond to and resolve incidents escalated from operational teams and performance tuning requests utilizing critical thinking skills. • Provide technical and advisory leadership as required to complete objectives. • Train and mentor for other personnel. • Perform other duties as assigned.
Job Requirements
- Education: Associate degree or combination of education and experience may substitute for minimum education.
- Work Experience: Information Technology (4 Years)
- Skills: Not Applicable
- Certifications /Professional Licenses: No specific certification/professional license is required for this position.
- Driver's license: A driver's license is not required for this position.
Benefits
- Up to 22 days of vacation, 10 recognized holidays, and sick time.
- Competitive health insurance packages with priority appointments and lower copays/coinsurance.
- Free Metro transit U-Pass for eligible employees.
- Defined contribution (403(b)) Retirement Savings Plan starting at 7% with employee and university contributions.
- Wellness challenges, annual health screenings, mental health resources, mindfulness programs and courses, employee assistance program (EAP), financial resources, access to dietitians, and more!
- 4 weeks of caregiver leave to bond with your new child.
- Tuition coverage for you and your family, including dependent undergraduate-level college tuition up to 100% at WashU and 40% elsewhere after seven years.
Related Guides
Related Categories
Related Job Pages
More Infrastructure Engineer Jobs
Cloud Infrastructure Engineer
RefinedScienceAdvance care by bringing together the best science, data and minds to discover pathways to life beyond disease.
Role Description We are seeking a Cloud Infrastructure Engineer to join our team. You will be responsible for designing, implementing, and maintaining cloud-based infrastructure and automation to meet the evolving needs of our organization. The ideal candidate has hands-on cloud engineering experience, is proficient with infrastructure as code, and is comfortable owning cloud networking, compute, and deployment pipelines in a regulated research environment. A key part of this role is serving as the infrastructure platform support layer for our software development and data science teams — helping them build, deploy, and observe their work reliably and at scale. You will collaborate closely with our development, data science, clinical, and informatics teams to deliver scalable, secure, and reliable infrastructure on Google Cloud Platform (GCP). Key Activities - Design and implement cloud infrastructure on GCP using infrastructure as code - Manage cloud networking components including VPCs, load balancers, DNS, Cloud Router, NAT, and firewall rules - Manage and optimize cloud compute resources including GCE instances, Cloud Run, and related services - Build and maintain CI/CD pipelines using tools such as GitHub Actions or Google Cloud Build to support reliable, repeatable deployments - Design and maintain observability infrastructure including metrics collection, log aggregation, and dashboards to surface actionable insights for engineering and research teams - Implement and maintain security and compliance controls appropriate to a regulated healthcare research environment - Manage cloud services including backup and disaster recovery - Support deployment and maintenance of internal applications and their underlying infrastructure - Champion containerization best practices and support containerized workload deployments - Collaborate with cross-functional teams to ensure seamless integration with existing systems and workflows - Troubleshoot and resolve cloud infrastructure and deployment issues - Create and maintain clear technical documentation, runbooks, and architecture diagrams - Stay current with cloud technology trends and evaluate new tools and approaches relevant to the organization Qualifications - Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience - Experience in cloud engineering, infrastructure engineering, or a closely related role - Strong proficiency with GCP; equivalent experience with AWS or Azure will be considered - Linux administration and comfort working in Linux-based environments via command line - Hands-on experience with infrastructure as code with Terraform and or Pulumi - Demonstrated experience building and maintaining CI/CD pipelines (GitHub Actions, Google Cloud Build, or similar) - Solid understanding of cloud networking: VPC design, load balancers, DNS, NAT, and firewall rule management - Hands-on experience with metrics collection and building operational dashboards - Experience scripting or automating infrastructure tasks in Python or a similar language - Strong written and verbal communication skills with an ability to produce clear technical documentation - Ability to work both independently and collaboratively within a cross-functional team Nice to Haves - Experience in a healthcare or life sciences environment; working knowledge of HIPAA compliance requirements - Experience with GCP-native observability tooling (Cloud Monitoring, Cloud Logging) in addition to Prometheus/Grafana - Experience with cloud IAM best practices and security controls (VPC Service Controls, service accounts, least-privilege patterns) - Experience with front-end application hosting infrastructure (load balancers, SSL/TLS termination, CDN) - Google Cloud Professional certification (Cloud Architect, DevOps Engineer, or similar) - Familiarity with containerization technologies, primarily Docker; Kubernetes exposure is a plus but not required Benefits - Medical, Dental and Vision insurance - Life, AD&D, Short-term and Long-term Disability Insurance (none is 100% covered) - HSA Spending Accounts - 22 Vacation days - 10 Paid Holidays and Sick Time (120 hours per year) - 401(K) Plan
Infrastructure Engineer – Storage Platform
TensorWaveGPU poor? Contact us for your AI cloud compute needs!
• Operate and maintain distributed storage platforms, including Ceph (RBD, CephFS, RGW), High-performance NAS platforms (e.g., Weka, VAST Data) • Manage storage lifecycle operations - cluster expansion, upgrades and migrations • Monitor and maintain storage health, including capacity utilization, data distribution and balance, cluster state and recovery operations • Analyze and troubleshoot storage performance across IOPS, throughput, and latency (including tail latency) • Identify and remediate bottlenecks across disk subsystems, network paths (including RDMA where applicable), client access patterns • Support incident response and root cause analysis for storage-related issues • Ensure storage platforms meet performance expectations for GPU and Kubernetes workloads • Operate and support Kubernetes-integrated storage - CSI drivers, StorageClasses, PersistentVolumes / PersistentVolumeClaims • Troubleshoot storage-related issues in Kubernetes environments, including stateful workloads, performance inconsistencies, scheduling and provisioning failures • Execute and improve automation for storage deployment and operations using Ansible, Terraform, Kubernetes manifests / Helm • Contribute to improving monitoring and alerting, operational workflows, runbooks and documentation • Partner with DevOps and Platform Engineering (automation and orchestration), Network Engineering (high-throughput and RDMA networking), Compute / Virtualization teams • Help ensure end-to-end performance across compute, network, and storage layers
• Manage and support infrastructure across Windows, Unix/Linux, networking, and cloud platforms • Monitor system alerts, triage issues, and respond quickly to maintain uptime • Lead and execute infrastructure projects, partnering with IT and business stakeholders • Support cross-functional teams with infrastructure-related incidents and escalations • Execute core processes including patching, monitoring, and system maintenance • Assist with M&A integrations, including system migrations and infrastructure support • Develop and maintain documentation (architecture, SOPs, standards, policies) • Continuously look for ways to improve performance, reliability, and scalability • Support additional initiatives and projects as needed
We are looking for someone to join our Infrastructure Engineering Team to work on the core infrastructure that underpins our platform. Our systems already process 1000’s of requests per second and with your help we will continue to scale with our clients. You will work closely with teams across engineering and in the wider organisation to build systems with reliability, security, and scalability at their heart. You care deeply about developer experience and will contribute to improving CI/CD pipelines, enabling new capabilities in GCP, and optimising performance of our Kubernetes cluster. If you love automating everything, treating infrastructure as code, and building resilient systems, we want to hear from you.




