Xsolla's video game business engine helps game developers and publishers operate more efficiently and sell more games.
AI Infrastructure Engineer
Location
Malaysia + 1 moreAll locations: Malaysia | Azerbaijan
Posted
122 days ago
Salary
0
Seniority
Senior
Job Description
AI Infrastructure Engineer
Xsolla
Title: AI Infrastructure Engineer Location: Kuala Lumpur Type: Full time Workplace: hybrid Category: Infrastructure Team Job Description: ABOUT YOU We are seeking a hands-on and forward-thinking AI Infrastructure Engineer to help build and operate the intelligent systems that power Xsolla's infrastructure. As part of our Infrastructure Team, you will implement AI-driven solutions across cloud optimization, security, automation, and developer support — helping us shift from manual and reactive operations to predictive, self-optimizing infrastructure management. The ideal candidate brings solid infrastructure engineering experience combined with practical knowledge of AI/ML integration. You are comfortable working with LLMs, ML pipelines, and AI automation frameworks, and you know how to apply them to real operational problems at scale. You thrive in environments that require both technical depth and the ability to experiment, iterate, and deliver. If you're passionate about using AI to transform how infrastructure is built and operated — and want to be part of a team that is driving that transformation at a global gaming company — we'd love to hear from you. ABOUT US Xsolla is a global commerce company with robust tools and services designed to help developers solve the inherent challenges of the video game industry. From indie to AAA, companies partner with Xsolla to help them fund, distribute, market, and monetize their games. Grounded in the belief in the future of video games, Xsolla is resolute in the mission to bring opportunities together, and continually make new resources available to creators. Headquartered and incorporated in Los Angeles, California, Xsolla operates as the merchant of record and has helped over 1,500+ game developers to reach more players and grow their businesses around the world. Responsibilities: - Design and implement AI/ML-powered solutions for infrastructure use cases, including predictive autoscaling, anomaly detection, intelligent cost optimization, and automated remediation across GCP and multi-cloud environments - Build and maintain AI-driven monitoring and observability systems that correlate logs, metrics, and traces to surface root causes, predict bottlenecks, and reduce mean time to resolution (MTTR) - Develop and operate automated incident response workflows using AI-powered playbooks that diagnose, contain, and resolve infrastructure issues with minimal manual intervention - Integrate AI tooling into CI/CD pipelines to improve deployment reliability, automate test prediction, score release health, and support rollback automation - Contribute to the development of internal AI agents and virtual assistants integrated into developer workflows (Slack, IDEs, Confluence) — enabling self-service for provisioning, troubleshooting, and infrastructure guidance - Implement AI/ML-based anomaly detection and automated vulnerability management workflows to enhance the security posture of Xsolla's infrastructure - Prototype and productionize Generative AI solutions for infrastructure automation, including auto-generation of Terraform/Puppet modules, IaC configurations, runbooks, and change documentation - Collaborate with senior engineers and leadership to evolve and execute the infrastructure AI strategy across its implementation phases - Maintain clear documentation of AI tools, integrations, and automated workflows; share knowledge and best practices across the team - Qualifications: - 5–7 years of experience in infrastructure engineering, DevOps, SRE, or a related field - Hands-on experience with GCP (priority) and/or AWS; solid understanding of cloud resource management, scaling, and cost structures - Practical experience building or integrating AI/ML-powered tools in an operational context (anomaly detection, predictive models, LLM-based automation, or similar) - Experience with infrastructure-as-code tools — Terraform, Puppet, Ansible, or equivalent - Proficiency in Python for scripting, automation, and AI/ML integration; Bash or Go a plus - Working knowledge of Kubernetes and container orchestration in production environments - Familiarity with observability and monitoring stacks (Prometheus, Grafana, ELK, Datadog, or similar) - Familiarity with LLM APIs (OpenAI, Anthropic, or similar) and prompt engineering for operational use cases - Strong problem-solving mindset with a bias toward automation and eliminating toil - Fluent in English (written and verbal) - Nice To Have: - Experience with AI workflow orchestration frameworks (LangChain, LlamaIndex, n8n, or similar) - Exposure to AIOps platforms (Dynatrace, Datadog AI, Moogsoft, BigPanda, or similar) - Background in FinOps or AI-driven cloud cost optimization - Familiarity with vector databases (Weaviate, Pinecone, Qdrant) for knowledge retrieval systems - Experience with VMware or hybrid cloud environments - GCP and/or AWS cloud certifications - Prior experience in gaming, high-growth tech, or SaaS platform environments - The duties and responsibilities of this position may evolve over time to support the organization's goals and individual growth - *Salary for this role is based on experience Note This job description is intended to outline the general nature and level of work being performed and is not intended to be an exhaustive list of all duties, responsibilities, and qualifications required. Benefits: We are passionate about fostering a supportive environment for our team, so we prioritize the physical, mental, and emotional well-being of our employees and their families through a comprehensive Benefits Program. This includes medical, dental, and vision, PTO, and a personalized career roadmap for each employee. By investing in professional development through training and educational opportunities, we ensure that our team thrives both personally and professionally. Together, we’re not just building a business; we’re cultivating a community that values creativity, collaboration, and the transformative power of play. Equal Employment Opportunity Statement: Xsolla is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate based on race, color, religion, sex, national origin, age, disability, sexual orientation, gender identity, or any other characteristic protected by law. We consider qualified applicants with criminal histories in accordance with the Fair Chance Act. Relevance to Job Responsibilities: The background check is relevant to this position because of the following role responsibilities: Accessing confidential company data Ensuring compliance with regulatory requirements
Related Guides
Related Categories
Related Job Pages
More Infrastructure Engineer Jobs
EUC Infrastructure Engineer
The Cigna GroupEvernorth Health Services, a division of The Cigna Group, creates pharmacy, care and benefit solutions to improve health and increase vitality. We relentlessly innovate to make the prediction, prevention and treatment of illness and disease more accessible to millions of people. Join us in driving growth and improving lives.
End User Computing Engineering Analyst Position Overview The Cigna International Health IT team are looking for an experienced and motivated EUC Engineering lead who will be responsible for deploying and managing large-scale end user computing solutions throughout the organization. Your contributions will be pivotal in shaping the technology strategy, sustaining a resilient infrastructure, and ensuring an exceptional digital experience for all end users. This opportunity is ideal for an EUC expert with proven experience managing environments across EMEA and Asia Pacific. Job Summary To be considered for this role, candidates must have at least eight years' experience as part of an EUC Engineering team within an enterprise setting. Demonstrate practical experience with most of the technologies below and a proven history of leading large-scale EUC projects and improvements. Fluency in both spoken and written English is essential, along with strong communication and collaboration skills and exceptional problem-solving abilities. Key Responsibilities - Manage and maintain the SCCM/MECM infrastructure, including software deployment, OS imaging, patch management, and compliance reporting. - Manage, maintain and optimize Citrix Cloud and Virtual Desktop Infrastructure (VDI) environments to ensure high availability and performance. - Manage and maintain Group Policy Objects (GPOs) for enforcing security configurations, user environment settings, and compliance standards. - Collaborate with security, networking, and infrastructure teams to deliver comprehensive solutions in line with enterprise standards. - Lead troubleshooting efforts for complex end user computing issues and serve as an escalation point for the wider team. - Contribute to roadmap planning, technology evaluations, and continuous improvement initiatives. - Produce and maintain technical documentation, runbooks, and standard operating procedures. - Proactively monitor and resolve infrastructure service issues before escalation occurs. - Ensure Service Now incident, request, problem, and change records are proactively updated. - Participate in change activities that may be scheduled outside normal working hours. - Drive project deliverables and initiatives to completion, collaborating with key stakeholders and project managers to implement solutions. - Adhere to Information Protection Policies and Business Continuity Plans. Skills and Experience - 5 to 6 years of demonstrable experience leading technical EUC infrastructure projects. - Deep expertise in Microsoft SCCM / MECM / Intune, including configuration, administration, software deployment, OS deployment (OSD/MDT), and patch management. - Proven experience implementing, designing, and supporting Virtual Desktop Infrastructure (VDI) at enterprise scale. - Strong troubleshooting skills in Microsoft Windows, O365, GPO, and FSLogix. - Experience with Service Now change and incident management workflows. - Service and customer-centric mindset. - Excellent problem-solving and critical thinking abilities. Additional desirable experience includes: - DevOps – Image Management and support of Azure Virtual Desktop. - Automation using PowerShell. - Knowledge of Nutanix HCI. - Agile Project Management methodology. - Thin Client Technology. - ITIL certifications. Why Join Cigna? Cigna offers the opportunity to work with a global, innovative, and flexible technology division that is expanding rapidly due to ongoing success and substantial transformation. The company invests in its employees, providing opportunities to advance your knowledge and skills through both internal and external training, as well as secondments to other teams or projects. Flexible working arrangements are available, including remote or home working and flexible shift times, to promote a true work-life balance for all employees. About The Cigna Group Cigna Healthcare, a division of The Cigna Group, advocates for better health throughout every stage of life. The organization guides customers through the healthcare system, empowering them with information and insights necessary to make informed decisions to improve their health and vitality. Join Cigna in driving growth and making a positive impact on lives About The Cigna Group Cigna Healthcare, a division of The Cigna Group, is an advocate for better health through every stage of life. We guide our customers through the health care system, empowering them with the information and insight they need to make the best choices for improving their health and vitality. Join us in driving growth and improving lives.
• Design, implement, and manage core infrastructure on Google Cloud Platform (GCP) • Build and operate resilient, highly available distributed systems using Kubernetes (GKE), Knative, Istio, and related technologies • Automate infrastructure life cycle using Terraform and Terragrunt • Implement and maintain CI/CD pipelines and deployment tools like Flux and Helm • Optimize and manage multi-tenant data layer on Postgres and Neon
Lead IT Infrastructure Engineer
Lumen TechnologiesLumen Technologies is self-described as a global company of 40,000+ professionals empowering businesses, government, and communities to “produce amazing thing
Lumen is the trusted network for AI. We’re transforming how businesses connect, secure, and scale in an AI-driven world. By connecting people, data, and applications quickly, securely, and effortlessly, we help organizations move faster and unlock what’s next. At Lumen, people power progress. Our culture is built on teamwork, trust, and transparency, giving you the flexibility, support, and opportunity to make a lasting impact. We’re looking for top-tier talent ready to take on the challenge. Join us in building the future. The Role Provide expert technical direction across multiple and complex systems, including planning, designing, implementing, and maintaining enterprise-level network infrastructure. This role combines advanced networking expertise with collaboration and leadership to ensure optimal network performance, security, and scalability across the organization. The position requires working with cross-functional teams, mentoring junior engineers, and processing service requests for complex networking challenges while maintaining high availability and security standards. You must be a US Citizen or Permanent Resident/Green Card for consideration for this position. This position will have support responsibilities for Government contracts and will additionally require successfully passing GSA suitability (public trust/Tier 2S) investigations. Location This is a work from home position within the US. The Main Responsibilities - Engineer and support routing & switching services with day-to-day and on-call responsibilities. - Deliver and maintain data center fabric / SDN environments including service requests, upgrades, problem remediation, and standardization across environments. - Perform complex network changes and migrations ensuring change planning, risk management, rollback readiness, and post-change validation. - Assist with load balancing operations and request fulfillment - Provide escalation and incident leadership for complex, cross-domain outages and degradations; drive root-cause isolation, restoration, and durable fixes while coordinating stakeholders. - Familiarity with hybrid connectivity patterns that integrate on-prem and cloud network services, ensuring seamless routing, segmentation, and operational supportability for migrations and steady-state operations. - Monitor and optimize network health using observability and diagnostic tooling, proactively identifying risks and improving reliability. - Develop and maintain engineering documentation including network/service diagrams, design references, operational runbooks, and standards. - Familiarity with automation and operational efficiency of operational tasks. - Collaborate with internal partners and external vendors to deliver business-focused solutions. - Mentor and provide technical leadership through peer reviews, knowledge sharing, and setting best practices for consistent engineering outcomes. What We Look For in a Candidate - Bachelor’s degree or equivalent education and experience with typically 5+ years Enterprise level support and design experience - Data Center Experience: Proven experience, or a willingness to learn, maintaining Cisco ACI. - Networking Technologies: Demonstrated expertise in implementing and managing enterprise-level services including advanced routing protocols (BGP and OSPF), virtual routing (VRFs), and network segmentation. Compensation This information reflects the anticipated base salary range for this position based on current national data. Minimums and maximums may vary based on location. Individual pay is based on skills, experience and other relevant factors. Location Based Pay Ranges $105,786 - $141,047 in these states: AL AR AZ FL GA IA ID IN KS KY LA ME MO MS MT ND NE NM OH OK PA SC SD TN UT VT WI WV WY $111,074 - $148,099 in these states: CO HI MI MN NC NH NV OR RI $116,364 - $155,152 in these states: AK CA CT DC DE IL MA MD NJ NY TX VA WA Lumen offers a comprehensive package featuring a broad range of Health, Life, Voluntary Lifestyle benefits and other perks that enhance your physical, mental, emotional and financial wellbeing. We're able to answer any additional questions you may have about our bonus structure (short-term incentives, long-term incentives and/or sales compensation) as you move through the selection process. Learn more about Lumen's:BenefitsBonus Structure#LI-Remote #LI-ZM1 Requisition #: 341755 Background Screening If you are selected for a position, there will be a background screen, which may include checks for criminal records and/or motor vehicle reports and/or drug screening, depending on the position requirements. For more information on these checks, please refer to the Post Offer section of our FAQ page. Job-related concerns identified during the background screening may disqualify you from the new position or your current role. Background results will be evaluated on a case-by-case basis. Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Equal Employment Opportunities We are committed to providing equal employment opportunities to all persons regardless of race, color, ancestry, citizenship, national origin, religion, veteran status, disability, genetic characteristic or information, age, gender, sexual orientation, gender identity, gender expression, marital status, family status, pregnancy, or other legally protected status (collectively, “protected statuses”). We do not tolerate unlawful discrimination in any employment decisions, including recruiting, hiring, compensation, promotion, benefits, discipline, termination, job assignments or training. Privacy Notice Lumen is committed to protecting the privacy and security of personal information collected during the recruitment and hiring process. Our Privacy Notice explains how we collect, use, disclose, and protect applicant information, as well as how individuals may request access to or deletion of their personal data. To review Lumen’s Privacy Notice, please visit: https://jobs.lumen.com/global/en/privacy-notice Disclaimer The job responsibilities described above indicate the general nature and level of work performed by employees within this classification. It is not intended to include a comprehensive inventory of all duties and responsibilities for this job. Job duties and responsibilities are subject to change based on evolving business needs and conditions. In any materials you submit, you may redact or remove age-identifying information such as age, date of birth, or dates of school attendance or graduation. You will not be penalized for redacting or removing this information. Please be advised that Lumen does not require any form of payment from job applicants during the recruitment process. All legitimate job openings will be posted on our official website or communicated through official company email addresses. If you encounter any job offers that request payment in exchange for employment at Lumen, they are not for employment with us, but may relate to another company with a similar name.
Staff ML Infrastructure Engineer - Embodied AI Offboard Perception
General MotorsGeneral Motors (GM), founded in 1908 by William "Billy" Durant in Flint, Michigan, began with the Buick Motor Company and later acquired brands like Oldsmobile
Description At General Motors, our product teams are redefining mobility. Through a human-centered design process, we create vehicles and experiences that are designed not just to be seen, but to be felt. We're turning today's impossible into tomorrow's standard -from breakthrough hardware and battery systems to intuitive design, intelligent software, and next-generation safety and entertainment features. Every day, our products move millions of people as we aim to make driving safer, smarter, and more connected, shaping the future of transportation on a global scale. Are you passionate about accelerating the future of autonomous driving? Join the Embodied AI team at General Motors. Our team is developing and deploying machine learning solutions that support safe and reliable autonomous vehicle behavior across real-world scenarios. As a Staff ML Infra Engineer on the Offboard Perception team within the Embodied AI organization, you will be a senior engineer responsible for developing and deploying offboard machine learning solutions that deliver ground-truth-quality world estimates for multiple partner teams, including onboard model teams, simulation, and evaluation. The models you build will influence every stage of autonomous vehicle development-from training and validation to testing and safety. You will work closely with cross-functional engineering teams, help shape technical direction in your domain, and support other engineers' growth through collaboration and mentorship. You will also help transition research into scalable onboard ML capabilities while continuously improving the autonomy stack. What You'll Do - Design, build, and maintain ML infrastructure that enables rapid development, training, evaluation, and deployment of offboard perception models. - Own the integration of models into production systems, including packaging, validation, deployment, rollout strategies. - Implement CI/CD pipelines for ML systems, including automated testing, model validation, performance regression checks, and deployment automation. - Establish model evaluation and observability frameworks, including training metrics, inference performance metrics, data quality checks, and production monitoring dashboards. - Develop infrastructure for experiment tracking and benchmarking, enabling teams to compare model architectures, datasets, hyperparameters, and training procedures in a reliable and repeatable way. - Support efficient dataset curation and ingestion pipelines that help prioritize high-value data, accelerate iteration cycles, and improve model performance on hard-edge cases. - Partner with ML engineers, researchers, and software teams to ensure models can be reliably integrated into larger autonomy stacks and production services at scale. - Define and enforce best practices for ML systems engineering, including reproducibility, configuration management, artifact management, security, and operational readiness. - Support technical collaboration through code reviews, design reviews, and mentorship, helping raise the quality and maintainability of ML infrastructure across the organization. Your Skills & Abilities - Strong software engineering fundamentals, including experience building reliable, maintainable, and scalable production systems. - Proficiency in Python, with experience using ML and scientific computing libraries such as PyTorch, NumPy, and related tooling. - Experience building and supporting ML training and deployment pipelines, including data processing, experiment execution, model packaging, and production rollout. - Experience deploying ML models into production environments, with understanding of end-to-end workflows such as validation, serving, monitoring, and lifecycle management. - Familiarity with distributed training and large-scale compute infrastructure, including GPUs, cluster scheduling, and performance optimization for training workloads. - Experience with containerization, orchestration, and automation tools such as Docker, Kubernetes, workflow schedulers, and CI/CD systems. - Experience with model observability and operational metrics, including training metrics, inference performance, reliability monitoring, and data/model drift detection. - Strong communication and collaboration skills, with the ability to work effectively across ML, software, data, and systems engineering teams. - Experience in robotics, perception systems, or autonomous driving is preferred. Remote/Hybrid: This role is based remotely but if you live within a 50-mile radius of Austin, Detroit, Warren, Milford or Mountain View, you are expected to report to that location three times per week, at minimum. Compensation: The compensation information is a good faith estimate only. It is based on what a successful applicant might be paid in accordance with applicable state laws. The compensation may not be representative for positions located outside of the California Bay Area. - The salary range for this role is $189,300.00 to $290,700.00. The actual base salary a successful candidate will be offered within this range will vary based on factors relevant to the position. - Bonus Potential: An incentive pay program offers payouts based on company performance, job level, and individual performance. Benefits: GM offers a variety of health and wellbeing benefit programs. Benefit options include medical, dental, vision, Health Savings Account, Flexible Spending Accounts, retirement savings plan, sickness and accident benefits, life insurance, paid vacation & holidays, tuition assistance programs, employee assistance program, GM vehicle discounts and more. Relocation: This job may be eligible for relocation benefits. Company Vehicle: Upon successful completion of a motor vehicle report review, you will be eligible to participate in a company vehicle evaluation program, through which you will be assigned a General Motors vehicle to drive and evaluate. Note: program participants are required to purchase/lease a qualifying GM vehicle every four years unless one of a limited number of exceptions applies. #LI-CX1 #GM-AV-1 About GM Our vision is a world with Zero Crashes, Zero Emissions and Zero Congestion and we embrace the responsibility to lead the change that will make our world better, safer and more equitable for all. Why Join Us We believe we all must make a choice every day - individually and collectively - to drive meaningful change through our words, our deeds and our culture. Every day, we want every employee to feel they belong to one General Motors team. Total Rewards | Benefits Overview From day one, we're looking out for your well-being-at work and at home-so you can focus on realizing your ambitions. Learn how GM supports a rewarding career that rewards you personally by visiting Total Rewards resources. Non-Discrimination and Equal Employment Opportunities (U.S.) General Motors is committed to being a workplace that is not only free of unlawful discrimination, but one that genuinely fosters inclusion and belonging. We strongly believe that providing an inclusive workplace creates an environment in which our employees can thrive and develop better products for our customers. All employment decisions are made on a non-discriminatory basis without regard to sex, race, color, national origin, citizenship status, religion, age, disability, pregnancy or maternity status, sexual orientation, gender identity, status as a veteran or protected veteran, or any other similarly protected status in accordance with federal, state and local laws. We encourage interested candidates to review the key responsibilities and qualifications for each role and apply for any positions that match their skills and capabilities. Applicants in the recruitment process may be required, where applicable, to successfully complete a role-related assessment(s) and/or a pre-employment screening prior to beginning employment. To learn more, visit How we Hire. Accommodations General Motors offers opportunities to all job seekers including individuals with disabilities. If you need a reasonable accommodation to assist with your job search or application for employment, email us [email protected] or call us at 1-800-865-7580. In your email, please include a description of the specific accommodation you are requesting as well as the job title and requisition number of the position for which you are applying.




