OpenNebula logo
OpenNebula

The Open Source Cloud & Edge Computing Platform 🚀

AI Platform Engineer – OneAI

Platform EngineerPlatform EngineerFull TimeRemoteSeniorTeam 11-50Since 2010H1B No SponsorCompany SiteLinkedIn

Location

Spain

Posted

106 days ago

Salary

0

Seniority

Senior

Bachelor Degree3 yrs expEnglishCloudKubernetes

Job Description

AI Platform Engineer – OneAI

OpenNebula

• Design, implement, and deploy advanced AI capabilities within the OneAI platform. • Shape the end-user experience by designing intuitive workflows for model management, deployment configuration, and job operation. • Streamline the model lifecycle by integrating public repositories (e.g., Hugging Face) for seamless discovery, import, versioning, and deployment. • Bridge the gap between systems engineering and product design to ensure a seamless transition from backend infrastructure to user features. • Integrate cutting-edge AI frameworks and engines, such as vLLM, NVIDIA Dynamo and Unsloth, into a secure and scalable environment. • Leverage OpenNebula to orchestrate high-performance inference and training workloads across diverse cloud and edge environments. • Develop and maintain reliable APIs for compute provisioning and workload scheduling. • Implement GPU-aware operations to ensure optimal resource allocation and hardware utilization. • Build comprehensive observability suites to monitor and track critical metrics, including latency, throughput, utilization, and failure rates. • Establish and refine deployment and workflow strategies to ensure AI workloads remain efficient and stable at scale. • Optimize system architecture to balance high performance with cost efficiency. • Research and integrate emerging AI tools and engines to keep the OneAI platform at the forefront of the industry. • Analyze performance bottlenecks to iterate on the efficiency of both training and inference processes.

Job Requirements

  • Bachelor’s or Master’s degree in Computer Science, Information Technology, or Engineering.
  • 3+ years of experience in applied AI, machine learning, or software engineering, with hands-on delivery of AI/ML solutions in production environments
  • Demonstrated experience designing and deploying high-performance AI infrastructure, specifically focusing on the scalability and reliability of inference and training workloads.
  • Proven track record of deploying Large Language Models (LLMs) at scale, with deep knowledge of serving engines (e.g., vLLM) and fine-tuning tools (e.g., Unsloth).
  • Experience building AI-centric platforms or toolchains that manage the model lifecycle (versioning, deployment, and discovery).
  • Experience with GPU orchestration and optimizing workloads for cloud, distributed or large-scale environments and collaborating with platform or infrastructure teams.
  • Hands-on experience with high-throughput inference engines (e.g., vLLM) and fine-tuning tools (e.g., Unsloth)
  • Proficiency in integrating with the Hugging Face ecosystem (Transformers, Hub, Datasets) for model and data management.
  • Experience implementing monitoring tools to track system-level AI metrics such as token throughput, latency, GPU utilization, and failure rates.
  • Experience designing and implementing scalable, reliable APIs for compute provisioning and workload scheduling.
  • Experience working with cloud platforms and containerized environments (e.g., OpenNebula, Kubernetes)
  • Advanced English level (B2 or higher) is required.

Benefits

  • Competitive compensation package and flexible remuneration: Meals, Transport, Nursery/Childcare
  • Customized workstation (macOS, Windows, Linux)
  • Private health insurance
  • Paid time off: Holidays, Personal Time, Sick Time, Parental leave
  • Afternoon-off working day every friday and during summer
  • Remote company with bright HQ centrally located in Madrid; offices in Boston (USA), Brussels (Belgium) and Brno (Czech Republic); and access to office space near your location when needed.
  • Healthy work-life balance: We encourage the right for Digital Disconnecting and promote harmony between employees personal and professional lives
  • Flexible hiring options: Full Time/Part Time, Employee (Spain/USA) / Contractor (other locations)

Related Categories

Related Job Pages

More Platform Engineer Jobs

Full TimeRemoteTeam 1,001-5,000Since 2008H1B Sponsor

• Lead the design and evolution of the cloud platform across AWS and GCP • Build and standardize platform capabilities (CI/CD, observability, service templates, and golden paths) • Develop and operate shared data platform services (operators in Go, Kafka, Neo4j, and related infrastructure) • Partner with product engineering, security, and SRE teams • Mentor other engineers and drive cloud cost optimization through automation

United Kingdom
MetLife logo

Software Platform Engineer I

MetLife

MetLife is a leading insurance and financial services company based in New York, New York. The company and its affiliates specialize in employee benefits and li

Platform Engineer106 days ago
Full TimeRemoteTeam 43,000Since 1868

Description and Requirements Position Summary One should have good hands-on exposure on Core Java development. One with minimal guidance should be able to design and implement complex solutions. One should be able to adopt with new technologies or technological changes in an Agile environment. Job Responsibilities - Design and develop complex cloud-based hybrid mobile applications from the scratch. - Design and develop scalable and resilient micro services - Create and configure CI/CD and build pipelines - Creation of highly reusable and reliable UI components - High levels of ownership of systems in your team - Collaborate within the team and with other stake holders. - Writing code based on widely accepted coding principles and standards. - Contribute in all phases of the product development life-cycle - High degree of professionalism, enthusiasm, autonomy, and initiative daily - Demonstrate high level of ownership, leadership, and drive in contributing innovative solutions. - Ability to work within a team environment - Experience interfacing with both external and internal customers at all levels - Demonstrated ability to analyze problems and understand the necessary components of a solution through analysis, planning, evaluation of cost/benefit, testing and reporting Knowledge, Skills and Abilities Education - Bachelor's degree in computer science, Engineering, Finance/Accounts, or related discipline Experience - 5 to 8 years of hands-on experience in Core Java, Advanced Java, Microservices Knowledge and skills (general and technical) - Java 8, Spring Framework, Spring Boot, Spring Cloud - Micro Services and Message Queues - MongoDB or other NoSQL databases - Redis, Docker, Bamboo or Jenkins - CI/CD, Gulp or Webpack, Maven or Gradle - Javascript (ES6), ReactJS or React Native or AngularJS, Node.JS About MetLife Recognized on Fortune magazine's list of the "World's Most Admired Companies" and Fortune World's 25 Best Workplaces™, MetLife, through its subsidiaries and affiliates, is one of the world's leading financial services companies; providing insurance, annuities, employee benefits and asset management to individual and institutional customers. With operations in more than 40 markets, we hold leading positions in the United States, Latin America, Asia, Europe, and the Middle East. Our purpose is simple - to help our colleagues, customers, communities, and the world at large create a more confident future. United by purpose and guided by our core values - Win Together, Do the Right Thing, Deliver Impact Over Activity, and Think Ahead - we're inspired to transform the next century in financial services. At MetLife, it's #AllTogetherPossible . Join us! #BI-Hybrid

India
Job Closed
Vancity logo

Senior Platform Engineer

Vancity

Put your money where your values are.

Platform Engineer106 days ago
Full TimeRemoteTeam 1,001-5,000Since 1946

• Implementing platform solutions that support lending workflows, leveraging SaaS and on-premises technologies • Building modular, API-first platforms that enable rapid launch of new lending products (e.g., BNPL, small business loans, green loans) • Designing and developing reusable components and APIs optimized for lending products and future Minimum Viable Products • Integrating security and compliance best practices aligned with financial regulations into platform architecture and operations • Collaborating closely with lending developers, product managers, and architecture teams to deliver scalable solutions • Automating and enhancing platform monitoring, availability, and performance for lending services using Splunk and other observability tools • Troubleshooting lending platform issues and communicating effectively with impacted business and technology teams • Supporting CI/CD pipelines for lending applications with version-controlled deployments • Mentoring and coaching junior team members • Defining design and software implementation best practices for the team

Canada
$101.2K - $125.0K / year
Job Closed
Fanatics Betting & Gaming logo

Senior Platform Engineer, Fanatics Markets - UK

Fanatics Betting & Gaming

Fanatics is building a leading global digital sports platform. We ignite the passions of global sports fans and maximize the presence and reach for our hundreds of sports partners globally by offering products and services across Fanatics Commerce, Fanatics Collectibles, and Fanatics Betting & Gaming, allowing sports fans to Buy, Collect, and Bet. Fanatics has an established database of over 100 million global sports fans. A global partner network with approximately 900 sports properties, including major national and international professional sports leagues, players associations, teams, colleges, college conferences, and retail partners. 2,500 athletes and celebrities, and 200 exclusive athletes. Over 2,000 retail locations, including its Lids retail stores. More than 22,000 employees committed to enhancing the fan experience and delighting sports fans globally.

Platform Engineer106 days ago
Full TimeRemoteTeam 10,001

About Us Fanatics is building a leading global digital sports platform. We ignite the passions of global sports fans and maximize the presence and reach for our hundreds of sports partners globally by offering products and services across Fanatics Commerce, Fanatics Collectibles, and Fanatics Betting & Gaming, allowing sports fans to Buy, Collect, and Bet. Through the Fanatics platform, sports fans can buy licensed fan gear, jerseys, lifestyle and streetwear products, headwear, and hardgoods; collect physical and digital trading cards, sports memorabilia, and other digital assets; and bet as the company builds its Sportsbook and iGaming platform. Fanatics has an established database of over 100 million global sports fans; a global partner network with approximately 900 sports properties, including major national and international professional sports leagues, players associations, teams, colleges, college conferences and retail partners, 2,500 athletes and celebrities, and 200 exclusive athletes; and over 2,000 retail locations, including its Lids retail stores. Our more than 22,000 employees are committed to relentlessly enhancing the fan experience and delighting sports fans globally. About the Team Fanatics Markets is the real-money prediction and trading app where you can invest in moments you care about. Built on a secure platform, we let users predict real-world outcomes and trade on events they actually follow - from sports and entertainment to political elections and beyond. Our mission is to redefine how fans engage with the moments and markets that matter most. We're looking for the right people to help us build the future of prediction markets. Role Overview As a Senior Platform Engineer, you will be a key driver of our core infrastructure and developer experience. You will design and implement durable platform abstractions and partner with engineering leads to ensure our systems enable rapid innovation without compromising stability or security. This role requires a balance of deep hands-on engineering and systems thinking. You will own complex features from ideation to production, championing best-in-class platform standards and mentoring mid-level engineers while remaining deeply embedded in the codebase. On the engineering side, we are pioneering the use of AI as a code collaborator. We need engineers who have moved beyond experimentation—who actively use tools like Claude Code, Cursor, or GitHub Copilot to ship production code faster while maintaining exceptional quality. Responsibilities - Implement team-level standards for AI-augmented development, utilizing prompt engineering and validation frameworks to increase velocity without sacrificing code quality. - Identify and mitigate AI-generated code pitfalls, such as security vulnerabilities or over-abstraction, and assist in tracking team productivity metrics. - Design, build, and maintain highly available cloud-native infrastructure, ensuring the platform scales reliably with our growing user base. - Build and optimize CI/CD pipelines and developer tooling to eliminate friction in the software delivery lifecycle. - Implement and maintain platform security safeguards and compliance standards to protect the integrity of the firm’s systems. - Contribute to the system observability roadmap by implementing monitoring, tracing, and alerting to ensure high operational reliability. - Partner with product and engineering teams to translate technical requirements into scalable platform solutions. - Mentor junior and mid-level engineers, lead incident response for team-owned services, and advocate for modern engineering best practices. Experience and Skills - 5 plus years of experience in platform engineering, DevOps, or backend engineering, with a track record of building scalable, cloud-native environments (AWS/Kubernetes). - Strong proficiency in using AI collaborators (Claude Code, Cursor, Copilot) to ship production code and an ability to optimize personal development workflows. - Ability to evaluate AI-generated code for security gaps or architectural "hallucinations" and implement rigorous testing strategies. - Strong hands-on experience with IaC frameworks such as Terraform, Pulumi, or CloudFormation. - High proficiency in Java and Python (Golang is a plus). You should be comfortable bridging the gap between infrastructure and application logic. - Experience building mature CI/CD pipelines and working with observability tools (e.g., Prometheus, Grafana, Datadog). - Proven ability to take ownership of complex projects in a fast-paced environment, balancing speed with architectural health. - Strong communication skills and the ability to collaborate effectively with distributed engineering and product teams. Preferred Qualifications - Experience in high-growth startup environments. - Experience with financial systems, real-money transactions, or regulatory compliance. - Background in building internal developer tools or CLI applications. - Experience leading small project groups or "squads" within a larger team. We are serious about AI-augmented development. During our technical interviews, be prepared to: - Demo your current AI-assisted workflow (e.g., how you use Cursor or Copilot). - Discuss specific examples of how you’ve used AI to solve complex technical hurdles. - Share your strategies for ensuring AI-generated code meets production-grade security and maintainability standards. Depending on the role, your interview and onboarding experience may include in-person components, such as onsite interviews or Launching into Better: LIVE—a multi-day cultural immersion in New York City for full-time, non-seasonal hires. These sessions are designed to build connection and bring our culture to life, though specific travel and participation requirements will be confirmed based on your role and location. Your recruiter will provide clear guidance at each stage of the process. For information about our benefits, please visit https://benefitsatfanatics.com/ #LI-Remote #LI-AL1

United Kingdom