AI Platform Engineer, Harness
Location
United States
Posted
3 hours ago
Salary
0
Seniority
Senior
Job Description
AI Platform Engineer, Harness
LTS
• Design, build, and maintain enterprise AI platform capabilities supporting Large Language Models (LLMs), AI agents, RAG, and Generative AI applications. • Develop reusable AI harnesses to automate testing, prompt evaluation, model benchmarking, regression testing, and quality assurance. • Build AI evaluation frameworks to measure model accuracy, retrieval quality, hallucination detection, latency, throughput, cost, and overall application performance. • Implement observability and monitoring solutions for AI applications, including telemetry, tracing, logging, dashboards, and operational metrics. • Build and maintain LLMOps pipelines supporting model deployment, versioning, evaluation, experimentation, rollback, and continuous improvement. • Design automated workflows for prompt testing, retrieval evaluation, AI system validation, and performance benchmarking. • Develop internal tools for prompt management, model experimentation, AI performance optimization, and developer productivity. • Build scalable backend services and APIs supporting AI platforms and enterprise AI integrations. • Collaborate with AI architects and engineering teams to integrate LLMs, RAG pipelines, vector databases, and agentic AI solutions into enterprise applications. • Support deployment of AI services across AWS, Azure, or Google Cloud using containerized and cloud-native architectures. • Implement CI/CD pipelines and infrastructure automation supporting enterprise AI development and deployment. • Apply security, governance, and Responsible AI controls throughout the AI development lifecycle. • Evaluate emerging AI frameworks, LLMOps technologies, evaluation methodologies, and automation tools to improve engineering productivity. • Troubleshoot production AI issues and continuously improve platform reliability, scalability, security, and user experience. • Document engineering standards, AI platform architecture, evaluation methodologies, and operational best practices.
Job Requirements
- Bachelor's degree in Computer Science, Software Engineering, Artificial Intelligence, Data Science, or a related technical field.
- 5+ years of experience in software engineering, platform engineering, backend engineering, DevOps, cloud engineering, or infrastructure engineering.
- 2+ years building or supporting Generative AI, Large Language Model (LLM), or machine learning applications.
- Strong programming experience in Python.
- Experience developing APIs, backend services, and distributed systems.
- Experience with cloud platforms including AWS, Azure, or Google Cloud Platform.
- Experience deploying applications using Docker and Kubernetes.
- Experience working with Git, CI/CD pipelines, Infrastructure as Code (IaC), and infrastructure automation.
- Strong understanding of Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), Prompt engineering, Embeddings, Vector databases, AI agents and agentic workflows
- Familiarity with AI evaluation techniques, automated testing, benchmarking, regression testing, and model validation.
- Experience building scalable, production-grade software platforms.
- Strong problem-solving, debugging, and performance optimization skills.
Benefits
- The Opportunity to support high-visibility federal missions
- A culture that values innovation, growth, and collaboration
- Access to cutting-edge tools and technologies
- Comprehensive benefits for you and your family
- A career path that rewards ambition and performance
Related Guides
Related Categories
Related Job Pages
More Platform Engineer Jobs
Platform Engineer – Security
Tiger Resourcing GroupIndependent Recruitment Agency Specialising in IT, Engineering, Defence, Security, Space Systems and ITS
• Security issue mitigation, such as CVEs and Pen test remediation • Hardening of the application, such as containers, Kubernetes clusters, and application services • Secret management and rotation • Improvements to security posture with focus on least privilege and zero trust • Monitoring of the infrastructure for incidents and control compliance • Automation of security updates, such as with AI, for issue mitigation, drift detection, compliance, and environment hardening • Conduct user access reviews and table-top exercises • Maintain and improve system and security documents such as system diagrams • Support current and new security certifications, such as SOC 2, ISO, NIST, etc. • Works closely with the technical and support teams
Software Engineer II – Data Platform
May MobilityTransforming cities through autonomous technology to create a safer, greener, more accessible world.
• Contribute to the design, implementation, and maintenance of real-time and historical data pipelines, ensuring adherence to standard procedures and best practices. • Build and maintain well-organized, documented libraries and APIs to manage, search, and analyze vehicle datasets. • Collaborate on the development of front-end interfaces to visualize data and improve fleet operations. • Partner with cross-functional autonomy and platform teams to address software needs, ensuring system compatibility and synchronization. • Independently determine technical approaches to assigned tasks, actively discussing improvements and potential alternatives with the team. • Apply rigorous testing and standard procedures to optimize data processing workflows for performance, scalability, and cost-effectiveness. • Implement data quality checks and monitoring tools, maintaining awareness of the impact of these systems on product and business outcomes. • Proactively invest in professional growth and technical mastery through research and application of industry standards.
Role Description We are looking for Platform Developers (Scala Developers) with a server-side focus to help us improve our core platform services. This is a fully remote role reporting to Engineering Management. - Developing and delivering new features for the existing platform - Implementing new third-party integrations within the platform - Collaborating with the team to migrate the platform to a new architecture and technology stack based on Scala - Building and maintaining services that support both customer-facing websites and internal administration tools, working closely with front-end developers where required - Diagnosing, troubleshooting, and resolving production issues to ensure platform stability - Participating in code reviews to maintain code quality and share knowledge across the team - Adapting to the varied and evolving challenges of working in a growing company - Demonstrating initiative and ownership, proactively identifying tasks rather than waiting for direction - Continuously seeking more efficient and effective ways of working, contributing ideas to improve processes and deliver Qualifications - Minimum of 3 years of experience with Scala - Experience with at least one additional JVM-based language - At least 1 year of experience writing complex SQL queries, beyond basic SELECT statements - Familiarity with automated testing practices and frameworks - Strong communication and interpersonal skills, with a collaborative team mindset - Ability to quickly learn new technologies and adapt in a fast-paced environment Requirements - Familiarity with Java Servlets, because there’s always legacy code - Working knowledge of functional programming and its advantages - An understanding of asynchronous/reactive programming - Exposure to integrating third-party APIs - Experience in performance profiling and tuning Java applications Benefits - A remote and flexible working schedule - Generous time off varied based on the country of residence - Discretionary annual performance bonus - Training and other learning & development opportunities to support you through your career progression - Hardware & Software allowance or work equipment is provided to make sure you have all the right tools to get the job done - Various well-being programmes and initiatives
• Design, develop and operationally manage automated, resilient, high availability, self-healing, secure platforms with native-AI capabilities for IT needs, serving both internal as well as customer business capabilities. • Develop, and manage the Observability OpenTelemetry Central Backend Stack: Grafana Enterprise, Mimir, Loki, Tempo, and Alertmanager on Kubernetes/RKE2 via Helm and GitLab CI-CD. • Build and manage iaC and CI-CD for automated provisiong and deployment, including Terraform modules for Infra/VM/storage provisioning, Ansible AWX playbooks for OS/App bootstrap, ArgoCD and Helm for Kubernetes configuration. • Develop and manage OpenTelemetry Prometheus scrape profile library including SNMP exporters, REST API exporters, and cloud provider exporters (CloudWatch, Azure Monitor, GCP) for multiple device classes. • Develop AIOps capabilities on platforms for e.g. Observability use-cases: anomaly detection integrations, event correlation rules in Alertmanager, and synthetic monitoring patterns to reduce alert noise. • Configure and maintain Zabbix auto-discovery: network range scanning, device classification, and Prometheus service discovery integration. • Build and harden Edge Stack deployments (Prometheus + OTel collector) per data center site using GitOps templates. • Integrate Alertmanager with ServiceNow: webhook routing, ticket enrichment, auto-close logic, and escalation policy configuration. • Maintain platform security: Conjur/CyberArk secret injection at runtime, mTLS between stack components, RBAC in Grafana Enterprise. • Author and maintain Grafana dashboards in JSON/GitLab — facility overview, network health, RED metrics, application telemetry. • Mentor mid-level engineers, lead code reviews, and establish engineering standards for the team. • Represent platform engineering in cross-functional architecture reviews and executive-level program updates. • Perform other duties as required and assigned.




