• Design, build, and maintain enterprise AI platform capabilities supporting Large Language Models (LLMs), AI agents, RAG, and Generative AI applications.
• Develop reusable AI harnesses to automate testing, prompt evaluation, model benchmarking, regression testing, and quality assurance.
• Build AI evaluation frameworks to measure model accuracy, retrieval quality, hallucination detection, latency, throughput, cost, and overall application performance.
• Implement observability and monitoring solutions for AI applications, including telemetry, tracing, logging, dashboards, and operational metrics.
• Build and maintain LLMOps pipelines supporting model deployment, versioning, evaluation, experimentation, rollback, and continuous improvement.
• Design automated workflows for prompt testing, retrieval evaluation, AI system validation, and performance benchmarking.
• Develop internal tools for prompt management, model experimentation, AI performance optimization, and developer productivity.
• Build scalable backend services and APIs supporting AI platforms and enterprise AI integrations.
• Collaborate with AI architects and engineering teams to integrate LLMs, RAG pipelines, vector databases, and agentic AI solutions into enterprise applications.
• Support deployment of AI services across AWS, Azure, or Google Cloud using containerized and cloud-native architectures.
• Implement CI/CD pipelines and infrastructure automation supporting enterprise AI development and deployment.
• Apply security, governance, and Responsible AI controls throughout the AI development lifecycle.
• Evaluate emerging AI frameworks, LLMOps technologies, evaluation methodologies, and automation tools to improve engineering productivity.
• Troubleshoot production AI issues and continuously improve platform reliability, scalability, security, and user experience.
• Document engineering standards, AI platform architecture, evaluation methodologies, and operational best practices.