Bright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.
AI Platform Engineer
Location
United States
Posted
3 days ago
Salary
$100K - $150K / year
Seniority
Mid Level
Job Description
AI Platform Engineer
Bright Vision Technologies
Role Description We are seeking an AI Platform Engineer to design, build, and operate high-performance, highly reliable inference platforms for serving large machine learning models in production. The role focuses on the systems engineering side of AI deployment, including: - Request routing - Batching - Caching - Autoscaling - GPU utilization - End-to-end observability across diverse model workloads The ideal candidate brings strong distributed systems and performance engineering expertise, has shipped serving systems at scale, and understands the trade-offs between latency, throughput, cost, and quality in ML serving. Qualifications - Bachelor’s or Master’s degree in Computer Science or a related field. - Six or more years of experience in distributed systems, infrastructure, or ML platform engineering. - Strong proficiency in Python and a systems language such as Go, Rust, or C++. - Deep experience operating high-throughput, low-latency services in production. - Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM. - Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization. - Familiarity with Kubernetes, autoscaling, and modern cloud platforms. - Experience with observability stacks including metrics, tracing, and structured logging. - Solid grounding in performance engineering and capacity planning. - Strong communication and incident response skills. Requirements - Design and operate model serving platforms supporting diverse workloads including LLMs, vision models, and recommendation systems. - Optimize inference performance using continuous batching, paged attention, speculative decoding, and request multiplexing. - Implement multi-tenant routing, rate limiting, and quality-of-service policies across model endpoints. - Build autoscaling and capacity management systems that balance latency, throughput, and cost. - Tune GPU utilization, memory management, and KV cache strategies for LLM serving workloads. - Integrate model serving with API gateways, identity systems, and observability platforms. - Implement caching, prompt deduplication, and response reuse strategies where appropriate. - Drive end-to-end observability including latency histograms, queue dynamics, GPU utilization, and error tracking. - Develop deployment workflows including canary releases, shadow testing, and automated rollback. - Operate incident response for high-availability AI services and drive durable reliability improvements. - Collaborate with ML and product teams to support new model releases and capability rollouts. - Implement security controls including request signing, content filtering, and abuse detection at the serving layer. - Document operational procedures, performance characteristics, and tuning guidance for internal teams. - Stay current with AI serving research and translate advances into production capabilities. Preferred Qualifications - Open-source contributions to model serving infrastructure. - Experience with multi-region or globally distributed AI serving. - Familiarity with model quantization, distillation, and compression techniques. - Exposure to FinOps for AI workloads and cost-efficient serving design. - Experience supporting external-facing AI APIs at scale. How to Apply Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3545. Learn more about Bright Vision Technologies at www.bvteck.com . Equal Employment Opportunity (EEO) Statement Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall. BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
Related Guides
Related Categories
Related Job Pages
More Platform Engineer Jobs
Kafka Platform Engineer
Bright Vision TechnologiesBright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.
Role Description We are seeking an experienced Kafka Engineer to architect, deploy, and operate large-scale Apache Kafka and Confluent platform environments supporting mission-critical event-driven workloads. In this role you will own the Kafka platform end-to-end, including: - Cluster sizing - Configuration - Security - Automation - Observability - Developer enablement The ideal candidate will combine deep Kafka internals knowledge with strong DevOps and SRE practices, and will partner with application teams to deliver a reliable, performant, and developer-friendly streaming platform. You will work closely with cross-functional partners — product, design, engineering, operations, and business stakeholders — to translate ambiguous requirements into well-engineered solutions. You will be expected to raise the bar through code review, design review, and mentorship of more junior engineers. The successful candidate brings strong engineering discipline, a clear communication style, and a track record of shipping meaningful work that holds up well in production. Qualifications - Bachelor’s degree in Computer Science, Engineering, or a related technical discipline. - Five or more years of experience operating Apache Kafka or Confluent Platform in production. - Deep, hands-on knowledge of Kafka internals (partitions, replication, ISRs, consumer groups). - Strong experience with Kafka security (SASL, mTLS, ACLs, RBAC). - Hands-on experience with Kafka Connect, Schema Registry, and either Kafka Streams or ksqlDB. - Experience with HA/DR strategies for Kafka. - Strong scripting skills in Python, Bash, or Go. - Hands-on experience with infrastructure-as-code (Terraform, Ansible). - Working knowledge of observability tooling for Kafka. - Excellent troubleshooting, communication, and documentation skills. Requirements - Architect, deploy, and operate large-scale Apache Kafka or Confluent Platform clusters across on-prem and cloud environments. - Design partitioning, replication, and topic strategies that balance throughput, durability, and operational simplicity. - Implement strong security on Kafka clusters using SASL, mTLS, ACLs, RBAC, and integration with corporate IdPs. - Operate Schema Registry, Kafka Connect, KSQL/ksqlDB, and Kafka Streams in production. - Build and operate Kafka Connect pipelines integrating sources and sinks across enterprise systems. - Design HA/DR strategies for Kafka, including MirrorMaker 2, Cluster Linking, and multi-region active-active patterns. - Build CI/CD pipelines for Kafka topic, ACL, and connector configurations using GitOps patterns. - Implement comprehensive observability using Prometheus, Grafana, Datadog, or Confluent Control Center. - Drive Kafka cost and capacity optimization through right-sizing and storage tiering. - Onboard application teams to Kafka with clear patterns, templates, and best practices. - Lead incident response and post-incident reviews for streaming workloads, applying disciplined engineering practices and partnering closely with stakeholders to ensure outcomes are durable, well-documented, and aligned with broader team and platform standards. - Mentor and coach junior and mid-level engineers through code review, design review, pair programming, and structured knowledge sharing, helping the broader team grow in technical maturity and confidence over time. - Maintain comprehensive, current technical documentation — including architecture diagrams, design decisions, configuration references, runbooks, and operational procedures — so that the system remains supportable, auditable, and easy to onboard new engineers onto over time. - Continuously evaluate emerging streaming technologies (Pulsar, Redpanda, AWS MSK, Azure Event Hubs). Benefits - Competitive salary range: $100,000–$150,000 Annually - Full-time, Direct W2 position - 100% Remote (U.S.) How to Apply Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3545. Learn more about Bright Vision Technologies at www.bvteck.com . Equal Employment Opportunity (EEO) Statement Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall. BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
• Own the ROS2 communications and middleware layer across drone and MHEV platforms, including the Zenoh-based comms layer, behavior trees, and system integration • Maintain and evolve the Docker containers our robotics software runs in, keeping them lean, reliable, and consistently deployable to the field • Build diagnostics and observability tooling that surfaces the health of the communications layer and running services, from low-level logging through dashboard visibility • Define and manage deployment processes for getting software changes out to field hardware reliably and safely • Develop the roadmap for converging our drone and MHEV architectures into a unified platform over time, and drive the execution of that convergence • Support both Autonomy team pods across drone and MHEV, flexing between their needs as priorities shift
Senior Platform Engineer
LifeMDLifeMD (Nasdaq: LFMD) is a rapidly growing direct-to-consumer telemedicine company.
• Design, build, and maintain AWS infrastructure as code using Terraform and Pulumi, with an emphasis on reusable modules, safe change management, and multi-account/multi-environment patterns • Operate and evolve container platforms across ECS and EC2, including cluster architecture, autoscaling, networking, and cost optimization • Drive our migration to microservices: decompose existing services, establish golden paths for service scaffolding, CI/CD, and deployment, and partner with product engineering teams through the transition • Build internal tooling and automation in a general-purpose language (Python, Go, or TypeScript) — this role writes real software, not just configuration • Own and advance our observability practice in Datadog: SLOs, dashboards, alerting strategy, distributed tracing, and driving down noise so on-call is humane • Lead incident response and post-incident reviews; turn findings into durable platform improvements • Mentor engineers, review designs and code, and set technical direction for platform standards across teams • Champion security, compliance, and reliability best practices appropriate to a healthcare environment
Senior Software Engineer, Platform
ECPClinical and operations software solutions for assisted living providers
• Build the tool surface an agent acts through, the retrieval and context layer that grounds it in the right clinical and operational data, the serving path product teams call to reach a model, and the evaluation that catches a regression before a customer does. • Take over shared components that every team depends on, providing interfaces worth depending on, and clarify the seams between the legacy platform and the newer services. • Enhance the messaging backbone; take the event bus the rest of the way, adding contract testing for service verification.


