We manage your technology and mitigate your cyber risk - so you don’t have to.
Principal Software Engineer, MDR Platform
Location
United States
Posted
3 days ago
Salary
0
Seniority
Lead
Job Description
Principal Software Engineer, MDR Platform
Harbor IT
• Co-own and maintain multiple Golang applications forming the core of our MDR platform • Maintain and enhance a high-performance log analytics engine analyzing events from thousands of sources across hundreds of customers • Maintain and enhance programs that transform engine output into actionable intelligence for SOC analysts • Maintain and enhance a multi-tenant cluster of servers that receive, buffer, and feed syslog-formatted logs to the engine • Maintain and enhance a cross-platform syslog agent that feeds syslog-formatted logs to the engine • Build and maintain a cross-platform security agent that provides visibility into and secures customer endpoints and servers • Make architectural decisions for various applications supporting the business • Influence technical design discussions and code reviews • Mentor and guide other development team members • Facilitate knowledge transfer during any transitionary periods; assisting with training and hiring as needed • Develop and maintain internal SOPs and best practices for software development • Collaborate with cross-functional teams to define, design, and ship new features
Job Requirements
- 6+ years of software engineering experience with at least 4+ years focused on Go development
- Bachelor's degree in computer science or equivalent practical experience
- Portfolio of delivered production systems and/or contributions to open-source projects
- Idiomatic fluency in Golang and deep familiarity with the standard library and package ecosystem
- Expertise in managing goroutine lifecycles and channel-based communication
- Mandatory use of context for deadline management, timeouts, and structured cancellation
- Implementation of thread-safe data structures and methods to manage shared state efficiently
- Mastery of Go paradigms and constructs, including interfaces and generics to build modular code
- Experience implementing worker pool patterns to manage resource-intensive tasks
- Deep understanding of memory management, including minimizing heap allocation, runtime profiling to identify memory leaks, and pre-allocating buffer memory
- Experience using Github Actions for continuous deployment of Docker containers on cloud infrastructure, i.e. AWS ECS or EC2 or equivalents
- Deep proficiency in interfacing with Redis, OpenSearch or similar, and SQL databases; optimizing queries for performance and atomicity
- Robust understanding of networking protocols, TLS, and firewalls, with practical experience implementing best practices at the application level
- Proficiency with Git version control and CI/CD pipelines
- Experience with automated testing, infrastructure monitoring, and observability practices
- Experience leveraging AI assistant tools for software development, such as Claude Code
Benefits
- Comprehensive health benefits
- Matching 401k
- Reimbursement for approved tuition, certifications, conference attendance, and more
- Unlimited PTO
Related Guides
Related Categories
Related Job Pages
More Platform Engineer Jobs
Senior Data Engineer, Platform
DraftKings Inc.Defining what it means to build and deliver the most extraordinary sports & entertainment experiences.The Crown is Yours
• Demonstrate leadership and ownership of the platform to deliver services for projects and users. • Provide infrastructure guidance of data platform capabilities to accommodate business/technical use cases. • Leverage your strong communication skills to keep users informed and provide excellent quality of service. • Automate and manage provisioning needs, such as Snowflake storage and compute, Role Based Access Control model, and permissions. • Configure and manage monitoring/alerting around replication latency, performance (cluster & query), and Airflow. • Coordinate and collaborate with dependent infrastructure and AWS services to implement Snowflake integration with services, such as S3, IAM, SSO, etc. • Provide technical expertise, troubleshooting, and support for change management, governance compliance, internal audits, and remediations.
• Contribute to building and improving platform services, writing code, reviewing changes, and supporting deployments. • Work alongside senior engineers to design and implement solutions that improve scalability, reliability, and developer experience. • Help maintain and operate core systems such as CI/CD pipelines, cloud infrastructure, and monitoring. • Participate in incident response and retrospectives, contributing to learning and improvement. • Collaborate with other engineers to understand needs and translate them into practical platform solutions. • Take ownership of tasks and smaller projects, while developing towards owning larger areas of the platform.
Staff Platform Engineer
PostscriptSMS marketing platform for ecommerce companies. Helping Shopify stores drive 30x ROI with text message marketing.
• Own infra topology, scaling and blast-radius boundaries. • Build the validation gates, idempotency checks, drift detectors, and runbooks. • Diagnose novel failures such as consumer lag, DLQs, materialized-view chains. • Own the 500K events/sec target: load testing, tuning, Database sizing, SLOs and alerting. • Own IAM/IRSA, secret rotation, and the auth chain. • Keep skills, runbooks, and agent context in sync with reality. • Lead incidents and capture every fix as a new runbook plus regression test.
AI Platform Engineer
Bright Vision TechnologiesBright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.
Role Description We are seeking an AI Platform Engineer to design, build, and operate high-performance, highly reliable inference platforms for serving large machine learning models in production. The role focuses on the systems engineering side of AI deployment, including: - Request routing - Batching - Caching - Autoscaling - GPU utilization - End-to-end observability across diverse model workloads The ideal candidate brings strong distributed systems and performance engineering expertise, has shipped serving systems at scale, and understands the trade-offs between latency, throughput, cost, and quality in ML serving. Qualifications - Bachelor’s or Master’s degree in Computer Science or a related field. - Six or more years of experience in distributed systems, infrastructure, or ML platform engineering. - Strong proficiency in Python and a systems language such as Go, Rust, or C++. - Deep experience operating high-throughput, low-latency services in production. - Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM. - Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization. - Familiarity with Kubernetes, autoscaling, and modern cloud platforms. - Experience with observability stacks including metrics, tracing, and structured logging. - Solid grounding in performance engineering and capacity planning. - Strong communication and incident response skills. Requirements - Design and operate model serving platforms supporting diverse workloads including LLMs, vision models, and recommendation systems. - Optimize inference performance using continuous batching, paged attention, speculative decoding, and request multiplexing. - Implement multi-tenant routing, rate limiting, and quality-of-service policies across model endpoints. - Build autoscaling and capacity management systems that balance latency, throughput, and cost. - Tune GPU utilization, memory management, and KV cache strategies for LLM serving workloads. - Integrate model serving with API gateways, identity systems, and observability platforms. - Implement caching, prompt deduplication, and response reuse strategies where appropriate. - Drive end-to-end observability including latency histograms, queue dynamics, GPU utilization, and error tracking. - Develop deployment workflows including canary releases, shadow testing, and automated rollback. - Operate incident response for high-availability AI services and drive durable reliability improvements. - Collaborate with ML and product teams to support new model releases and capability rollouts. - Implement security controls including request signing, content filtering, and abuse detection at the serving layer. - Document operational procedures, performance characteristics, and tuning guidance for internal teams. - Stay current with AI serving research and translate advances into production capabilities. Preferred Qualifications - Open-source contributions to model serving infrastructure. - Experience with multi-region or globally distributed AI serving. - Familiarity with model quantization, distillation, and compression techniques. - Exposure to FinOps for AI workloads and cost-efficient serving design. - Experience supporting external-facing AI APIs at scale. How to Apply Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3545. Learn more about Bright Vision Technologies at www.bvteck.com . Equal Employment Opportunity (EEO) Statement Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall. BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.



