Axiad

Axiad is a cybersecurity company that provides enterprise-grade identity and access management (IAM) solutions, helping organizations securely manage credentials and authentication

AI-First SRE/DevOps Engineer

Location

United States

Posted

1 day ago

Salary

$120K - $160K / year

Seniority

Mid Level

No structured requirement data.

Job Description

AI-First SRE/DevOps Engineer

Axiad

Role Description Axiad is seeking a skilled AI-First SRE/DevOps Engineer with 5–8 years of hands-on infrastructure and platform engineering experience to help build and run Mesh, our Identity Visibility and Intelligence Platform (IVIP) — a cloud-native microservices platform on Kubernetes spanning human identity, non-human identity (NHI), post-quantum cryptography, and agentic AI identity risk. The ideal candidate has a builder mentality and a strong AI-First mindset: automation and AI are the default, not the afterthought, and infrastructure is something you create, not just maintain. This is a startup environment. You will own real surface area end-to-end, move fast, and ship. The role requires deep operational expertise in Kubernetes, CI/CD, and infrastructure-as-code, along with practical experience running AI/LLM systems in production. If your instinct when facing a repetitive task is to script it, agent-ify it, or delete it entirely — you'll fit right in. Responsibilities - Own reliability, observability, and delivery for a multi-tenant, cloud-native Kubernetes platform — from design through production, yours to run and yours to improve. - Build (not just operate) CI/CD pipelines, infrastructure-as-code, and GitOps-driven progressive delivery that let a small team ship many times a day, safely. - Embrace and advocate AI-First operations: automate incident response, runbooks, and remediation, and put AI agents in the loop to triage, diagnose, and propose fixes where it makes sense. Treat toil as a bug. - Build the infrastructure that AI-native features run on: inference gateways, LLM cost/latency observability, prompt/version pipelines, eval harnesses, and guardrails for agentic workloads. - Instrument everything — SLOs, error budgets, and distributed tracing across services and data pipelines. - Harden the platform: secrets management, supply-chain security, and least-privilege everywhere. - Troubleshoot and resolve production issues, leveraging AI-powered debugging and observability tooling. - Collaborate directly with product and platform engineers to translate requirements into resilient infrastructure — no throwing tickets over a wall; if you see a problem, it's yours to solve. - Mentor engineers in adopting AI-first operational practices and automation-by-default culture. Qualifications - 5–8 years of professional experience in SRE, DevOps, or platform engineering roles. - Builder mentality: you'd rather create a tool, platform, or automation than run a manual process twice. You ship things and stand behind them. - Ownership: you take problems from ambiguity to resolution without waiting for a ticket, a spec, or permission. When something you own breaks, you're the first to know and the first to act. - Strong Kubernetes operational experience — running it in production, not just deploying to it. - Demonstrable adoption of an AI-First mindset and tools (Claude Code, Cursor, or Windsurf). Daily use of at least one AI development tool is a must. - Fluency with infrastructure-as-code, GitOps, and modern CI/CD; comfortable scripting and building tooling (Go or Python preferred). - Cloud-native depth on at least one major cloud provider. - Solid observability expertise and SLO-driven operations experience. - Experience with containerization (Docker) and service mesh concepts. - Strong problem-solving skills and a collaborative mindset; excellent communication within Agile teams. - A bias for shipping — startup pace energizes you rather than stresses you. Preferred Qualifications - Experience building or operating LLM infrastructure: inference gateways, eval/observability tooling, agentic orchestration. - Data-pipeline and streaming/CDC experience. - Security or identity background; familiarity with post-quantum cryptography or supply-chain security. - Prior experience at an early-stage startup. Benefits - 120,000 - 160,000 OTE + Equity + Benefits

Related Categories

Related Job Pages

More DevOps Engineer Jobs

ContractRemoteTeam 201-500Since 2014H1B No Sponsor

• Diseñar APIs/servicios backend que conecten y sincronicen datos entre AWS y Azure. • Construir conectores entre servicios cloud-nativos (ej. S3/SQS con Blob Storage/Event Grid). • Implementar arquitecturas orientadas a eventos (colas, webhooks, pub/sub) • Elegir el stack más adecuado por integración (Python, Node.js/TS, Java, Go u otros). • Diseñar contratos de API (REST/GraphQL/gRPC): versionado, auth (OAuth2/JWT), seguridad. • Escribir código probado y listo para producción; participar en code reviews. • Usar IA generativa (Copilot, Claude u otros) para acelerar desarrollo y documentación. • Diseñar/mantener pipelines CI/CD (GitHub Actions, ArgoCD) para los servicios que desarrolla. • Gestionar infraestructura como código (Terraform) en AWS y Azure. • Desplegar y operar en Kubernetes con Helm. • Dar observabilidad (Prometheus, Grafana, Datadog); rotaciones de guardia si aplica. • Explorar AIOps: IA/ML para detección de anomalías e incidentes. • Ser referente técnico ante el cliente en desarrollo y DevOps.

Mexico
Full TimeRemoteTeam 51-200Since 2003

• The RTE leads all Program Increment planning events, coordinating capacity, sequencing, and feature prioritization across the ART in alignment with program management, customer direction and COR guidance. • The RTE maintains the Integrated Master Schedule (IMS) on a weekly basis, manages release-level Project Process Agreements (PPAs) in accordance with FDA EPLC requirements, and ensures all sprint plans reflect realistic team capacity and accurate task-level estimates. • The RTE prepares and maintains release plan summaries, JIRA/ALM release plans mapped to sprint-level effort, and sprint status dashboards for review by the FDA Government IT PM. • The RTE serves as the primary integration point across all internal and external teams with dependencies on the program. • The RTE proactively identifies cross-team blockers, facilitates resolution, and escalates impediments to executive leadership when necessary. • The RTE drives participation from all impacted teams in end-to-end integration, regression, and UAT testing cycles for every FSDX release. • The RTE facilitates all ART-level ceremonies, including PI planning, System Demos, Inspect and Adapt workshops, and Scrum of Scrums. • The RTE coaches individual Scrum Masters and team-level Agile practices, fosters a culture of inspect and adapt, and drives measurable velocity improvements across the ART. • The RTE tracks and reports sprint burn rates, planned versus actual story point delivery, and release velocity trends, escalating potential overages to the Government IT PM immediately upon detection. • The RTE supports the Project Manager in preparing bi-weekly status reports, Monthly Status Reports (MSRs), and Monthly Financial Reports (MFRs) with accurate sprint metrics, burn rates, and risk indicators. • The RTE maintains the program-level Risk Management Plan, the RACI chart, and the Change Request Log, ensuring all items are tracked through closure and reflected in subsequent governance reporting. • The RTE maintains the program risk register and leads proactive risk identification, prioritization, and mitigation planning across all workstreams. • The RTE facilitates cross-team risk reviews during weekly status meetings and ensures newly identified risks are surfaced to FDA stakeholders immediately. • The RTE facilitates the integration of agency-approved AI tools and DevSecOps practices into the ART's delivery pipeline, supporting the program's commitment to leveraging AI for development efficiency, automated testing, intelligent workflow optimization, and documentation.

Maryland
Stellar Health logo

Senior DevOps Engineer

Stellar Health

Empowering providers to deliver high-quality care through real-time notifications and meaningful incentives.

Full TimeRemoteTeam 51-200Since 2018H1B Sponsor

• Collaborate daily with fellow DevOps engineers and cross-functional engineering teams to build a scalable, high-performing foundation. • Design and maintain HIPAA-compliant, multi-account AWS infrastructure using Terraform and Terragrunt with a strict GitOps mindset. • Build resilient GitHub Actions pipelines and implement safe release strategies (e.g., Blue/Green, automated rollbacks) for Python/Django applications. • Ensure high availability, automated backups, and version upgrades for core data stores, including Postgres, ElastiCache (Redis), and OpenSearch. • Mature full-stack Datadog/CloudWatch telemetry, refine on-call workflows, lead blameless postmortems, and drive key metrics (MTTR, deployment frequency). • Enforce least-privilege access and automated security guardrails across AWS, database layers, and GitHub. • Create paved-path tooling and self-serve infrastructure that reduce friction, shorten test cycles, and lower cognitive load for product teams. • Author clear design docs, conduct thorough code reviews, mentor team members, and proactively identify operational risks before they impact production.

United States
$150K - $175K / year
Full TimeRemoteTeam 51-200H1B No Sponsor

Role Description You're probably not looking for another DevOps job. You're looking for an opportunity to build the platform that everything else depends on. At NIVA Health, we're building AI-powered healthcare solutions, automation platforms and cloud-native applications that improve how healthcare is delivered. Now we're looking for the engineer who will make all of it reliable, scalable and production-ready. If you enjoy building infrastructure instead of firefighting it, keep reading. Here's what you'll be building: - Own the cloud platform behind our engineering and AI ecosystem. - Design CI/CD pipelines. - Automate deployments. - Improve security. - Manage Kubernetes environments. - Build Infrastructure as Code. - Ensure our engineers can deploy with confidence. You'll work closely with our AI Architect and Engineering team to create an environment where great software can move from idea to production quickly and safely. You’ll probably enjoy this role if you: - Automate everything you can. - Believe infrastructure should be repeatable, secure and scalable. - Enjoy solving reliability problems before they happen. - Love cloud architecture and modern DevOps practices. - Enjoy enabling developers rather than slowing them down. - Like building platforms that people trust. Qualifications - Experience with Google Cloud Platform (preferred) - Kubernetes - Docker - Terraform - GitHub Actions - CI/CD pipelines - Infrastructure as Code - Cloud Run - Networking & Security - Monitoring & Observability - Python or Bash scripting Benefits You'll join a small engineering team where you'll build the cloud platform that supports everything we create—from AI systems to internal healthcare applications. This isn't a maintenance role. You'll help shape how engineering operates for years to come.

India