Your Full Stack Technology Partner
DevOps Architect – Azure, IaC
Location
United States
Posted
61 days ago
Salary
$160K - $190K / year
Seniority
Lead
Job Description
DevOps Architect – Azure, IaC
Emergent Software
• Own the technical direction of Emergent’s DevOps practice, including the IaC standards, pipeline patterns, code review expectations, and reference architectures that the team applies across client engagements. • Author and maintain the internal documentation, templates, and modules that the team relies on day to day. This is a core part of the role, not a side task. • Mentor a team of engineers from mixed backgrounds, including software developers growing into cloud roles and traditional cloud engineers building DevOps fluency. Pair, review pull requests, run office hours, and raise the technical bar across the group. • Lead the most complex DevOps engagements end to end, from discovery through delivery, with a focus on environments where scale, regulatory requirements, or legacy constraints demand strong technical judgment. This includes grooming backlog work and feature & task assignments to project team as well as influence technical direction across practice areas. • Represent the DevOps practice in pre-sales. Partner with sales and account teams to scope work, lead technical discovery sessions, contribute to proposals, and build client confidence in Emergent’s approach. • Continuously evolve Emergent’s DevOps practice by evaluating new tools, patterns, and approaches, and bringing the right ones into the team’s standard playbook. • Diagnose and improve existing client systems that may be partially implemented or inconsistently managed, balancing pragmatic delivery with long-term maintainability. • Accurately track time and work performed to support client billing, project planning, and overall project health, as part of a professional services delivery model. • Strong time management skills to be able to prioritize workload, meet deadlines, and provide technical support and guidance for other DevOps focused team members. • Share your knowledge at regular talk shop and lunch & learn sessions to help build a stronger team.
Job Requirements
- 8+ years of hands-on experience with cloud infrastructure, with deep expertise in Microsoft Azure across compute, storage, and governance
- Deep networking expertise in Azure, including designing topologies from the ground up (segmentation strategy, hub-and-spoke or virtual WAN, cross-region connectivity, on-premises integration), as well as working fluently with virtual networks, ExpressRoute, VPN Gateways, and private endpoints.
- Comfortable diagnosing connectivity issues across hybrid environments.
- Strong identity and security fundamentals, including Entra ID, RBAC, managed identities, and secure access patterns
- Advanced experience authoring production Infrastructure as Code, including modular Terraform or Bicep with strong opinions on idempotency, reuse, environment parity, and safe change management
- Proven track record designing and operating CI/CD pipelines at scale using Azure DevOps, GitHub Actions, or comparable tooling
- Experience designing disaster recovery and high availability strategies in Azure, including defining RTO and RPO targets, multi-region patterns, and failover testing
- Experience designing for cost efficiency at scale, including tagging strategies, reservation and savings plan planning, and ongoing FinOps practices
- Demonstrated experience writing and maintaining technical standards and internal documentation that other engineers actively use. Be prepared to share examples.
- Experience mentoring engineers, including those transitioning into cloud or DevOps from adjacent backgrounds
- Comfortable leading client-facing technical conversations, including pre-sales discovery and architecture reviews
- Previous consulting or professional services experience
Benefits
- Medical Insurance: up to 84% of your monthly medical premium (HSA options available)
- HSA Contribution: up to $200/month
- Dental & Vision Insurance: up to 50% of your monthly dental and vision premium costs
- 401(k) plan: company match up to 4% of salary
- Profit sharing bonus: up to 15% of salary paid quarterly
- Extra compensation: get paid extra for work over 40 hours/week
- Employee referral & customer referral bonuses
- Flex Spending Account (FSA) for Dependent Care & Healthcare Costs
- Short Term Disability: $500/week for 12 weeks
- Long Term Disability: up to $6,000/month
- Group term life and AD&D insurance: $50k
- PTO, standard holidays, 2 floating holidays
- Paid parental leave
- Staff development program: 100 hours/year plus training costs
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
SRE / DevOps / Cloud Platform Engineer
Built TechnologiesAn award-winning FinTech startup, Built Technologies is on a mission to power smarter construction by transforming the construction finance ecosystem. Past flexible jobs at Built T
Role Description Built is hiring a Mid or Senior Cloud Platform Engineer to join our Cloud Platform Team in Mendoza, Argentina. This role is for the hands-on engineer who keeps AWS environments running smoothly, improves Terraform every week, and is the person product teams want in the room when an infrastructure problem needs a practical, durable fix. You will help operate and improve the cloud platform that product engineering teams build on: - Triaging AWS issues - Hardening infrastructure as code - Supporting production databases - Improving delivery pipelines - Raising the bar on reliability This is a strong fit for someone who enjoys digging into AWS console errors, tracing IAM denials, untangling Terraform state, and tuning the database or deployment path behind a production slowdown. At the L3 end of the band, this person will increasingly lead well-scoped platform initiatives while staying close to the operational craft. What You'll Do - Solve AWS problems across the services Built depends on, including EC2, EKS, ECS, RDS, S3, IAM, VPC, Route 53, CloudWatch, KMS, Secrets Manager, and the surrounding service ecosystem. - Author, review, and improve Terraform across Built's AWS estate by extending modules, resolving drift, managing state safely, and lifting the quality of reusable infrastructure patterns. - Support production databases as a hands-on partner to engineering teams, including: - Schema reviews - Index tuning - Slow-query investigations - Backup and restore validation - Replication or failover checks - Capacity planning - Migration support across RDS PostgreSQL, MySQL, Aurora, DynamoDB and related data stores. - Own practical improvements to GitHub Actions CI/CD workflows, making day-to-day builds and deploys faster, safer, and easier for product teams to debug. - Improve observability by building useful Datadog or CloudWatch dashboards, writing meaningful monitors, tuning noisy alerts, and helping teams instrument services with metrics, logs, traces, and basic SLO thinking. - Participate in incident response and on-call for AWS-related failures: investigate quickly, communicate clearly, write clean postmortems, and turn findings into Terraform-backed or operational fixes. - Partner with product engineering teams on infrastructure design by reviewing TRDs, recommending AWS patterns, and helping teams choose services that fit the problem, reliability needs, and operational cost. - Help raise the standard for security and compliance through least-privilege IAM, secret hygiene, patching cadence, vulnerability remediation, and data-protection practices appropriate for a regulated fintech environment. - Grow into larger ownership over time by leading bounded reliability, infrastructure, database, or developer-experience initiatives from problem definition through delivery. Qualifications - Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience. - 3-7 years of hands-on experience in DevOps, platform engineering, SRE, or cloud infrastructure in a production environment. - Strong AWS troubleshooting skills, including comfort with CloudTrail, CloudWatch Logs, VPC flow logs, IAM policy analysis, AWS Support tooling, and root-cause work across compute, networking, storage, and identity. - Solid working knowledge of core AWS services, including EC2, RDS, S3, IAM, VPC networking, Security Groups, Route 53, KMS, Secrets Manager, CloudWatch, and at least one of EKS or ECS. - Hands-on Terraform experience writing and refactoring modules, working with remote state, managing environments or workspaces, and resolving drift and provider issues with care. - Experience building or maintaining CI/CD pipelines in GitHub Actions and shipping production changes safely through reusable workflows, secrets management, and failed-run debugging. - Working knowledge of at least one observability platform such as Datadog, Grafana, or CloudWatch, including dashboards, monitors, alert quality, and basic SLO concepts. - Ability to script in Python, Go, Node or Bash for automation, operational tooling, and cloud-platform glue work. - Clear written communication, with the ability to explain an AWS issue, Terraform change, database finding, or incident timeline to engineers outside the platform team. Requirements - Production DBA or database engineering experience, especially with MySQL and/or Postgres: query tuning, EXPLAIN plan analysis, indexing strategy, backup and restore, replication, failover testing, schema migrations, and partitioning. - Experience operating DynamoDB in production, including data modeling, partition and sort key strategy, GSIs and LSIs, capacity mode selection, streams, DAX, and tuning hot-partition or throttling issues. - Experience supporting database migrations and major version upgrades on RDS or Aurora. - Familiarity with data security and compliance practices around PII, PCI, and regulated production environments. Benefits - The rare opportunity to radically disrupt a $1.5T industry. - Competitive benefits including: uncapped vacation [US ONLY], health, dental & vision insurance. - Robust compensation package, including equity in the form of stock options. - Learning Grant program to support ongoing professional development. - 401k with match and expedited vesting [US ONLY]. Operational Details - Team size: You will join a team of 6+ platform engineers, with senior ICs on the team to learn from and partner with. - On-call: Rotations are a week at a time, with frequency based on team size, expected roughly once every two months. The team works continuously to reduce incident volume so on-call stays manageable and low-noise. - Location: Mendoza, Argentina, with flexible working hours.
• Owning the SRE lifecycle for NodeBalancer and Network Load Balancer — from design reviews and pre-rollout readiness assessments through production sign-off and ongoing reliability management • Designing and implementing SLO/SLI frameworks that reflect true customer experience for L4 and L7 load balancing services, and driving action when error budgets are at risk • Building and maintaining observability pipelines for NB/NLB infrastructure, including Prometheus metrics from load balancing components and system-level sources, and Grafana dashboards that enable rapid incident triage • Leading technical incident response for complex NB/NLB failures — BGP/VIP issues, failover failures, data plane degradations, and configuration problems — acting as the technical commander and driving root cause analysis and preventive follow-through • Developing and automating safe deployment workflows for phased NB/NLB releases, including bake period monitoring, feature flag management, and GO/NO-GO validation across global datacenter rollouts • Reviewing design documents, product requirement Documents and producing actionable SRE input on operational risks, capacity implications, Day-2 concerns, and product strategy gaps • Building automation and tooling using Python or Go that reduces operational toil and improves team-wide operational capability • Mentoring SRE II engineers on the NB team, providing hands-on technical guidance, code/config reviews, and raising the bar for the team's SRE practice • Participating in an on-call rotation for NB/NLB production systems, responding to incidents and driving resolution for customer-facing load balancing infrastructure • Participate in a scheduled, daytime-only on-call rotation to spearhead technical incident response and resolve complex NB/NLB failures.
• Own developer experience end-to-end: deployments, staging, dev environments, CI/CD pipelines. • Find what’s painful, fix it, then make it stay fixed. • Be a force multiplier for the engineering team. • Maintain, improve, and sometimes rebuild our CI/CD pipelines. • Work alongside our existing DevOps engineer on the broader infra picture: cluster management, observability, secrets, networking, cost. • Build and maintain templates and tooling for spinning up new services and dev environments. • Recommend and implement efficiency wins where you see them — right-sizing, autoscaling, scheduling, GPU utilization. • Write real code when the situation calls for it. • Care about the small stuff: clear runbooks, sane defaults, good error messages, deploys that fail safely.
Sr Staff Site Reliability Engineer - Veza
ServiceNowAs the AI platform for business transformation, we're putting AI to work across organizations — freeing people for work that matters. Making old tech work with new tech. Reaching across departments, from the front office to the back office and every office in between. Our ambition? To become the AI defining enterprise software company of the 21st century (or "AI DESCO21C," as we like to call it). With more than 8,400+ customers, we serve approximately 90% of the Fortune 500®, and we're proud to be a Fortune 100 Best Companies to Work For® and World's Most Admired Companies™. Explore your future career with us, visit www.careers.servicenow.com From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.
Company Description Veza is the pioneer in identity security, purpose-built to answer the fundamental question enterprises face: who can and should take what action on what data. Veza's Access Graph platform maps an organization's entire identity ecosystem across users, groups, roles, policies, permissions, and resources providing deep visibility and control over human, non-human, and agentic identities across SaaS, cloud, on-prem, and custom applications. With over 30 billion access permissions under management, global enterprises including Blackstone, Expedia, and Wynn Resorts trust Veza to manage privileged access monitoring, non-human identity security, access entitlement management, and next-generation identity governance. Founded in 2020 and headquartered in Redwood City, California, Veza is now part of the ServiceNow family, with the acquisition closing in March 2026. The combination brings together Veza's AI-native Access Graph with ServiceNow's AI Control Tower and agentic workflows, enabling organizations to enforce end-to-end identity security rooted in the principle of least privilege across applications, data, cloud environments, and AI agents. For engineers joining Veza today, this means the scale and resources of an enterprise platform company, with the product velocity and mission-driven focus of a security innovator at a pivotal moment in the industry. Job Description 3 days onsite at the Redwood City office We are seeking an exceptional Sr Staff Site Reliability Engineer to lead critical infrastructure initiatives and drive innovation across our organization. You'll architect scalable solutions, navigate complex technical challenges independently, and deliver results under tight deadlines in a fast-paced environment. You'll work cross-functionally alongside builders who have helped shape the success of companies such as Google, Okta, AWS, and Snowflake. We are building the next generation identity security platform for the multi-cloud era - will you join us? You will: Strategic Leadership & Technical Execution - Lead enterprise-wide reliability and infrastructure projects across multiple teams with high autonomy - Navigate ambiguous problem spaces and deliver innovative solutions under tight deadlines - Architect and deploy solutions for Cloud Prem and SaaS customers at scale - Drive technical innovation and establish SRE best practices across the organization - Respond to critical incidents, lead root cause analysis, and implement long-term resolutions - Develop automation solutions to streamline operations and reduce manual workload - Participate in on-call rotation and ensure effective incident handoff and documentation Cross-Functional Collaboration & Communication - Partner with Engineering, Product, and Customer Success teams to align reliability goals with business objectives - Communicate complex technical concepts effectively to technical and non-technical audiences, including executives - Influence technical decisions across teams through thought leadership and demonstrated expertise - Build consensus and drive adoption of new tools, processes, and architectural patterns Customer-Facing Technical Leadership - Provide tier 2/3 technical support to enterprise customers for complex troubleshooting - Work directly with customer technical teams to resolve deployment, configuration, and integration challenges - Conduct technical onboarding and provide expert guidance on platform architecture and best practices - Create customer-facing documentation, troubleshooting guides, and run-books - Lead customer calls and technical discussions as a trusted advisor - Team Development - Mentor SRE and engineering team members, elevating technical capabilities - Foster a culture of reliability, operational excellence, and continuous improvement Qualifications Required Experience - BS degree in Computer Science or related field (or equivalent practical experience) - 7+ years in Site Reliability Engineering, DevOps, or Infrastructure Engineering - Proven track record leading large-scale, cross-team infrastructure projects from conception to production - Demonstrated ability to work autonomously on ambiguous projects with tight deadlines Technical Expertise - 5+ years with AWS (VPC, EC2, RDS, EKS, CloudFormation) and cloud automation - Expert-level experience with Kubernetes, Helm, Linux, and Terraform - Strong experience with GitOps model, distributed version control, and CI/CD pipelines - Proficiency with monitoring tools (Prometheus, Grafana, DataDog) - Strong programming/scripting skills (Python, Go, Bash) for automation - Deep understanding of distributed systems, microservices, and reliability patterns - Experience with Bazel and CueLang a plus - Hands-on experience with at least one major compliance framework (SOC 1/2, ISO 27001, FedRAMP Moderate/High) through an audit cycle Leadership & Communication - Exceptional ability to articulate complex technical concepts to diverse audiences - Track record of driving technical change across organizational boundaries - Successfully delivered multiple complex projects under tight deadlines - Strong customer service orientation with patience and empathy Work Style: - Thrives in ambiguous environments and makes progress without perfect information - Hands-on, "can do" attitude with bias for action - Low ego and high intellectual curiosity - Comfortable working across time zones - Self-motivated with strong ownership mentality For positions in this location, we offer a base pay of $165,500 - $289,600, plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location. Additional Information Work Personas We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here . To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service. Equal Opportunity Employer ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, national origin or nationality, ancestry, age, disability, gender identity or expression, marital status, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements. Accommodations We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact globaltalentss@servicenow.com for assistance. Export Control Regulations For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. From Fortune. ©2025 Fortune Media IP Limited. All rights reserved. Used under license.




