Senior Site Reliability Engineer
Location
Poland
Posted
123 days ago
Salary
0
Seniority
Senior
No structured requirement data.
Job Description
Senior Site Reliability Engineer
Akamai Technologies
Do you enjoy solving complex reliability challenges for cutting-edge technology? Do you have a passion for automation and building systems that scale? Join the Akamai Inference Cloud Team The Akamai Inference Cloud team is part of Akamai's Cloud Technology Group. We design, implement, deploy and operate AI platforms that enable customers to run inference models and developers to create AI applications with unmatched performance, compliance, and economics. Partner with the best As a Senior SRE, responsibilities include owning reliability workstreams for Akamai's serverless inference platform, building automation and tooling, and contributing to architecture and operational decisions. Opportunities exist to take ownership of critical reliability problems end-to-end, partner with product engineering teams, and develop expertise in GPU infrastructure, Kubernetes at scale, and AI inference workloads. As a Site Reliability Engineer, you will be responsible for: - Building and maintaining observability for AI workloads, including telemetry, dashboards, alerts, SLO/SLI tracking, and driving improvements when targets are missed - Writing automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response - Integrating AI workloads into Akamai's existing incident management processes, building runbooks, participating in on-call rotations, and conducting blameless post-mortems - Building and maintaining CI/CD integrations, deployment safety checks, and rollback automation - Collaborating with product engineering teams to improve reliability, contribute to architecture decisions, and ensure operational readiness for product releases - Contributing to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure Do what you love To be successful in this role you will: - Demonstrate expertise in SRE, infrastructure, or platform engineering, managing large-scale distributed systems with extensive operational experience. - Demonstrate expertise in Kubernetes and large-scale containerization systems. - Define SLOs and work with observability tools like Prometheus, Grafana, and distributed tracing to enhance system monitoring. - Demonstrate proficiency in Python or Go for automation, CI/CD pipelines, deployment safety, and infrastructure-as-code like Terraform. - Interest in or experience with AI/ML infrastructure, model serving, or GPU workloads - Resolve issues independently while maintaining accountability throughout the process. - Demonstrate accountability for reliability, develop automation and monitoring, and collaborate effectively with an engineering team unfamiliar with SRE practices. Work in a way that works for you FlexBase, Akamai's Global Flexible Working Program, is based on the principles that are helping us create the best workplace in the world. When our colleagues said that flexible working was important to them, we listened. We also know flexible working is important to many of the incredible people considering joining Akamai. FlexBase, gives 95% of employees the choice to work from their home, their office, or both (in the country advertised). This permanent workplace flexibility program is consistent and fair globally, to help us find incredible talent, virtually anywhere. We are happy to discuss working options for this role and encourage you to speak with your recruiter in more detail when you apply. Learn what makes Akamai a great place to work Connect with us on social and see what life at Akamai is like! We power and protect life online, by solving the toughest challenges, together. At Akamai, we're curious, innovative, collaborative and tenacious. We celebrate diversity of thought and we hold an unwavering belief that we can make a meaningful difference. Our teams use their global perspectives to put customers at the forefront of everything they do, so if you are people-centric, you'll thrive here. Working for you At Akamai, we will provide you with opportunities to grow, flourish, and achieve great things. Our benefit options are designed to meet your individual needs for today and in the future. We provide benefits surrounding all aspects of your life: - Your health - Your finances - Your family - Your time at work - Your time pursuing other endeavors Our benefit plan options are designed to meet your individual needs and budget, both today and in the future. About us Akamai powers and protects life online. Leading companies worldwide choose Akamai to build, deliver, and secure their digital experiences helping billions of people live, work, and play every day. With the world's most distributed compute platform from cloud to edge we make it easy for customers to develop and run applications, while we keep experiences closer to users and threats farther away. Join us Are you seeking an opportunity to make a real difference in a company with a global reach and exciting services and clients? Come join us and grow with a team of people who will energize and inspire you! #LI-Remote
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Senior GCP Data Specialist
AvengaA global IT engineering and consulting company specializing in custom software development.
This is us At Avenga, we believe that human creativity empowers technology that matters. Operating globally, our 6000+ specialists provide a full spectrum of services, including business and tech advisory, enterprise solutions, CX, UX and Ul design, managed services, product development, and software development. This is the job In Bulgaria, we are actively seeking for an experienced GCP Data Specialist to strengthen our team dedicated to to build and further develop a scalable cloud data platform. The goal is to implement high-performance, maintainable, and cost-efficient data architectures for analytics and reporting use cases in an international environment. This is you - At least 4 years of practical experience with GCP in a data environment - BigQuery - Dataflow - Pub/Sub - Cloud Storage - Cloud Composer (or comparable orchestration) - Very good SQL skills - Very good Python skills (alternatively Java/Scala) - Experience in data modeling (e.g., Star Schema, Data Vault) - Experience with Infrastructure as Code (e.g., Terraform) is desirable - Independent, structured, and proactive Working Methods - Experience in international projects is an advantage This is your role - Design and implementation of modern data architectures on Google Cloud Platform - Building and optimizing scalable data pipelines (batch & streaming) - Developing data warehouse/lakehouse structures in BigQuery - Implementing and optimizing ETL/ELT processes - Performance and cost optimization within the GCP environment - Implementing data governance, security, and access concepts - Technical coordination with international stakeholders (English-speaking) What awaits you at Avenga? - Through our values, Better Minds, Bolder Ideas, and Bigger Hearts, we strive to provide you with the tools, autonomy, trust, and assistance you need to excel. Enjoy benefits like private health insurance, well-being programs, flexible and hybrid work models, laptops and gear, training, language classes, social events, great offices, and more. At Avenga, everyone matters. We provide equal opportunities in recruitment, career development, and leadership, regardless of race, ethnicity, gender identity, sexual orientation, disability, age, religion, or any other characteristic. We are committed to fostering a work environment where our diverse community of employees, candidates, and business partners actively shapes our growth. By bringing together people from different backgrounds and experiences, we build a workplace where everyone feels free to be themselves while honoring the boundaries of others.
DevOps Engineer, 3+ Years
Codvo.aiBuilding Advance AI & Cloud Native Software Using The "Virtual Silicon Valley" Model. Let’s Talk AI, Cloud and Outcomes.
• Lead a team of DevOps engineers responsible for ensuring the seamless operation of platforms and applications. • Implement and maintain infrastructure, automate deployments, and optimize systems for performance and reliability. • Champion DevOps best practices and drive continuous improvement within the team.
Junior Ops Analyst
ElasticSelf-described as the leading platform for search-powered solutions, Elastic helps organizations, their customers, and their employees find what they need faste
Role Description We are seeking for a Junior Operations Analyst – Professional Services to support the operational and analytical needs of our Professional Services organization. This role will report to Senior Operations Analyst and work closely with the Operations Director, Resource Managers, and cross-functional stakeholders to ensure accurate reporting, strong operational hygiene, and efficient execution across forecasting, resource planning, compliance, and month-end activities. This position is ideal for an early-career professional looking to develop skills in services operations, analytics, and global business support within a fast-paced, international environment. What you will be doing: - Operations & Business Support - Help document, maintain, and improve operational processes and workflows. - Maintain analytical reports covering services revenue, bookings, utilization and delivery mix (e.g., partner vs. internal consultants). - Support continuous improvement of business reporting processes to align with organizational changes. - Provide regular operational analysis to support business objectives and assist with services team enablement sessions (e.g., compliance processes, global updates, and services attach reporting). - Assist with improvements to operational management systems and tools. - Assist with reviewing and processing backlog, utilization, and forecast data. - PTO coverage for the Sr Ops Analyst. - Timecard & Utilization Management - Track and follow up on missing timecards for consultants and partners. - Monitor submitted but unapproved timecards and coordinate with Delivery Managers, Project Coordinators, or approvers to ensure timely approval. - Assist Sr Ops Analyst in reviewing Estimated vs. Actuals by comparing planned billable hours against recorded time. - Identify trends or significant variances, summarize findings, and escalate issues where appropriate. - Perform weekly checks to ensure required billable and non-billable hours are submitted (adjusted for public holidays). - Support weekly utilization reviews and reporting. - Project & PSA Operations - Review newly created projects to ensure accurate setup, assignment, and data quality (dates, milestones, region, currency, ownership). - Assist with project triage, including assigning projects to the appropriate Delivery Manager or Project Coordinator. - Monitor and report on: - Expired projects - Projects with zero hours logged - Project extensions - Support PSA contact onboarding, record maintenance, and data accuracy. - Verify targets and cost accuracy within PSA systems. - Month-End & Billing Support - Assist with pre-month-end and month-end close activities. - Audit non-billable milestones to ensure billable time is not incorrectly logged. - Support reporting of end-of-month (EOM) and end-of-quarter (EOQ) metrics. - Assist with setup and monthly review of Services Credits projects. - Review project milestone reports to ensure planned hours are not exceeded. - Resource Planning Support - Assist with Resource Planner activities, including creating internal assignments and identifying potential double bookings. - Assist Resource Managers with ad hoc administrative tasks related to the Resource Planner, including maintaining assignments, supporting public holiday planning, and ensuring data accuracy. - Support the creation and upkeep of resource management documentation like RM meeting planner playbooks, and process guides. - Support coordination between Operations and Resource Management teams to align staffing with forecasted demand. Qualifications - Bachelor’s degree in Business, Finance, Operations, Analytics, or a related field. - 3+ years of experience in Revenue Operations, Sales/Services Operations, Strategy, FP&A, Consulting, or similar roles within SaaS or Professional Services. - Proven experience driving operational performance in a regional or international context. - Advanced analytical expertise leveraging SQL for data extraction and transformation, paired with advanced Excel techniques (Power Query, pivot tables, complex formulas, and modeling) to deliver accurate forecasting, utilization analysis, and executive-level reporting. - Excellent stakeholder management and cross-functional collaboration skills. - Experience with Tableau or other BI tools. Benefits - Competitive pay based on the work you do here and not your previous salary. - Health coverage for you and your family in many locations. - Ability to craft your calendar with flexible locations and schedules for many roles. - Generous number of vacation days each year. - Increase your impact - We match up to $2000 (or local currency equivalent) for financial donations and service. - Up to 40 hours each year to use toward volunteer projects you love. - Embracing parenthood with a minimum of 16 weeks of parental leave.
DevOps Engineer
NSW GovernmentThe New South Wales (NSW) Government serves as the governing body for Australia’s most populous state, dedicated to delivering programs and services that enha
DevOps Engineer (Operations) Employment type: Ongoing Grade: Clerk Grade 7/8 salary range $113,574 - $125,720 + super Location: Sydney with weekly office attendance and hybrid working Job Description: About the Role The DevOps Engineer is part of the Development Operations team within the broader Platform, Development and Operations function. This role blends platform engineering, DevOps practices, and operational support, with a strong focus on Azure-based infrastructure, automation, and data platforms. You will contribute to the design, automation, support, and continuous improvement of cloud-native services while also playing a key role in operational support, ticket investigation, and platform reliability. This is an ideal opportunity for someone who is developing toward a DevOps Engineer level, with hands-on exposure to both engineering and production operations in a real-world environment. Key Responsibilities Operational Support, Ticketing & Investigation - Investigate, troubleshoot, and resolve Level 1-3 support tickets across: Azure platform services Data tools (ADF, Snowflake, Power BI, Tableau) - Perform incident management, root cause analysis, and system debugging - Support ticket-driven work across both platform and data environments - Contribute to improving operational processes and response times DevOps, Automation & CI/CD - Build, maintain, and enhance CI/CD pipelines using Azure DevOps and YAML pipelines - Automate infrastructure provisioning, deployments, and operational processes - Contribute to automation initiatives that improve platform efficiency and reduce manual work - Support and evolve DevOps best practices across the team Data Platforms & Workflows - Support and maintain data platforms and workflows - Assist with data pipeline reliability and performance - Work across data and platform boundaries to ensure end-to-end service delivery Monitoring, Maintenance & Reliability - Monitor system performance, availability, and security - Support patching, upgrades, and platform maintenance - Contribute to logging, monitoring, and alerting improvements - Maintain configuration records and documentation (e.g., CMDB) Collaboration & Technical Contribution - Work closely with teams across platform, development, and operations - Participate in agile delivery and contribute to shared outcomes - Provide support and guidance to junior team members where appropriate - Collaborate on cross-functional platform enhancements and projects About You You are a technically capable and motivated DevOps or Platform Engineer with a strong interest in cloud automation, infrastructure, and operational excellence. You are comfortable working across both engineering and support contexts, and you enjoy solving real world platform and data challenges. Essential Skills & Experience - Experience in DevOps, Platform Engineering, or Cloud Engineering roles - Strong exposure to Azure (preferred) or AWS/GCP - Experience with: CI/CD pipelines (Azure DevOps, YAML pipelines) Infrastructure as Code (IaC) (Bicep, Terraform, ARM) - Experience handling ticket-based operational support and investigation - Strong troubleshooting and problem-solving skills - Knowledge of cloud networking concepts (VNets, subnets, DNS, security groups) - Experience with scripting (PowerShell, Python, Bash) - Familiarity with version control systems (Git) - Strong communication and collaboration skills Nice to have - Experience with: Azure Data Factory (ADF), Snowflake, Power BI, Tableau Web applications and APIs Containerization (Docker, Kubernetes) - Exposure to: Cloud security practices Monitoring and logging tools - Software engineering experience (e.g., React, Express, TypeScript) - Experience supporting data workflows and analytics platforms Why Join Us? - Be part of a high-performing, supportive engineering team that values collaboration and knowledge sharing. - Work on interesting, technically challenging projects that genuinely matter. - Access ongoing learning and development opportunities, including exposure to cloud, DevOps, and emerging AI technologies. - Enjoy a flexible hybrid working environment with one day per week in the office. - Receive a competitive salary and superannuation package aligned to Clerk Grade 9/10. What We Need From You To apply, please attach your resume (max 5 pages) and cover letter (max 2 pages). In your cover letter, please share your motivation for applying and your relevant skills. Salary Grade 7/8, with the base salary for this role starting at $113574 base plus superannuation



