Senior Site Reliability Engineer
Autodesk · Croatia, EMEA
- Location
- Croatia, EMEA
- Experience
- 7+ years
- Funding
- ~$48.8B
- Posted
- Oct 6, 2026
Autodesk is hiring a Senior Site Reliability Engineer based in Croatia, EMEA. Every apply link on Engg.space goes straight to the company's own careers page - no recruiter middleman, no generic job-board form.
Apply directly at AutodeskRole details
Job Requisition ID # 26WD101322 Senior Site Reliability Engineer, Reliability Automation Position Overview Autodesk is the global leader in design and make technology, including industry-leading 3D design, engineering, and entertainment software and services, that offer customers better outcomes through automation and insights for their design and make processes. If you’ve ever driven a high-performance car, admired a towering skyscraper, used a smartphone, or watched a great film, chances are you’ve experienced what millions of Autodesk customers are doing with our software. Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk products and customers. As part of an SRE team focused on Reliability Automation, you will have a unique opportunity to help shape how Autodesk detects, responds to, and engineers away production issues at scale. This is a project-oriented, engineering-heavy role where you will build the automation, tooling, and platform capabilities that other engineering teams rely on to run their services reliably. You will combine software engineering and production operations to automate how Autodesk services are operated, monitored, recovered, and validated. You will partner closely with product engineering, platform, infrastructure, and operations teams to turn manual operational work into durable, reusable automation. The ideal candidate is a strong software engineer who is drawn to operational problems. You should be comfortable designing and shipping production-grade software, and equally comfortable reasoning about how distributed systems fail. Success in this role requires strong technical judgment, the ability to drive projects end to end, and a passion for replacing manual operational work with durable engineering. Responsibilities Lead reliability automation projects end to end, from problem definition and design through delivery, adoption, and measurable impact on availability, performance, and operational efficiency. Design, develop, and maintain shared automation, tooling, and platform services that improve the reliability, scalability, and operability of production systems across Autodesk. Partner with engineering teams to understand their operational pain, and drive adoption of the automation and reliability capabilities your team builds. Instrument and automate reliability signals such as SLOs/SLIs, error budgets, and alerting-as-code, so teams get reliability measurement by default rather than by manual effort. Build automation that improves deployment safety, operational efficiency, and service recovery, including self-healing and auto-remediation capabilities. Automate incident response workflows, diagnostics, and runbook execution to reduce time to detect, time to mitigate, and manual intervention. Extend monitoring, alerting, logging, and tracing capabilities, and automate their coverage and consistency across services. Own the reliability, availability, and operability of the automation and platform services your team runs, including on-call response when they fail. Turn incident and post-incident findings from across the organization into automation, tooling, and durable engineering fixes. Design and scale automated resilience testing, chaos engineering, Gameday, and disaster recovery tooling to validate system behavior and recovery capabilities. Continuously identify and eliminate operational toil through software engineering, automation, and process improvement. Ensure the automation and services your team builds meet Autodesk security, privacy, and operational risk requirements. Participate in a 24x7 on-call rotation for the automation, tooling, and platform services owned by the team. Function effectively in a fast-paced environment while helping mature reliability automation practices across engineering. Basic Qualifications B.S. or higher in Computer Science, Engineering, or a related technical discipline, or equivalent practical experience. 7+ years of experience in Software Engineering, Site Reliability Engineering, Platform Engineering, or a closely related discipline. Strong software engineering fundamentals, including API and service design, testing, code review, and shipping production-grade software. Working knowledge of reliability engineering concepts such as SLOs/SLIs, observability, incident response, and toil reduction, and experience applying them in practice. Experience building on AWS, Azure, or another public cloud platform. Strong programming skills in languages such as Python, Go, Java, or similar. Experience with Infrastructure as Code, CI/CD pipelines, and deployment automation. Ability to work independently, scope ambiguous problems, and drive projects to completion across team boundaries. Strong written and verbal communication skills. Preferred Qualifications 7+ years of experience building software and automation for production systems. Experience building self-healing, auto-remediation, or event-driven operational automation. Experience building internal platforms, developer tooling, or services adopted by other engineering teams. Experience building or adopting AI-assisted operational tooling, such as LLM-based incident triage, diagnostics, or agentic remediation workflows. Experience with containers, Kubernetes, cloud-native architectures, APIs, and distributed systems. Experience integrating with observability platforms such as Splunk, Dynatrace, Datadog, CloudWatch, or similar. Experience with event-driven architectures, workflow orchestration, or job scheduling systems. Experience with incident management platforms such as incident.io, FireHydrant, or PagerDuty, and automating workflows on top of them. Experience designing and implementing operational automation at scale. Experience building or automating Gamedays, chaos experiments, disaster recovery exercises, or resilience testing. Experience
More roles at Autodesk
- Software Engineer6 days agoSingapore, SGP~$48.8B
- Senior Developer Advocate (API's)30-09-2026Bengaluru, IND~$48.8B
- Software Development Engineer30-09-2026Singapore, SGP~$48.8B
- Senior Data Engineering30-09-2026Toronto, ON, CAN~$48.8B
- Senior Software Developer30-09-2026Alberta CAN Remote~$48.8B
- Senior Software Developer (Full stack)30-09-2026Toronto ON CAN~$48.8B
- Principal Full Stack Engineer24-09-2026AMER United States Oregon Portland~$48.8B