Senior Site Reliability Engineer

Amtechsoftware · India

Location
India
Funding
N/A
Posted
Sep 8, 2026

Amtechsoftware is hiring a Senior Site Reliability Engineer based in India. Every apply link on Engg.space goes straight to the company's own careers page - no recruiter middleman, no generic job-board form.

Apply directly at Amtechsoftware

Role details

About Vista Equity Partners Vista Equity Partners is a leading global investment firm focused exclusively on enterprise software, data, and technology-enabled businesses. With over $100B in assets under management and a portfolio of 90+ software product companies worldwide, Vista accelerates growth through operational excellence, shared expertise, and long-term partnership. In India, Vista’s presence continues to expand with 45+ portfolio companies employing more than 17,000 professionals across technology, product, customer success, and operations — reinforcing India’s strategic role as a hub of innovation and talent within the Vista ecosystem. Through its Agentic AI Factory, Vista is embedding Generative AI across its global portfolio — enabling companies to integrate intelligent, responsible AI into products, operations, and decision-making. This initiative is strengthened through portfolio-wide learning programs, leadership workshops, and AI hackathons that foster innovation, build fluency, and accelerate practical AI adoption across teams. About Amtech Amtech is a leading provider of enterprise software solutions for the packaging, printing, and manufacturing industries. Our integrated systems streamline order management, production planning, scheduling, inventory, and business analytics — empowering customers to drive efficiency, reduce costs, and improve operational performance. With a strong commitment to innovation and customer success, Amtech delivers reliable technology backed by deep industry expertise. With Vista’s investment and strategic guidance, we combine the agility of a growing technology organization with the scale, stability, and career mobility of a global software ecosystem. Our Employee Value Proposition At Amtech, our people are our greatest differentiator. We create an environment where you can: Purpose Shape the future of manufacturing and supply chain operations by delivering mission-critical enterprise software used by industry-leading organizations. Growth Access continuous learning, leadership development, and cross-portfolio opportunities through Vista’s global network — accelerating both technical and managerial career paths. Culture Work in a collaborative, transparent, and people-first environment where values, accountability, and integrity guide every decision. Innovation Engage with cutting-edge technologies, including AI-driven automation, and contribute to modernizing financial systems and operational processes across the business. Role Description Amtech is scaling its Platform Engineering organization as our products move to a fully AWS-hosted, multi-tenant SaaS model. Our SRE practice keeps Encore, LabelTraxx, and supporting platforms reliable, observable, and secure across a multi-account AWS estate. As a Senior Platform Engineer on the SRE track, you will own reliability for major production domains end to end: defining SLOs, building the observability and automation that defend them, and leading incident command for the most complex events. You will shape Amtech's global reliability standards, mentor the SRE team, and partner with Cloud-track engineers to make new platforms operable from day one. KEY RESPONSIBILITIES Reliability & Performance Own SLIs, SLOs, and error budgets for major production domains, and drive engineering priorities from them. Design the monitoring, alerting, and reliability frameworks the global team standardizes on (OpenTelemetry, OpenObserve, CloudWatch, PagerDuty). Engineer capacity management, autoscaling, and self-healing so systems recover without human intervention. Lead root cause analysis for the highest-severity incidents and verify permanent fixes land. Automation & Platform Engineering Design reusable Terraform modules, deployment patterns, and progressive delivery (blue/green, canary, automated rollback) in GitHub Actions. Operate and optimize ECS Fargate, EKS, Lambda, and RDS PostgreSQL workloads at production scale. Set the toil-reduction agenda: quantify operational load and eliminate it through engineering. Build reliability into the Encore-on-AWS and LabelTraxx platforms as they scale customer counts. Incident Response & On-Call Serve as senior incident commander for cross-service, customer-impacting incidents. Own the on-call program's health for your domains: escalation policies, alert quality, and rotation sustainability in PagerDuty. Drive game days and failure testing to validate runbooks and recovery paths. Security & Compliance Engineer security into reliability tooling: IAM boundaries, secrets, and network controls. Ensure operations satisfy SOC 2 and ISO 27001 obligations with automated audit evidence. AI Competency Apply AI in operations and incident response with risk tiering: AI assistance for triage, analysis, and hypothesis generation; human gates for production changes. Design validation gates and guardrails for AI-assisted operational changes, including rollback and audit trails. Measure whether AI tooling improves reliability work against baselines (MTTR, rework, defect escape), and enforce data classification policy in all AI-assisted work. Technical Leadership Mentor SRE I-III engineers through design reviews, paired incident response, and career coaching input. Set and document global reliability standards adopted across U.S. and India teams. Represent reliability in architecture reviews and migration planning. QUALIFICATIONS 6+ years of hands-on SRE, DevOps, or platform engineering experience, including ownership of customer-facing production services. Deep AWS operational expertise (ECS/EKS, Lambda, RDS, IAM, VPC, multi-account organizations). Strong Terraform and CI/CD (GitHub Actions or equivalent) skills, including progressive delivery patterns. Proven software engineering ability in Python (or similar) for automation and tooling at team scale. Demonstrated incident command experience for high-severity, multi-service incidents. Experience designing and operating SLO-driven obse

More roles at Amtechsoftware

Search all engineering roles