Workday logo

Principal, Software Engineer (Distributed Systems)

Workday · USA CA Pleasanton

Location
USA CA Pleasanton
Experience
12+ years
Funding
Public Company • not captured
Posted
Sep 23, 2026

Workday is hiring a Principal, Software Engineer (Distributed Systems) based in USA CA Pleasanton. Every apply link on Engg.space goes straight to the company's own careers page - no recruiter middleman, no generic job-board form.

Apply directly at Workday

Role details

Your work days are brighter here. We’re obsessed with making hard work pay off, for our people, our customers, and the world around us. As a Fortune 500 company and a leading AI platform for managing people, money, and agents, we’re shaping the future of work so teams can reach their potential and focus on what matters most. The minute you join, you’ll feel it. Not just in the products we build, but in how we show up for each other. Our culture is rooted in integrity, empathy, and shared enthusiasm. We’re in this together, tackling big challenges with bold ideas and genuine care. We look for curious minds and courageous collaborators who bring sun-drenched optimism and drive. Whether you're building smarter solutions, supporting customers, or creating a space where everyone belongs, you’ll do meaningful work with Workmates who’ve got your back. In return, we’ll give you the trust to take risks, the tools to grow, the skills to develop and the support of a company invested in you for the long haul. So, if you want to inspire a brighter work day for everyone, including yourself, you’ve found a match in Workday, and we hope to be a match for you too. About the Team Data Platform and Observability team is based in Pleasanton and Atlanta in the US, Dublin in Ireland and Chennai in India. Our focus is on the development of large scale distributed data systems to support critical Workday products and provide real-time insights across Workday’s platforms, infrastructure and applications. The team provides platforms that process 100s of terabytes of data that enable core Workday products and use cases like core HCM, Fins, AI/ML skus, internal data products and Observability. If you enjoy writing efficient software or tuning and scaling large distributed systems you will enjoy working with us. Do you want to tackle exciting challenges at massive scale across private and public clouds for our 10000+ global customers? Do you want to work with world class engineers and facilitate the development of the next generation Distributed systems platforms? If so, we should chat. About the Role The Messaging, Streaming and Caching team is a full-service Distributed Systems Engineering team. We architect and provide async messaging, streaming, and NoSQL platforms and solutions that power the Workday products and SKUs ranging from core HCM, Fins, Integrations, and AI/ML. We develop client libraries and SDK’s that make it easy for teams to build Workday products. We develop automation to deploy and run hundreds of clusters, and we also operate and tune our clusters as well. As a team member you will play a key role in improving our services and encouraging their adoption within Workday's infrastructure both in our private cloud and public cloud. As a member of this team you will design and build new capabilities from inception to deployment to exploit the full power of the core middleware infrastructure and services, and work hand in hand with our application and service teams! Primary Responsibilities Design, build, and enhance critical distributed services, including Kafka, Redis, RabbitMQ etc. Design, develop, build, deploy and maintain core distributed services using a combination of open source and proprietary stacks across diverse infrastructure environments (Kubernetes, OpenStack, Bare Metal, etc.) Design and develop core software modules for streaming, messaging and caching. Build observability modules, alerts and automation for Dashboard lifecycle management for the distributed services. Build, deploy and operate infrastructure components in production environments. Champion all aspects of streaming, messaging and caching with a focus on resiliency and operational excellence. Evaluate and implement new open-source and cloud-native tools and technologies as needed. Participate in the on-call rotation to support the distributed systems platforms. Manage and optimize Workday distributed services in AWS, GCP & Private cloud env. About You Basic Qualification 12+ years experience in software development engineering. 6+ years experience specifically focused on designing, building, and operating distributed systems like Redis, Kafka, RabbitMQ or NoSQL solutions. 5+ experience in designing and implementing complex distributed system architectures, evidenced by successful deployment of systems with high availability (e.g., 99.99% uptime) and fault tolerance. 8+ years experience with at least two of the following programming languages (e.g., Java, Python, Go, C/C++), including experience in writing production-level code for distributed systems. Expertise with configuration management using Chef and service deployment on Kubernetes via Helm and ArgoCD Bachelor’s degree in a relevant field such as Computer Science, Engineering, or a related discipline; a Master's degree (e.g., MS in Computer Science, Distributed Systems, or related field) is strongly preferred or equivalent practical experience. Other Qualification Expert-level ability in Algorithmic Thinking, including CAP theorem, queuing theory, consensus protocols etc, to architect highly efficient and scalable solutions for complex distributed systems implementations. Deep expertise in API and Client Library Development, including understanding of RESP protocol, Kafka wire protocol etc and extensive experience in designing and building API layer as well as client libraries. Experience building cloud native controllers for distributed systems. Familiarity with operator like Strimzi would be a bonus Strong understanding of modern Code Testing methodologies like consistency / linearizability testing, and experience in leading chaos and and fault injection testing strategies. Deep understanding of Distributed Systems Software principles, including fault tolerance, high availability, and extensive experience in replication / sharding techniques Proven ability to design and implement High Availability strategies for critical distributed systems, including global rep

More roles at Workday

Search all engineering roles