Senior Software Engineer I (Observability Platform)
Smartsheet · -REMOTE, USA-
- Location
- -REMOTE, USA-
- Salary
- $161,250 - $193,750 USD
- Funding
- N/A
- Posted
- Sep 24, 2026
Smartsheet is hiring a Senior Software Engineer I (Observability Platform) based in -REMOTE, USA-. Every apply link on Engg.space goes straight to the company's own careers page - no recruiter middleman, no generic job-board form.
Apply directly at SmartsheetRole details
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Smartsheet’s Observability Engineering team owns how the company sees itself: the collection, modeling, storage, and analysis of metrics, logs, distributed traces, and events across a global, multi-region infrastructure. As a Senior Software Engineer I on this team, you will be hands-on-keyboard building that platform (instrumentation libraries, telemetry pipelines, SLO and alerting systems, dashboards as code), and you will be the engineer who connects it to everything around it. Observability only pays off when signals flow into the systems where engineers actually work: CI/CD, the service catalog, incident management, ticketing, chat, and the automated remediation that closes the loop without waking anyone up. We run observability as an internal product, and the engineers of Smartsheet are its users. That means published interfaces, versioned client libraries, honest deprecation paths, and SLOs on the platform itself. It also means we measure ourselves on adoption rather than on components shipped: a beautifully engineered tracing pipeline that three teams use is a failure. You will own a slice of that product end to end, including the golden path that makes correct instrumentation the easy choice, the documentation and hands-on sessions that get teams there, and the numbers that tell you which part to fix next. This is a build role, not a configure role, and it is engineering first. You will write the services, the integrations, and the automation that turn a collection of separate tools into one coherent platform, and you will own the architecture decisions that make it hold up as it grows. Where tooling and teaching compete for the same problem, we prefer the tool: a CI check that rejects an unlabeled metric works for every team forever, while a workshop works for the people in the room. You will report to the Team Lead, Observability Engineering, and partner closely with the Principal Engineer setting telemetry data platform direction. You Will Architect and build the observability platform: Design and ship end-to-end capability for metrics, logs, distributed traces, and events across multiple regions and environments, owning components from collection through storage, query, and presentation Run observability as an internal product: Build for the engineers who depend on you, with published interfaces, versioned client libraries, clear deprecation paths, and SLOs on the platform itself, so teams rely on it the way they rely on any production service Drive instrumentation with OpenTelemetry: Build and maintain shared instrumentation libraries, collector deployments, semantic conventions, context propagation, and sampling strategies so service teams get correlated signals by default rather than by effort Engineer telemetry pipelines at scale: Build high-volume collection, enrichment, redaction, and routing pipelines with the reliability, backpressure handling, and tiered retention that multi-region, high-cardinality traffic requires Connect the platform end to end: Integrate the observability toolchain with CI/CD, service catalog, incident management, ticketing, chat, and feature-flag systems through REST APIs, webhooks, and event-driven services, so that a deploy, an alert, an incident, and a ticket form one continuous thread instead of four disconnected ones Build automated and self-healing remediation: Design event-driven and agent-assisted workflows that detect, diagnose, and resolve known failure modes automatically, with human-in-the-loop approval gates as the safety mechanism for anything consequential Make reliability measurable: Implement SLOs, error budgets, and golden-signal alerting as code, and drive down alert noise so that a page means something is genuinely wrong Build the golden path: Own the paved road for instrumenting a new service, including scaffolding and templates, versioned Terraform modules, and CI checks that catch missing or malformed telemetry before merge. Make the correct path the easy path, so coverage comes from good defaults rather than from chasing teams Instrument the platform itself: Track coverage, onboarding time, time to first useful dashboard, query performance, and cost per service, and build the guardrails that keep cardinality, sampling, and retention proportional to the value of the signal. Let those numbers decide what you build next instead of guessing Build the enablement layer: Write documentation as code, reference architectures, and worked examples that scale past the conversations you can personally have, build the onboarding path that takes a team from zero to instrumented without a meeting, and run the workshops, office hours, and game days that exercise dashboards and alerts under realistic failure Raise the technical bar: Lead code reviews and architecture discussions, author the instrumentation standards other teams build against, mentor engineers on signal design and cost-aware instrumentation, and grow a group of instrumentation champions who carry the practice inside their own teams Apply AI where it earns its place: Use AI tooling to improve your own and the team’s efficiency across coding, testing, design, and troubleshooting, and help instrument Smartsheet’s AI and agentic systems so their behavior is as observable as any other service Turn incidents into durable improvements: Join the team’s on-call rotation, drive root-cause analysis, and close every incident with an instrumentation change, an automation, or a documented lesson that reaches the teams who need it You Have 5+
More roles at Smartsheet
- Principal AI Engineer - Hybrid in Bangalore6 days agoBangalore, INDIA
- Manager, People Systems Engineering (Remote Eligible)17-09-2026 -REMOTE, USA-$122,000 - $175,000 USD
- Sr. Software Engineer II - Billing & Subscriptions Engineering (Remote Eligible)15-09-2026 -REMOTE, USA-$175,000 - $245,000 USD
- Staff Engineer - Databricks 15-09-2026Bangalore, INDIA
- Solutions Architect03-09-2026-REMOTE, USA-$115,000 - $152,500 USD
- Sr. Software Engineer II - AI Engineering (Remote Eligible) 03-09-2026 -REMOTE, USA-$175,000 - $245,000 USD
- Principal Software Engineer - Observability & Telemetry Data01-09-2026Bellevue, WA, USA$222,500 - $257,500 USD
- Principal Software Engineer (Hybrid in Bangalore)24-08-2026Bangalore, INDIA