STAFF SOFTWARE ENGINEER, Frontier Security Team
Snowflake · US-CA-Menlo Park
- Location
- US-CA-Menlo Park
- Salary
- $236K - $295K
- Funding
- Public Company
- Posted
- Sep 16, 2026
Snowflake is hiring a STAFF SOFTWARE ENGINEER, Frontier Security Team based in US-CA-Menlo Park. Every apply link on Engg.space goes straight to the company's own careers page - no recruiter middleman, no generic job-board form.
Apply directly at SnowflakeRole details
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. At Snowflake, we are powering the era of the agentic enterprise. We seek AI-native engineers who are curious about emerging capabilities, rigorous about measured results, and motivated to redefine how software is built and operated. Snowflake’s Frontier Security AI teams develop production-grade LLM applications, intelligent agents, AI infrastructure, and evaluation systems for enterprise customers. Our products must meet a high bar for quality, security, reliability, and efficiency while operating over sensitive data at large scale. ABOUT THE ROLE We are looking for a Staff Software Engineer to lead the design and development of our Agentic Harness and agent evaluation platform. The Agentic Harness provides the runtime, tools, context, state, policies, and observability required to build and operate production AI agents. The evaluation platform measures how well those agents complete real customer tasks and detects regressions before they reach production. This is a hands-on technical leadership role. You will build production systems, establish architecture across team boundaries, and define the metrics and engineering practices used to improve agent quality. You will work with product, infrastructure, applied AI, security, and modeling teams to take new AI capabilities from prototype to dependable customer value. WHAT YOU WILL DO IN THIS ROLE • Architect and build the Agentic Harness that executes complex, multi-step AI workflows across models, tools, data, and services. • Design stable interfaces for tool execution, context construction, state management, memory, permissions, retries, fallbacks, and human review. • Own agent quality end to end by building evaluation harnesses, representative datasets, automated graders, experiment pipelines, and release gates. • Convert ambiguous reports such as “the agent feels worse” into measurable failure modes, reproducible tests, and durable fixes. • Analyze production agent trajectories to identify failures in reasoning, retrieval, tool use, context, orchestration, and application code. • Close the loop between production incidents, root-cause analysis, evaluation coverage, and regression prevention. • Develop offline and online measurements for task completion, correctness, groundedness, safety, latency, reliability, and cost. • Build simulation and replay infrastructure for golden-set tests, adversarial scenarios, model comparisons, and large-scale experiments. • Improve agent efficiency through model routing, prompt and semantic caching, context compaction, tool-result management, and token optimization. • Productionize new model capabilities as secure, observable, multi-tenant services with clear operational controls. • Establish standards for evaluation design, including sampling, ground-truth quality, grader calibration, leakage prevention, and statistical significance. • Define technical direction across multiple teams and lead projects whose scope extends beyond a single service. • Mentor engineers, raise the quality of architecture reviews, and remain directly involved in implementation and debugging. REQUIREMENTS • 9+ years of software engineering experience, including technical leadership of complex production systems. • Direct experience shipping and operating LLM applications, AI agents, or model-backed workflows in production. • Strong background in distributed systems, service architecture, high-throughput APIs, concurrency, and failure handling. • Experience building an agent runtime, workflow engine, developer platform, evaluation system, or similar infrastructure. • Demonstrated ability to evaluate nondeterministic systems without relying on a single aggregate score. • Fluency in Python and strong proficiency in at least one systems or application language such as Java, Go, Rust, or TypeScript. • Hands-on knowledge of tool calling, structured generation, retrieval, context engineering, prompt management, and model APIs. • Experience with production observability, including structured traces, replay, metrics, logs, and incident diagnosis. • Ability to balance agent quality with latency, reliability, security, and inference cost. • Track record of setting technical direction and delivering results across organizational boundaries. • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience. • Clear written and verbal communication with engineering, product, and leadership audiences. BONUS EXPERIENCE • Building evaluation or observability infrastructure for agentic coding, data engineering, or analytics systems. • Designing human-evaluation programs, scoring rubrics, annotation workflows, or grader-calibration methods. • Working with multi-agent orchestration, long-running agents, asynchronous workflows, or durable execution. • Developing synthetic tasks, simulations, adversarial tests, red-team exercises, or safety guardrails. • Building retrieval systems that use vector search, hybrid search, semantic indexing, ranking, or caching. • Operating multi-tenant systems that process sensitive enterprise data. • Working with model training, fine-tuning, reinforcement learning, or feedback-driven optimization. • Evaluating and onboarding frontier models based on measured product outcomes. • Exp
More roles at Snowflake
- Senior Software Engineer - Security Foundations5 days agoUS-WA-Bellevue$200K - $287.5KPublic Company
- Software Engineer5 days agoPL-Warsaw-Lixa CPublic Company
- Engineering Manager6 days agoUS-CA-Menlo Park$236,000 - $339,250Public Company
- Principal Software Engineer - Semantic Views6 days agoUS-CA-Menlo Park$264K - $379.5KPublic Company
- Senior/Staff Software Engineer – LLM Inference & Reinforcement Learning Platform 6 days agoUS-WA-Bellevue$236K - $330KPublic Company
- Principal Software Engineer - Performance Engineering (Cloud Infrastructure)04-09-2026US-CA-Menlo Park$264K - $379.5KPublic Company
- Solutions Architect, Frontier Engineering03-09-2026JP-TokyoPublic Company
- Senior Software Engineer - Cloud Security02-09-2026US-CA-Menlo Park$200K - $287.5KPublic Company
- Solutions Architect, Observe02-09-2026US, Remote$160K - $210KPublic Company
- Senior Software Engineer - Cloud Efficiency02-09-2026US-WA-Bellevue$200K - $287.5KPublic Company
- Senior AI/ML Architect, Applied Field Engineering02-09-2026US-GA-Atlanta$165,000 - $216,562Public Company
- Staff Software Engineer - Query Processing (Snowtrail)01-09-2026US-CA-Menlo Park$236K - $339.2KPublic Company