ML Infrastructure Engineer
Clera · San Mateo
- Location
- San Mateo
- Experience
- 5+ years
- Funding
- $3M
- Posted
- Aug 25, 2026
Clera is hiring a ML Infrastructure Engineer based in San Mateo. Every apply link on Engg.space goes straight to the company's own careers page - no recruiter middleman, no generic job-board form.
Apply directly at CleraRole details
ABOUT THE ROLE This is a hands-on ML Infrastructure Engineer role at an early-stage enterprise AI startup, where you'll own the end-to-end inference and model-serving infrastructure that keeps production AI agents running reliably and at scale. You'll sit at the intersection of ML and platform engineering, directly shaping the systems that power real-world, high-stakes deployments in regulated industries like insurance, banking, and healthcare. WHAT YOU'LL DO - Own inference and model-serving infrastructure end to end, from design through production deployment. - Build and scale systems that enable AI agents to run reliably and efficiently under increasing concurrency. - Collaborate closely with ML and infrastructure teams to ensure seamless integration and performance optimization. - Identify infrastructure bottlenecks and drive cross-functional solutions across engineering teams. WHAT WE'RE LOOKING FOR - 5+ years of experience building and operating ML inference systems, model-serving platforms, or ML infrastructure in production. - Hands-on experience designing and scaling inference-serving infrastructure using frameworks such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems. - Strong track record optimizing production ML systems for latency, throughput, and reliability at scale. - Experience with containerization and orchestration (Docker, Kubernetes) for deploying and scaling ML workloads. - Experience building distributed systems that handle concurrent requests and manage resource allocation under load. - Proficiency with observability and debugging tooling for production systems (e.g., Prometheus, Grafana, ELK, distributed tracing). - Cloud platform experience on AWS, GCP, or Azure for deploying and managing ML systems. - Proficiency in at least one systems or backend language — Python, Go, Rust, C++, or Java. - Nice to have: experience with knowledge graphs, semantic search, or graph databases (e.g., Neo4j, Amazon Neptune); real-time or low-latency inference systems; agentic or multi-step reasoning pipelines; enterprise data infrastructure or integration platforms. LOCATION On-site in San Mateo, CA. No visa sponsorship is available for this role.
More roles at Clera
- Software Engineeryesterdayremote$75,000 to $265,000 USD$3M
- Founding Engineer (Full Stack)yesterdaySan Francisco$100,000 to $180,000 USD$3M
- Founding Full-Stack iOS EngineeryesterdayLondon$3M
- Software Engineer - AIyesterdayBerlin$3M
- Applied AI Engineer - AI & AutomationyesterdayBerlin$3M
- Full Stack Software EngineeryesterdayNew York$120,000 to $200,000 USD$3M
- Data Engineer2 days agoBerlin$3M
- Founding Engineer2 days agoNew York$170,000 to $220,000 USD$3M
- Founding Engineer2 days agoBerlin$3M
- Founding Engineer2 days agoSan Francisco$3M
- Founding AI Engineer2 days agoAmsterdam$225,000 to $255,000 USD$3M
- Founding Engineer, Infrastructure2 days agoSan Francisco$175,000 to $300,000 USD$3M