Cisco
US - California - San Jose
View Company Profile /
<< Go Back
Own application-level reliability for AI-powered features by defining SLIs, SLOs, and error budgets, building observability dashboards, and analyzing usage and operational data. Develop LangGraph diagnostic and remediation agents, evaluation tooling, and Python automation while leading application-focused incident response and partnering with engineering and SRE teams.
Requirements: Requires 7+ years of software engineering experience focused on reliability, observability, or production operations, plus a bachelor's or master's degree in a relevant technical discipline. Candidates need strong Python and SQL skills, production GCP experience with GKE, BigQuery, and Bigtable, and experience operating application-level SLI/SLO frameworks and debugging distributed applications.
Key Skills: Python, Kubernetes, Google Cloud Platform, BigQuery, Bigtable, SQL, LangGraph, Looker, Application Reliability, Observability, Site Reliability Engineering, SLI/SLO Frameworks, Distributed Tracing, Incident Response, Agent Evaluation, Generative AI
Benefits: Medical Insurance, Dental Insurance, Vision Insurance, 401(k) With Employer Matching, Paid Parental Leave, Short-Term Disability Coverage, Long-Term Disability Coverage, Life Insurance, Restricted Stock Units, Paid Holidays, Floating Holiday, Birthday Leave, Year-End Holiday Shutdown, Personal Wellness Leave, Paid Vacation, Sick Leave, Family Emergency Leave, Volunteer Leave, Annual Bonus
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX